Text recognition method and related apparatus, device, and storage medium

By combining visual and linguistic feature decoding methods during text recognition and dynamically updating the decoding state, the problem of inaccurate text recognition on OOV is solved, and the recognition accuracy is improved.

CN116935404BActive Publication Date: 2026-04-10IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-09
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing text recognition technologies perform poorly on OOV (Out Of Vocabulary) text, making it difficult to improve accuracy.

Method used

By extracting image features from the image to be identified, combining visual and linguistic features for decoding, and utilizing techniques such as attention mechanisms and long short-term memory networks, the decoding state and features are dynamically updated to achieve fusion decoding of vision and language.

Benefits of technology

It improves the accuracy of text recognition, especially in cases of out-of-vocabulary (OOV) text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935404B_ABST
    Figure CN116935404B_ABST
Patent Text Reader

Abstract

The application discloses a text recognition method and related device, equipment and storage medium, wherein the text recognition method comprises: extracting image features of an image to be recognized; performing the following first decoding operation based on the image features: extracting first visual features of a current decoding moment from the image features based on decoding information of a previous decoding moment; and obtaining language features of the current decoding moment based on the first visual features of the current decoding moment and the decoding information of the previous decoding moment; and decoding based on the first visual features and the language features to obtain decoded characters of the current decoding moment; wherein the decoding information comprises at least one of the decoded characters and a decoding state, and the decoded characters of each decoding moment are combined to obtain candidate recognized texts of the first decoding operation; and based on the candidate recognized texts of several decoding operations, target recognized texts of the image to be recognized are obtained. The above scheme can improve the accuracy of text recognition, especially the accuracy of OOV.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of visual understanding, and in particular to a text recognition method and related device, equipment and storage medium. BACKGROUND

[0002] With the rapid development of information technology, text recognition technology has been applied in more and more scenarios.

[0003] However, the existing text recognition technology usually performs poorly on OOV (Out Of Vocabulary). However, in real-world scenarios, OOV words are common and very important, such as place names, company names, and URLs (Uniform Resource Locator). Therefore, how to improve the accuracy of text recognition, especially on OOV, has become a problem to be solved. SUMMARY

[0004] The technical problem solved by the present application is to provide a text recognition method and related device, equipment and storage medium, which can improve the accuracy of text recognition, especially on OOV.

[0005] To solve the above technical problem, the first aspect of the present application provides a text recognition method, comprising: extracting image features of a to-be-recognized image; wherein the image features are used for a plurality of decoding operations, and the plurality of decoding operations at least include a first decoding operation; performing the following first decoding operation based on the image features: extracting first visual features of a current decoding time from the image features based on decoding information of a previous decoding time; and obtaining language features of the current decoding time based on the first visual features of the current decoding time and the decoding information of the previous decoding time; and decoding based on the first visual features and the language features to obtain decoded characters of the current decoding time; wherein the decoding information includes at least one of the decoded characters and a decoding state, and the decoded characters of each decoding time are combined to obtain a candidate recognition text of the first decoding operation; and obtaining a target recognition text of the to-be-recognized image based on the candidate recognition texts of the plurality of decoding operations.

[0006] To solve the above technical problems, the second aspect of the present application provides a text recognition device, comprising: an encoding module, a decoding module and a determination module, the encoding module is used for extracting image features of an image to be recognized; wherein the image features are used for a plurality of decoding operations, and the plurality of decoding operations at least include a first decoding operation; the decoding module is used for performing the following first decoding operation based on the image features: extracting first visual features of a current decoding time from the image features based on decoding information of a previous decoding time; and obtaining language features of the current decoding time based on the first visual features of the current decoding time and the decoding information of the previous decoding time; and decoding based on the first visual features and the language features to obtain decoded characters of the current decoding time; wherein the decoding information includes at least one of the decoded characters and a decoding state, and the candidate recognition text of the first decoding operation is obtained by combining the decoded characters of each decoding time; the determination module is used for obtaining target recognition text of the image to be recognized based on the candidate recognition text of each decoding operation of the plurality of decoding operations.

[0007] To solve the above technical problems, the third aspect of the present application provides an electronic device, comprising a memory and a processor coupled to each other, the memory stores program instructions, and the processor is used to execute the program instructions to realize the text recognition method of the first aspect.

[0008] To solve the above technical problems, the fourth aspect of the present application provides a computer readable storage medium, which stores program instructions capable of being executed by a processor, and the program instructions are used to realize the text recognition method of the first aspect.

[0009] The above scheme extracts image features of an image to be recognized, and the image features are used for a plurality of decoding operations, the plurality of decoding operations at least include a first decoding operation, and then the following first decoding operation is performed based on the image features: first visual features of a current decoding time are extracted from the image features based on decoding information of a previous decoding time, and language features of the current decoding time are obtained based on the first visual features of the current decoding time and the decoding information of the previous decoding time, and decoding is performed based on the first visual features and the language features to obtain decoded characters of the current decoding time, and the decoding information includes at least one of the decoded characters and a decoding state, and the candidate recognition text of the first decoding operation is obtained by combining the decoded characters of each decoding time, and finally the target recognition text of the image to be recognized is obtained based on the candidate recognition text of each decoding operation of the plurality of decoding operations. Since the visual features and the language features are modeled based on the image features at each decoding time and combined for decoding, the features information in the visual dimension and the language dimension can be combined for decoding, thereby helping to improve the accuracy of text recognition, especially the accuracy of OOV. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1is a flowchart of an embodiment of the text recognition method of the present application.

[0011] Figure 2 is a process diagram of an embodiment of the first decoding operation.

[0012] Figure 3 is a process diagram of an embodiment of the text recognition method of the present application.

[0013] Figure 4 is a process diagram of an embodiment of the text recognition model training.

[0014] Figure 5 is an effect diagram of an embodiment of the text recognition method of the present application.

[0015] Figure 6 is a framework diagram of an embodiment of the text recognition apparatus of the present application.

[0016] Figure 7 is a framework diagram of an embodiment of the electronic device of the present application.

[0017] Figure 8 is a framework diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION

[0018] The scheme of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0019] In the following description, specific details such as specific system structures, interfaces, techniques, etc. are presented in order to provide a thorough understanding of the present application for the sake of explanation, but not for the purpose of limiting the present application.

[0020] The terms "system" and "network" are often used interchangeably herein. The term "and / or" herein is merely an associative relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which means that there can be three cases: A alone, A and B together, and B alone. In addition, the segment " / " herein generally means that the associated objects before and after are an "or" relationship. In addition, "multiple" herein means two or more than two.

[0021] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the text recognition method of the present application.

[0022] Specifically, it can include the following steps:

[0023] Step S11: Extracting image features of the image to be recognized.

[0024] In the embodiments of the present disclosure, the image features are used for several decoding operations, and the several decoding operations at least include a first decoding operation. It should be noted that the several decoding operations can refer to one decoding operation (i.e., the aforementioned first decoding operation), two decoding operations (including the aforementioned first decoding operation), or three or more decoding operations (including the aforementioned first decoding operation), and the number of decoding operations is not limited herein.

[0025] In one implementation scenario, the to-be-recognized image can contain text. Specifically, according to different actual application scenarios, the to-be-recognized image can also be different. For example, in a retail scenario, the to-be-recognized image can be obtained by shooting the goods on the shelf, that is, the to-be-recognized image contains text such as the name of the goods on the packaging of the goods (for example, XX mineral water, XX potato chips); or in a financial scenario, the to-be-recognized image can be obtained by shooting a voucher such as an invoice, that is, the to-be-recognized image contains text such as the name of the voucher on the voucher (for example, fixed-rate invoice, value-added tax invoice); or in a conference scenario, the to-be-recognized image can be an image formed by the user writing on the terminal, that is, the to-be-recognized image contains text such as handwritten words (for example, conference highlights, to-do list). It should be noted that the above examples are only several possible cases of the to-be-recognized image in the actual application process, and do not limit the specific information contained in the to-be-recognized image.

[0026] In one implementation scenario, in the embodiments of the present disclosure, the text recognition model can be used to perform text recognition on the to-be-recognized image to obtain the target text in the to-be-recognized image. The text recognition model can include an encoding network for extracting image features from the to-be-recognized image.

[0027] In one specific implementation scenario, the encoding network can include but is not limited to VGG, ResNet, etc., which is not limited herein.

[0028] In another specific implementation scenario, in order to improve the accuracy of the image features, the encoding network can also be set as a neural network combining CNN (Convolutional Neural Network, CNN) and Transformer. Specifically, for the to-be-recognized image (i.e., the to-be-recognized image I contains 3 channels, and the resolution of each channel is H*W), down-sampling can be performed first to obtain a feature map with a size of Specifically, two convolution blocks with a step of 2 can be used for fast down-sampling. On this basis, the feature map with a size of can be flattened to Specifically, for each feature map with a size of , the first and last of each row can be spliced to obtain a feature map with a size of the feature vector. On this basis, it can be further input to the encoder in the Transformer for encoding, and reshaped into a feature map with a size of the feature vector. On this basis, it can be further input to the encoder in the Transformer for encoding, and reshaped into a feature map with a size of

[0029] In another implementation scenario, unlike the foregoing implementation manner of extracting image features through a neural network, image features can also be extracted through a traditional manner such as HOG, SIFT, etc., which is not limited herein.

[0030] It should be noted that the foregoing examples are only several possible implementation manners of extracting image features in actual application, and do not limit the extraction of image features of the image to be recognized through other manners.

[0031] Step S12: performing the following first decoding operation based on the image features: extracting, from the image features, a first visual feature at a current decoding time based on decoding information at a previous decoding time; obtaining a language feature at the current decoding time based on the first visual feature at the current decoding time and the decoding information at the previous decoding time; and decoding based on the first visual feature and the language feature to obtain a decoded character at the current decoding time.

[0032] In the embodiments of the present disclosure, the decoding information includes at least one of the decoded character and the decoding state. For example, in order to improve the decoding accuracy as much as possible, the decoding information can include the decoded character and the decoding state. It should be noted that the decoding state at the previous decoding time includes feature information recorded when the decoding operation at the previous decoding time is performed. Therefore, by combining the decoded character and the decoding state to jointly perform decoding, the decoding accuracy can be improved. In addition, when the decoding operation is performed for the first time, the previous decoding time does not exist, and the decoding information at the previous decoding time can be a default value. For example, the decoded character can be a default value <start> 、 <s>and so on, the decoding state can be all-zero parameters by default; or, as mentioned above, text recognition can be performed on the to-be-recognized image by using the text recognition model to obtain the target recognition text of the to-be-recognized image, in which case the decoding information at the previous decoding time when the decoding operation is performed for the first time can be adjusted together with the network parameters of the text recognition model as a special network parameter until the training converges, and can be used in the text recognition stage in subsequent embodiments of the present disclosure. The training process of the text recognition model can be referred to in the related description below, and will not be described here.

[0033] In one implementation scenario, the first visual feature at the current decoding time can be obtained from the image feature based on the decoding information at the previous decoding time by using the attention mechanism. Specifically, the first query feature of the attention mechanism can be obtained based on the decoding information at the previous decoding time, and the first key feature and the first value feature of the attention mechanism can be obtained based on the image feature, so that the attention processing can be performed based on the first query feature, the first key feature and the first value feature to obtain the first visual feature at the current decoding time. It should be noted that the specific process of attention processing can be referred to the technical details of the attention mechanism (Attention), which will not be described here. In the above manner, the first query feature of the attention mechanism is obtained based on the decoding information at the previous decoding time, and the first key feature and the first value feature of the attention mechanism are obtained based on the image feature, so that the attention processing is performed based on the first query feature, the first key feature and the first value feature to obtain the first visual feature at the current decoding time, and then the decoding information at the previous decoding time can be used to mine the pure visual feature (i.e., the first visual feature) that is helpful for decoding at the current decoding time from the image feature by using the attention mechanism, which helps to improve the accuracy of pure visual feature mining.

[0034] In one specific implementation scenario, the embedding representation (i.e., embedding) corresponding to the decoded character at the previous decoding time can be obtained, and the decoding state at the previous decoding time can be obtained, and the decoding state and the embedding representation of the decoded character are spliced as the first query feature.

[0035] In one specific implementation scenario, the image feature can be directly used as the first key feature and the first value feature.

[0036] In one specific implementation scenario, in order to facilitate the distinction, the previous decoding time can be denoted as t-1, the current decoding time can be denoted as t, the decoded character at the previous decoding time can be denoted as y t-1 , the decoding state at the previous decoding time can be denoted as h t-1 , and the image feature can be denoted as F, then the first visual feature a t can be represented as:

[0037] a t = Attention([y t-1 ; h t-1 ], F, F) …… (1)

[0038] In the above formula (1), Attention(*query, *key, *value) represents attention processing, where *query represents a first query feature, which is the decoding state h t-1 at the previous decoding moment t-1 corresponding to the decoding character y t-1 ; h t-1 ] in formula (1) is directly recorded as [y t-1 ; h t-1 ] for convenience of representation. In addition, *key represents a first key feature, which is the image feature F, and *value represents a first value feature, which is the image feature.

[0039] In a specific implementation scenario, in particular, the above attention processing can adopt a coverage mechanism (coverage attention). The specific process of the attention processing using the coverage mechanism can be referred to technical details of the coverage mechanism, which will not be described here.

[0040] In another implementation scenario, unlike the above-mentioned extraction of the first visual feature at the current decoding moment from the image feature using the attention mechanism, other ways can also be used to extract the first visual feature, which is not limited by the embodiments of the present disclosure. Illustratively, the correlation between the decoding information at the previous decoding moment and each element in the image feature can also be obtained. The element in the image feature that is closer to the decoding character at the previous decoding moment has a higher corresponding correlation, and vice versa. In addition, the element in the image feature that contains more feature information recorded when the decoding operation at the previous decoding moment is performed has a higher corresponding correlation, and vice versa. Based on this, a weight matrix with the same resolution as the image feature can be formed, so that the image feature can be weighted element by element using the weight matrix, and the first visual feature can be obtained.

[0041] In yet another implementation scenario, unlike the above-mentioned two implementation manners, a weight prediction model can also be trained in advance. The weight prediction model can include but is not limited to a convolutional neural network, a recurrent neural network, etc., which is not limited here. On this basis, the decoding information at the previous decoding moment and the image feature can be input into the weight prediction model, and the weight prediction model can output a weight matrix with the same resolution as the image feature, so that the image feature can be weighted element by element using the weight matrix, and the first visual feature can be obtained.

[0042] It should be noted that the above embodiments are only several possible cases of extracting the first visual feature, and do not limit other embodiments. For example, in the case where the accuracy requirement of the first visual feature is relatively loose, any one of the decoded character of the previous decoding moment and the decoding state of the previous decoding moment can be referred to, such as only the decoded character of the previous decoding moment or only the decoding state of the previous decoding moment, which is not limited herein.

[0043] In one implementation scenario, after obtaining the first visual feature, the language feature of the current decoding moment can be obtained based on the decoding information of the previous decoding moment. It should be noted that, unlike the first visual feature which only contains pure visual feature information, the language feature contains language feature information required at the current decoding moment. Specifically, the decoding state of the previous decoding moment can be updated based on the decoded character of the previous decoding moment and the first visual feature of the current decoding moment to obtain the decoding state of the current decoding moment. On this basis, the language feature of the current decoding moment can be fused based on the decoded character of the previous decoding moment and the decoding state of the current decoding moment. The above method first updates the decoding state of the previous decoding moment based on the decoded character of the previous decoding moment and the first visual feature of the current decoding moment to obtain the decoding state of the current decoding moment, which can timely and accurately update the decoding state of the current decoding moment. Based on this, the language feature of the current decoding moment is fused based on the decoded character of the previous decoding moment and the decoding state of the current decoding moment, so that the decoded character of the previous decoding moment can be enhanced based on the decoding state of the current decoding moment, and the accuracy of the language feature of the current decoding moment can be improved.

[0044] In one specific implementation scenario, as one possible implementation, the decoding state of the current decoding moment can be updated using a long short-term memory (Long Short-Term Memory, LSTM) network. Specifically, the embedding corresponding to the decoded character of the previous decoding moment can be concatenated with the first visual feature of the current decoding moment, and then the concatenated feature and the decoding state of the previous decoding moment are input into the long short-term memory network, and the long short-term memory network outputs the decoding state of the current decoding moment. That is, the decoding state of the current decoding moment h t can be represented as:

[0045] h t =LSTM([y t-1 ;a t ],h t-1 )…… (2)

[0046] In the above formula (2), LSTM represents the long short-term memory network, y t-1 a represents the decoded character at the previous decoding time point t a represents the decoded character at the previous decoding time point t-1 a represents the decoded character at the previous decoding time point t-1 a represents the decoded character at the previous decoding time point t a represents the decoded character at the previous decoding time point t-1 a represents the decoded character at the previous decoding time point t a represents the decoded character at the previous decoding time point

[0047] In a specific implementation scenario, the embedding of the decoded character at the previous decoding time point can be obtained, and then the embedding is concatenated with the decoding state at the current decoding time point to obtain the language feature at the current decoding time point. That is, the language feature r t at the current decoding time point can be represented as:

[0048] r t = Concat(h t ; y t-1 ) … … (3)

[0049] In the above formula (3), Concat represents a concatenation operation, h t represents the decoding state at the current decoding time point, and y t-1 represents the decoded character at the previous decoding time point. It should be noted that, in formula (3), the embedding of the decoded character at the previous decoding time point is directly represented as y t-1 , that is, Concat(h t ; y t-1 ) actually represents concatenating the embedding of the decoded character at the previous decoding time point with the decoding state at the current decoding time point.

[0050] In another implementation scenario, different from the foregoing implementation, in order to improve the efficiency of extracting the language feature, a language feature extraction model can be pre-trained, which can include but is not limited to a recurrent neural network, and the network structure of the language feature extraction model is not limited herein. Based on this, the decoded character at the previous decoding time point, the decoding state, and the first visual feature at the current decoding time point can be input into the language feature extraction model to obtain the language feature at the current decoding time point and the decoding state, so as to reference the decoding state at the current decoding time point for decoding at the next decoding time point.

[0051] It should be noted that the above examples are only a few possible implementations of extracting language features, and do not limit the specific way of extracting language features. For example, in the case where the accuracy requirement of the language feature is relatively loose, the weight of each element in the decoding state at the previous decoding time can be predicted directly using the first visual feature at the current decoding time, and then the weight of each element in the decoding state at the previous decoding time is obtained. The decoding state at the previous decoding time is weighted based on each element and its weight at the previous decoding time, as the decoding state at the current decoding time, thereby fusing the embedding representation corresponding to the decoding character at the previous decoding time and the decoding state at the current decoding time to obtain the language feature at the current decoding time.

[0052] In one implementation scenario, after obtaining the first visual feature and the language feature, the decoding character at the current decoding time can be obtained based on the first visual feature and the language feature. Specifically, the weight information at the current decoding time can be obtained by predicting based on the first visual feature and the language feature, and then the first visual feature and the language feature are weighted based on the weight information to obtain the fusion feature at the current decoding time, so that the decoding character at the current decoding time is obtained by decoding based on the fusion feature at the current decoding time. The above-mentioned method, since in the decoding process, the weight information at the current decoding time is obtained by predicting based on the first visual feature and the language feature, so that the feature weighting is adaptively performed through the weight information, and then the decoding is performed, thereby dynamically referring to the visual feature and the language feature in the decoding process. Decoding helps to improve the decoding accuracy.

[0053] In a specific implementation scenario, in order to improve the weight prediction efficiency and accuracy, a weight prediction model can be trained in advance, and the weight prediction model can include but is not limited to a fully connected layer, and the network structure of the weight prediction model is not limited herein. In addition, in order to limit the weight information predicted by the weight prediction model to the range of 0~1, the weight information predicted by the weight prediction model can be normalized after the weight prediction model predicts the weight information. For example, the normalization operation can be performed through a normalization function such as sigmoid. That is, the weight information g t at the current decoding time can be represented as:

[0054] g t = sigmoid(W m [r t ; a t ]) … … (4)

[0055] In the above formula (4), g t represents the weight information at the current decoding time, sigmoid represents the normalization function, W m represents the network parameters of the weight prediction model, and r t a represents the language feature of the current decoding time point, t a represents the first visual feature of the current decoding time point.

[0056] In one specific implementation scenario, after obtaining the weight information g t of the current decoding time point, the first visual feature and the language feature of the current decoding time point are weighted and subsequently processed (such as dimension reduction, etc.), to obtain the fusion feature o t of the current decoding time point.

[0057] o t = W o [ g t ⊙ [ r t ; a t ] ]... (5)

[0058] In the above formula (5), o t represents the fusion feature of the current decoding time point, g t represents the weight information of the current decoding time point, r t represents the language feature of the current decoding time point, and a t represents the first visual feature of the current decoding time point.

[0059] In one specific implementation scenario, after obtaining the fusion feature of the current decoding time point, decoding can be performed based on the fusion feature of the current decoding time point, to obtain the decoding probability of each preset character at the current decoding time point, so that at least one preset character can be selected as the decoding character of the current decoding time point based on the decoding probability. Illustratively, each preset character can be sorted in descending order of the decoding probability, and the first two characters can be selected as the decoding characters. Of course, the selected preset characters are not limited to this, and for example, only the preset character with the highest decoding probability can be selected as the decoding character, which is not limited herein. That is, the decoding character y t of the current decoding time point can be represented as:

[0060] y t = softmax (MLP (o t ))... (6)

[0061] In the above formula (6), y t represents the decoding character of the current decoding time point, softmax represents a normalization operation, MLP represents a multi-layer perception, and o t represents the fusion feature of the current decoding time point. Further, MLP (o t ) represents the decoding probability of each preset character at the current decoding time based on the fusion feature decoding at the current decoding time. In addition, when multiple preset characters are selected as the decoding characters at the current decoding time, at the next decoding time, a new round of decoding can be performed by referring to the decoding information of each current decoding time (at the next decoding time, the current decoding time is the new previous decoding time) respectively, until the decoding is completed, so that multiple decoding paths can be obtained, and the decoding character combination at each decoding time on each decoding path is the candidate recognition text of the first decoding operation. For details, please refer to the technical details of beam search, which will not be described here.

[0062] In another implementation scenario, different from the foregoing implementation manner of directly decoding based on the first visual feature and the language feature, in order to further improve the decoding accuracy, the second visual feature of the current decoding time can be extracted from the image feature based on the position feature of the current decoding time before that, and the position features of different decoding times are encoded based on the corresponding decoding times. On this basis, the decoding character of the current decoding time can be obtained based on the first visual feature, the second visual feature and the language feature. In the above manner, in addition to the first visual feature and the language feature, the second visual feature of the current decoding time is further extracted from the image feature based on the position feature of the current decoding time, so that the first visual feature, the second visual feature and the language feature are combined to decode the decoding character of the current decoding time, and the pure visual feature (i.e. the second visual feature) queried by the expected position can be further referred to in the decoding process, which can help to further improve the decoding accuracy.

[0063] In a specific implementation scenario, the position feature of the current decoding time can be obtained by embedding the current decoding time. Specifically, the current decoding time (such as the foregoing t) can be one-hot encoded to obtain a one-hot encoding representation of the current decoding time, and then the one-hot encoding representation can be input into an embedding layer to obtain the position feature of the current decoding time. It should be noted that the position features of different decoding times are different. Of course, during the text recognition of different to-be-recognized images, the position features of the same decoding time remain unchanged. For example, the position feature of the decoding time t during the text recognition of to-be-recognized image A is the same as the position feature of the decoding time t during the text recognition of to-be-recognized image B. In addition, for details of one-hot encoding, please refer to the technical details of one-hot encoding, and for the implementation principle of the embedding layer, please refer to the technical details of embedding representation, which will not be described here. For ease of description, the position feature of the current decoding time t can be denoted as p t That is, for the embedding layer, each decoding time can output the corresponding position feature P = [p1, p2, …, pt, …, pT] through the embedding layer. t′ ].

[0064] In one specific implementation scenario, after obtaining the position feature of the current decoding time, the second query feature of the attention mechanism can be obtained based on the position feature of the current decoding time, and the second key feature and the second value feature of the attention mechanism can be obtained based on the image feature, and then the second visual feature of the current decoding time can be obtained by performing attention processing based on the second query feature, the second key feature and the second value feature. Similar to the foregoing description of the first visual feature, the position feature of the current decoding time can be directly taken as the second query feature, and the image feature can be directly taken as the second key feature and the second value feature. Of course, in order to further improve the position awareness, the image feature can also be position enhanced to obtain the second key feature. For example, two layers of long short-term memory networks can be used to process each row of the image feature to capture the global context. For each row, the long short-term memory networks share network parameters to overcome overfitting and reduce the amount of parameters. After that, two layers of convolutional networks can be used to extract high-dimensional information, so as to obtain the position enhanced feature as the second key feature. In this case, the second visual feature q t of the current decoding time t can be represented as:

[0065] q t = Attention(p t , F', F) …… (7)

[0066] In the foregoing formula (7), q t represents the second visual feature of the current decoding time, p t represents the position feature of the current decoding time, i.e., as the second query feature of the attention mechanism, F' represents the position enhanced feature, i.e., as the second key feature of the attention mechanism, and F represents the image feature, i.e., as the second value feature of the attention mechanism. In addition, Attention represents the attention processing, and the specific processing process can be referred to the technical details of the attention mechanism, which will not be described here.

[0067] In another specific implementation scenario, unlike the foregoing implementation of obtaining the second visual feature through the attention mechanism, other manners can also be used to extract the second visual feature, which are not limited herein. For example, the correlation between the position feature at the current decoding moment and each element in the image feature or the position enhanced feature can also be obtained. The higher the correlation, the more useful the feature information contained in the position is for the decoding at the current decoding moment. Conversely, the lower the correlation, the less useful the feature information contained in the position is for the decoding at the current decoding moment. Based on this, a weight matrix with the same resolution as the image feature can be formed, so that the image feature can be weighted element by element using the weight matrix, and the second visual feature can be obtained.

[0068] In yet another specific implementation scenario, unlike the foregoing two implementation manners, a weight prediction model can also be pre-trained. The weight prediction model can include but is not limited to a convolutional neural network, a recurrent neural network, etc., which are not limited herein. On this basis, the image feature / position enhanced feature and the position feature at the current decoding moment can be input into the weight prediction model, and the weight prediction model can output a weight matrix with the same resolution as the image feature, so that the image feature can be weighted element by element using the weight matrix, and the second visual feature can be obtained.

[0069] It should be noted that the foregoing implementation is only one of several possible manners of extracting the second visual feature, and other implementation manners are not limited herein. For example, in the case of relatively loose requirements on the accuracy of the second visual feature, the operation of position enhancement on the image feature can be omitted, that is, only the image feature and the position feature can be used, which are not limited herein.

[0070] In one specific implementation scenario, after obtaining the second visual feature, the weight information at the current decoding moment can be predicted based on the first visual feature, the second visual feature, and the language feature, and the first visual feature, the second visual feature, and the language feature can be weighted based on the weight information to obtain the fusion feature at the current decoding moment, so that the decoding character at the current decoding moment can be obtained based on the fusion feature at the current decoding moment. In order to improve the weight prediction efficiency and accuracy, a weight prediction model can be pre-trained, and the weight prediction model can include but is not limited to a fully connected layer, and the network structure of the weight prediction model is not limited herein. In addition, in order to limit the weight information predicted by the weight prediction model to the range of 0-1, the weight information predicted by the weight prediction model can be normalized after the weight prediction model predicts the weight information. For example, the normalization operation can be performed through a normalization function such as sigmoid. That is, the weight information gt at the current decoding moment can be represented as:

[0071] g t = sigmoid(W m [ r t ; a t ; q t ]) … … (8)

[0072] The specific meaning of the above formula (8) can be referred to the aforementioned formula (4) and its related description, which will not be repeated here. After obtaining the weight information of the current decoding moment, the weight information g t The first visual feature, the second visual feature and the language feature of the current decoding moment are weighted and processed (such as dimension reduction, etc.), and the fusion feature o t of the current decoding moment is obtained.

[0073] o t = W o [ g t ⊙ [ r t ; a t ; q t ]] … … (9)

[0074] The specific meaning of the above formula (8) can be referred to the aforementioned formula (5) and its related description, which will not be repeated here. After obtaining the fusion feature of the current decoding moment, decoding can be performed based on the fusion feature of the current decoding moment, and the decoding probability of each preset character at the current decoding moment is obtained, so that at least one preset character can be selected as the decoding character of the current decoding moment based on the decoding probability. For details, please refer to formula (6) and its related description, which will not be repeated here.

[0075] In one implementation scenario, please refer to Figure 2 , Figure 2 is a process schematic diagram of an embodiment of the first decoding operation. As shown in Figure 2 , during the execution of the first decoding operation process, the image feature can be taken as the first key feature and the first value feature respectively, and based on the decoding information (i.e. the decoding character y t-1 and the decoding state h t-1 ) of the previous decoding moment, the first query feature is obtained, so that the attention processing (such as visual perception attention in Figure 2 ) is performed based on the first query feature, the first key feature and the first value feature, and the first visual feature a t of the current decoding moment is obtained. On this basis, the decoding character y t-1 of the previous decoding moment, the first visual feature a t of the current decoding moment and the decoding state h t-1 of the previous decoding moment are input into the long short-term memory network LSTM to utilize the decoding character y t-1 of the previous decoding moment.and the first visual feature a t the decoding state h of the previous decoding time t-1 is updated to obtain the decoding state h of the current decoding time t , so that the decoding character y of the current decoding time can be obtained based on the decoding state h of the current decoding time t and the decoding character y of the previous decoding time t-1 , and the language feature r of the current decoding time is obtained t Meanwhile, the image feature can also be taken as the second value feature, and the position enhancement is performed based on the image feature as the second key feature, and the second query feature is obtained based on the position feature of the current decoding time, so that the attention processing (such as Figure 2 the position-aware attention) can be performed based on the second query feature, the second key feature and the second value feature to obtain the second visual feature q t of the current decoding time. On this basis, adaptive fusion can be performed based on the first visual feature a t , the second visual feature q t and the language feature r t to obtain the fusion feature o t of the current decoding time, and the fusion feature o t of the current decoding time is input into the multi-layer perception machine MLP for decoding to obtain the decoding character y t of the current decoding time.

[0076] In the embodiments of the present disclosure, the decoding characters of each decoding time are combined to obtain the candidate recognition text of the first decoding operation. Exemplarily, at the decoding time T, if the decoding character is an end character (such as, <end> 、 <eos>If the first decoding operation is completed, the decoding character of the current decoding time can be selected based on the decoding probability of each preset character as the decoding character of the current decoding time. For example, if the decoding probability of each preset character as the decoding character of the current decoding time is the same (or substantially the same), the decoding operation of the current decoding time can be ended. On this basis, the decoding characters of each decoding time can be combined in the order from the first to the last according to the decoding time to obtain the candidate recognition text of the first decoding operation.

[0077] Of course, in order to improve the decoding accuracy, in addition to the first decoding operation, several decoding operations can also include a second decoding operation. It should be noted that the decoding states of the first decoding operation and the second decoding operation at each decoding time are not shared with each other. In addition, the decoding probabilities of the first decoding operation and the second decoding operation at each decoding time can be combined or not combined. For details, please refer to the relevant description below. It should be noted that the second decoding operation can be implemented by a neural network such as Decoder in Transformer, recurrent neural network, long short-term memory network, etc., which is not limited here.

[0078] In one implementation scenario, the decoding probabilities of the first decoding operation and the second decoding operation at each decoding time can be combined. As mentioned earlier, in the first decoding operation, the decoding is performed based on the first visual feature and the language feature to obtain the decoding probability of each preset character as the decoding character of the current decoding time, which can be referred to as the first probability for ease of distinction. At the same time, in the second decoding operation, the decoding can be performed based on the image feature and the decoding information of the previous decoding time to obtain the decoding probability of each preset character as the decoding character of the current decoding time, which can be referred to as the second probability for ease of distinction. On this basis, at least one preset character can be selected as the decoding character of the current decoding time based on the first probability and the second probability of each preset character as the decoding character of the current decoding time. For example, for each preset character, the first probability and the second probability corresponding thereto can be subjected to any fusion operation such as addition, weighting, averaging, etc. to obtain the decoding probability of the preset character as the decoding character of the current decoding time, which can be referred to as the third probability for ease of distinction, so that at least one preset character can be selected as the decoding character of the current decoding time based on the third probability of each preset character as the decoding character of the current decoding time. For details of the process of determining the decoding character based on the decoding probability, please refer to the foregoing relevant description, which will not be repeated here. It should be noted that in this case, the decoding characters of each decoding time are combined to obtain the candidate recognition text commonly output by the first decoding operation and the second decoding operation. In addition, in the case of selecting multiple preset characters as decoding characters, please refer to the foregoing relevant description and the technical details of beam search, which will not be repeated here.

[0079] In another implementation scenario, the decoding probabilities of the first decoding operation and the second decoding operation at respective decoding moments can also not be combined. As described above, in the first decoding operation, decoding is performed based on the first visual feature and the language feature to obtain a decoding probability of each preset character as a decoding character at a current decoding moment, which can be referred to as a first probability for the sake of distinction, so that at least one preset character can be selected as the decoding character at the current decoding moment based on the first probability of each preset character as the decoding character at the current decoding moment, and then the decoding characters at respective decoding moments can be combined to obtain a candidate recognized text of the first decoding operation. Meanwhile, in the second decoding operation, decoding is performed based on the image feature and the decoding information of the previous decoding moment to obtain a decoding probability of each preset character as a decoding character at a current decoding moment, which can be referred to as a second probability for the sake of distinction, so that at least one preset character can be selected as the decoding character at the current decoding moment based on the second probability of each preset character as the decoding character at the current decoding moment, and then the decoding characters at respective decoding moments can be combined to obtain a candidate recognized text of the second decoding operation. In addition, in the case of selecting multiple preset characters as decoding characters, reference can be made to the related descriptions and technical details of beam search described above, which will not be repeated here.

[0080] Step S13: obtaining the target recognized text of the to-be-recognized image based on the candidate recognized texts of the decoding operations.

[0081] In one implementation scenario, for each candidate recognized text, a fusion operation such as summation, averaging, or weighting can be performed on the decoding probabilities of the decoding characters in the candidate recognized text to obtain a fusion probability of the candidate recognized text, so that the candidate recognized texts can be sorted in descending order of the fusion probability, and then the candidate recognized texts located in a preset sequence position (for example, the first position, the first two positions, etc.) can be selected as the target recognized text of the to-be-recognized image.

[0082] In an implementation scenario, to further improve the decoding accuracy for out-of-vocabulary (OOV) words, the first decoding operation can also be set to include a first forward decoding operation and a first reverse decoding operation. The first forward decoding operation is the first decoding operation that decodes character by character in the forward order, and the first reverse decoding operation is the first decoding operation that decodes character by character in the reverse order. Exemplarily, taking the text "The weather is great" contained in the image to be recognized as an example, the decoding order of the first forward decoding operation is: "天 (tiān)", "气 (qì)", "真 (zhēn)", "好 (hǎo)", and the decoding order of the first reverse decoding operation is "好 (hǎo)", "真 (zhēn)", "气 (qì)", "天 (tiān)". Other cases can be deduced by analogy and will not be elaborated here one by one. That is to say, except for the different decoding orders, the technical essence of the first forward decoding operation and the second reverse decoding operation is the same, and the decoding process can refer to the relevant description of the前述 first decoding operation and will not be repeated here. Based on this, the candidate recognition text of the first forward decoding operation can be reversed to obtain the decoding characters respectively referred to at each decoding moment when the first reverse decoding operation is re-executed; and / or, the candidate recognition text of the first reverse decoding operation can be reversed to obtain the decoding characters respectively referred to at each decoding moment when the first forward decoding operation is re-executed, and the first forward decoding operation is re-executed. Based on this, at least one candidate recognition text can be selected as the target recognition text based on the decoding probability distributions of the first decoding operation executed at least twice on the same candidate recognition text, and the first decoding operations executed twice have opposite decoding directions. That is to say, for the candidate recognition text of the first forward decoding operation, the first decoding operation executed at least twice on it includes the first forward decoding operation executed for the first time and the reverse decoding operation executed for the second time. On the contrary, for the candidate recognition text of the first reverse decoding operation, the first decoding operation executed at least twice on it includes the first reverse decoding operation executed for the first time and the first forward decoding operation executed for the second time. Through the above method, by reversing the candidate recognition text and re-executing the decoding operation with the opposite decoding order when the first decoding is performed for the first time by referring to the reversed candidate recognition text, that is, performing mutual re-decoding on the forward decoding and the reverse decoding, the reversed candidate recognition text can be used as the target output of the re-executed decoding operation, so as to select the target recognition text by combining the decoding probability distributions of the two first decoding operations, and further improve the decoding accuracy for OOV words. In addition, it can also improve the decoding accuracy for OOV + IV (In Vocabulary, that is, within the vocabulary).

[0083] In a specific implementation scenario, for the candidate recognition text of the first forward decoding operation, for the sake of description, it is denoted as Therefore, reversing it can obtain the decoding characters at each decoding moment when the first reverse decoding operation is re-executed Thus, in terms of the decoding time t = 1 at which the first reverse decoding operation is re-executed, the relevant steps described above in relation to the first decoding operation can be performed based on the image features and the decoding information of the previous decoding time (as described above, this can be a default parameter since there is no previous decoding time), to obtain the decoding probabilities of each preset character at time t = 1. Since at time t = 1, the target output of the re-executed first reverse decoding operation is the decoded character , the decoding probability of the decoded character can be taken. Similarly, in terms of the decoding time t = 2 at which the first reverse decoding operation is re-executed, the relevant steps described above in relation to the first decoding operation can be performed based on the image features and the decoding information of the previous decoding time (at this time, the decoded character of the previous decoding time is ). Thus, the decoding probabilities of each preset character at time t = 2 can be obtained. Since at time t = 2, the target output of the re-executed first reverse decoding operation is the decoded character , the decoding probability of the decoded character can be taken. In this way, until the decoded character is obtained, the decoding probability distribution of the re-executed first reverse decoding operation can be obtained. On this basis, the decoding probabilities of the same decoded characters in the decoding probability distributions of the two executions of the first decoding operation can be fused by summation, averaging, weighting, etc., to obtain the final probability distribution of the candidate recognized text

[0084]

[0085] In the above formula (10), y ored represents the candidate recognized text of the first forward decoding operation, Reverse(y pred ) represents the reverse sequence of the candidate recognized text of the first forward decoding operation, U L2R represents the first forward decoding operation, Y L2R (y pred ) represents the decoding probability distribution of the candidate recognized text y pred in the first execution of the first forward decoding operation, Y R2L represents the first reverse decoding operation, Y R2L (Reverse(y pred )) represents the decoding probability distribution of the candidate recognized text y pred after the reverse sequence in the second execution of the first reverse decoding operation.

[0086] In one specific implementation scenario, for the candidate recognition text of the first backward decoding operation, after being reversed, it can be taken as the target output of re-executing the first forward decoding operation, obtain the decoding probability distribution of the candidate recognition text re-executing the first forward decoding operation, so as to obtain the final probability distribution of the candidate recognition text in combination with its decoding probability distribution in the first backward decoding operation. For details, please refer to the foregoing description of re-executing the first backward decoding operation on the candidate recognition text of the first forward decoding operation after being reversed, which will not be repeated here.

[0087] In one specific implementation scenario, after re-executing the decoding operation, the decoding probability distributions of the same candidate recognition text in at least two executions of the first decoding operation can be fused to obtain the final probability distribution of the corresponding candidate recognition text, so that at least one candidate recognition text can be selected as the target recognition text based on the final probability distribution of each candidate recognition text. For example, after obtaining the final probability distribution of the candidate recognition text, the decoding probabilities of each decoding character in the candidate recognition text in the final probability distribution can be summed, averaged, weighted, or fused in other ways to obtain the fusion probability of the candidate recognition text, and finally at least one candidate recognition text can be selected as the target recognition text based on the fusion probability of the candidate recognition text. For details, please refer to the foregoing description, which will not be repeated here. The above-mentioned manner fuses the decoding probability distributions of the same candidate recognition text in at least two executions of the first decoding operation to obtain the final probability distribution of the corresponding candidate recognition text, so that at least one candidate recognition text can be selected as the target recognition text based on the final probability distribution of each candidate recognition text, and then the accuracy of the candidate recognition text being recognized can be determined in combination with the decoding probability distributions of the two first decoding operations, which helps to improve the accuracy of text recognition.

[0088] In one implementation scenario, as described previously, in order to improve the text recognition efficiency, the text recognition model can be used to recognize the to-be-recognized image, so as to obtain the target recognized text. In this case, in order to improve the model accuracy of the text recognition model, the text recognition model can be pre-trained. Specifically, the sample image features of the sample image can be extracted, and a first forward decoding operation is performed based on the sample image features to obtain a first predicted forward text decoded in a forward order, and a first backward decoding operation is performed based on the sample image features to obtain a first predicted backward text decoded in a backward order, so that the network parameters of the text recognition model can be adjusted based on at least the difference between the first predicted forward text after reverse order and the first predicted backward text, and / or the difference between the first predicted backward text after reverse order and the first predicted forward text. The above-mentioned manner adjusts the network parameters of the text recognition model by measuring the difference between the predicted texts obtained by the forward and backward decoding respectively, so as to constrain the decoding results of the forward and backward branches to be as consistent as possible, thereby forcing the forward and backward branches of the text recognition model to learn from each other, and further helping to improve the recognition accuracy of the text recognition model.

[0089] In one specific implementation scenario, the difference between the predicted texts of the forward and backward branches can be measured by, for example, KL divergence, without limitation. Taking the simultaneous measurement of the difference between the first predicted forward text after reverse order and the first predicted backward text, and the difference between the first predicted backward text after reverse order and the first predicted forward text as an example, in order to facilitate description, the first predicted forward text decoded in a forward order by the first forward decoding operation can be denoted as Y L2R , and the first predicted backward text decoded in a backward order by the first backward decoding operation can be denoted as Y R2L , then the mutual learning loss measured based on the above two differences can be denoted as: , which can be expressed as:

[0090]

[0091] In the above formula (11), KL represents the KL divergence function, and the specific calculation process can be referred to the calculation details of the KL divergence, which will not be described herein. In addition, RS represents the text reverse operation, and the specific meaning of the text reverse operation can be referred to the foregoing description, which will not be described herein.

[0092] In a specific implementation scenario, the sample image has sample text, in order to further improve the recognition accuracy of the text recognition model, the sample image can also be annotated with sample forward text obtained by arranging the sample characters in the sample text in a forward order one by one and sample reverse text obtained by arranging the sample characters in the sample text in a reverse order one by one, on this basis, a first loss (also referred to as a main loss) can be obtained based on the difference between the first predicted forward text and the sample forward text and / or the difference between the first predicted reverse text and the sample reverse text, and a second loss (i.e., the aforementioned mutual learning loss) can be obtained based on the difference between the first predicted forward text after being reversed and the first predicted reverse text and / or the difference between the first predicted reverse text after being reversed and the first predicted forward text, so that the network parameters of the text recognition model can be adjusted based on the first loss and the second loss. It should be noted that the first loss can be measured by a cross-entropy function, which is not limited herein. For ease of description, the sample characters contained in the sample text in the sample image can be denoted as [s1, s2, …, s L ], so the sample forward text arranged in a forward order can be denoted as S L2R = [s1, s2, …, s L ], the sample reverse text arranged in a reverse order can be denoted as S R2L = [s L , …, s2, s1], and the main loss (i.e., the aforementioned first loss) measured based on the above two differences can be represented as:

[0093]

[0094] In the above formula (12), CE represents a cross-entropy loss function, and its specific calculation process can be referred to the technical details of the cross-entropy loss function, which will not be described herein. The above method further measures the difference between the annotated text of the sample image and the predicted text of the text recognition model based on the mutual learning loss, so as to adjust the network parameters together, and then not only forces the forward branch and the reverse branch of the text recognition model to learn from each other, but also forces the forward branch and the reverse branch to learn the annotated text respectively. Therefore, it is helpful to further improve the accuracy of the text recognition model.

[0095] ​In one implementation scenario, as mentioned above, in order to further improve the decoding accuracy, the decoding operations can further include a second decoding operation, and the first decoding operation and the second decoding operation are independent of each other and have the same decoding direction. For example, when the first decoding operation includes a first forward decoding operation and a second backward decoding operation, the second decoding operation can also include a second forward decoding operation which is independent of the first forward decoding operation and has the same decoding direction, and similarly, when the first decoding operation includes a first backward decoding operation, the second decoding operation can also include a second backward decoding operation which is independent of the first backward decoding operation and has the same decoding direction. Further, as mentioned above, during the first decoding operation, decoding can be performed based on the first visual feature and the language feature to obtain a first probability of each preset character being a decoding character at a current decoding time, and during the second decoding operation, decoding can be performed based on the image feature and the decoding information at the previous decoding time to obtain a second probability of each preset character being a decoding character at the current decoding time. On this basis, at least one preset character can be selected as a decoding character at the current decoding time based on the first probability and the second probability of each preset character being a decoding character at the current decoding time, and the decoding characters at each decoding time can be combined to obtain a candidate recognition text output by the first decoding operation and the second decoding operation with the same decoding direction. Specifically, for the first forward decoding operation, forward decoding can be performed based on the first visual feature and the language feature to obtain a first probability of each preset character being a decoding character at a current decoding time, and similarly, for the second forward decoding operation, forward decoding can be performed based on the image feature and the decoding information at the previous decoding time to obtain a second probability of each preset character being a decoding character at the current decoding time, so that at each decoding time of the forward decoding, a fusion operation (such as addition, averaging, weighting, etc.) can be performed based on the first probability and the second probability of the same preset character to obtain a fusion probability of the corresponding preset character being a decoding character at the current decoding time, and then at least one preset character can be selected as a decoding character at the current decoding time based on the fusion probability, and the specific selection process can be referred to the foregoing related description, which will not be described here. Therefore, the decoding characters at each decoding time of the forward decoding (for example, if multiple preset characters are selected as decoding characters at each decoding time, then the decoding characters on the same decoding path at each decoding time of the forward decoding can be combined, i.e., common beam decoding, which can be referred to the foregoing related description and the technical details of beam search, which will not be described here) can be combined to obtain a candidate recognition text output by the first forward decoding operation and the second forward decoding operation.Similarly, for the first reverse decoding operation, the first reverse decoding operation can be performed based on the first visual feature and the language feature to obtain a first probability of each preset character being a decoding character at a current decoding time, and similarly, for the second reverse decoding operation, the second reverse decoding operation can be performed based on the image feature and the decoding information at the previous decoding time to obtain a second probability of each preset character being a decoding character at the current decoding time. Thus, at each decoding time of the reverse decoding, a fusion operation (such as addition, averaging, weighting, etc.) can be performed based on the first probability and the second probability of the same preset character to obtain a fusion probability of the corresponding preset character being a decoding character at the current decoding time, and then at least one preset character can be selected as a decoding character at the current decoding time based on the fusion probability. The specific selection process can be referred to the foregoing related description, which will not be repeated here. Therefore, the decoding characters at each decoding time of the reverse decoding (for example, if multiple preset characters are selected as decoding characters at each decoding time, the decoding characters located on the same decoding path at each decoding time of the reverse decoding can be combined, that is, the common beam decoding, which can be referred to the foregoing related description and the technical details of beam search, which will not be repeated here) can be combined to obtain the candidate recognition text output by the first reverse decoding operation and the second reverse decoding operation. The above method combines the first decoding operation and the second decoding operation in the same decoding direction to obtain the candidate recognition text in the decoding direction, which can make the first decoding operation and the second decoding operation complementary to each other, and is helpful to further improve the text recognition accuracy.

[0096] In a specific implementation scenario, as described above, the first decoding operation includes the first forward decoding operation and the first reverse decoding operation, and the second decoding operation includes the second forward decoding operation and the second reverse decoding operation. In this case, the first forward decoding operation and the second forward decoding operation can jointly output the candidate recognition text, and similarly, the first reverse decoding operation and the second reverse decoding operation can also jointly output the candidate recognition text. On this basis, the first forward decoding operation and the second forward decoding operation, which are both forward sequential decoding, and the first reverse decoding operation and the second reverse decoding operation, which are both reverse sequential decoding, can be mutually re-decoded to determine the target recognition text of the image to be recognized. Please refer to Figure 3 , Figure 3 is a process schematic diagram of an embodiment of the text recognition method of the present application. As Figure 3 As shown, the to-be-recognized image is encoded to obtain image features, the image features are jointly decoded by the first forward decoding operation and the second forward decoding operation to jointly output candidate recognized texts in a forward decoding order, and meanwhile, the image features are jointly decoded by the first backward decoding operation and the second backward decoding operation to jointly output candidate recognized texts in a backward decoding order. On this basis, the candidate recognized texts in the forward decoding order are mutually re-decoded, and / or the candidate recognized texts in the backward decoding order are mutually re-decoded, so as to obtain the target recognized text. Specifically, when the candidate recognized texts in the forward decoding order are mutually re-decoded, the candidate recognized texts jointly output by the first decoding operation and the second decoding operation in the forward decoding order can be reversed to obtain decoding characters respectively referenced by the first decoding operation and the second decoding operation in the backward decoding order when re-executed, and the first decoding operation and the second decoding operation in the forward decoding order are re-executed. When the candidate recognized texts in the backward decoding order are mutually re-decoded, the candidate recognized texts jointly output by the first decoding operation and the second decoding operation in the backward decoding order can be reversed to obtain decoding characters respectively referenced by the first decoding operation and the second decoding operation in the forward decoding order when re-executed, and the first decoding operation and the second decoding operation in the backward decoding order are re-executed. On this basis, at least one candidate recognized text is selected as the target recognized text based on decoding probability distributions of the same candidate recognized text in at least two times of decoding operation execution, and the two times of decoding operation execution have opposite decoding directions. The specific process of mutual re-decoding can refer to the technical details of mutual re-decoding described above, which will not be described here again.

[0097] In one specific implementation scenario, as mentioned above, the target recognition text can be recognized by a text recognition model from the to-be-recognized image, and in the case that the plurality of decoding operations further include a second decoding operation, and the second decoding operation further includes a second forward decoding operation and a second backward decoding operation, in order to improve the recognition accuracy of the text recognition model, the text recognition model can also be pre-trained. Specifically, the sample image features of the sample image can be extracted, and the first decoding operation and the second decoding operation, both of which are decoding in the forward order, can be performed based on the sample image features respectively to obtain a second predicted forward text, and the first decoding operation and the second decoding operation, both of which are decoding in the reverse order, can be performed based on the sample image features respectively to obtain a second predicted backward text, so that the network parameters of the text recognition model can be adjusted based on at least the difference between the second predicted forward text after being reversed and the second predicted backward text and / or the difference between the second predicted backward text after being reversed and the second predicted forward text, so as to force the forward decoding branch and the backward decoding branch to learn from each other and improve the recognition accuracy of the text recognition model. For example, based on both the difference between the second predicted forward text after being reversed and the second predicted backward text and the difference between the second predicted backward text after being reversed and the second predicted forward text to jointly adjust the network parameters of the text recognition model, for the convenience of description, the second predicted forward text of the first forward decoding operation can be denoted as Y L2R the second predicted forward text of the second forward decoding operation can be denoted as Y' L2R the second predicted backward text of the first backward decoding operation can be denoted as Y R2L the second predicted backward text of the second backward decoding operation can be denoted as Y' R2L In addition, as mentioned above, the above-mentioned difference can be measured by the KL divergence function, so the mutual learning loss based on the above-mentioned two differences can be denoted as:

[0098]

[0099] In the above formula (13), KL(Y L2R ||RS(Y R2L )) represents the difference between the second predicted backward text of the first backward decoding operation after being reversed and the second predicted forward text of the first forward decoding operation, KL(Y R2L , RS(Y L2R )) represents the difference between the second predicted forward text of the first forward decoding operation after being reversed and the second predicted backward text of the first backward decoding operation, KL(Y' L2R ||RS(Y' R2L )) represents the difference between the second predicted backward text of the second backward decoding operation after being reversed and the second predicted forward text of the second forward decoding operation, and KL(Y' R2L ​, RS(Y′ L2R )) represents the difference between the second predicted forward text after the reverse order of the second forward decoding operation and the second predicted reverse text of the second reverse decoding operation.

[0100] In one specific implementation scenario, as mentioned previously, in order to further improve the recognition accuracy of the text recognition model, a main loss can be further added on the basis of the aforementioned mutual learning loss. On this basis, the mutual learning loss and the main loss can be combined to jointly adjust the network parameters of the text recognition model. Specifically, the sample image has sample text, and the sample image is labeled with the sample forward text in which each sample character in the sample text is arranged in a forward order and the sample reverse text in which each sample character in the sample text is arranged in a reverse order. Then, the main loss can be obtained based on at least one of the difference between the second predicted forward text of the first forward decoding operation and the sample forward text, the difference between the second predicted forward text of the second forward decoding operation and the sample forward text, the difference between the second predicted reverse text of the first reverse decoding operation and the sample reverse text, and the difference between the second predicted reverse text of the second reverse decoding operation and the sample reverse text. For ease of description, the sample forward text can be denoted as S L2R , and the sample reverse text can be denoted as S R2L The main loss obtained based on the four differences can be denoted as:

[0101]

[0102] In the above formula (14), Y L2R represents the second predicted forward text of the first forward decoding operation, Y R2L represents the second predicted reverse text of the first reverse decoding operation, Y′ L2R represents the second predicted forward text of the second forward decoding operation, Y′ R2L represents the second predicted reverse text of the second reverse decoding operation. In addition, CE represents a cross-entropy loss function, and the specific process can be referred to the technical details of the cross-entropy loss function, which will not be described herein.

[0103] In one specific implementation scenario, please refer to Figure 4 , Figure 4 is a process schematic diagram of one embodiment of training the text recognition model. As Figure 4 ​As shown, the sample image is encoded to obtain sample image features, and then is subjected to first and second forward decoding operations to obtain second predicted forward texts, and is subjected to first and second reverse decoding operations to obtain second predicted reverse texts. On this basis, a main loss can be obtained based on differences between the sample forward text and the second predicted forward text of the first forward decoding operation, between the sample forward text and the second predicted forward text of the second forward decoding operation, between the sample reverse text and the second predicted reverse text of the first reverse decoding operation, and between the sample reverse text and the second predicted reverse text of the second reverse decoding operation Meanwhile, a mutual learning strategy can be adopted to obtain a mutual learning loss based on differences between the second predicted forward text of the first forward decoding operation and the second predicted reverse text of the first reverse decoding operation after being reversed, between the second predicted reverse text of the first reverse decoding operation and the second predicted forward text of the first forward decoding operation after being reversed, between the second predicted forward text of the second forward decoding operation and the second predicted reverse text of the second reverse decoding operation after being reversed, and between the second predicted reverse text of the second reverse decoding operation and the second predicted forward text of the second forward decoding operation after being reversed On this basis, a total loss can be obtained based on the main loss and the mutual learning loss

[0104]

[0105] In the above formula (15), λ represents a weight, that is, the main loss (i.e., the first loss mentioned above) and the mutual learning loss (i.e., the second loss mentioned above) can be weighted to obtain the total loss, so that the network parameters of the text recognition model can be adjusted based on the total loss. It should be noted that the weight λ can be set to 0.4 after applying the grid search method. In addition, during the training process, the sample image can be uniformly adjusted to 32*100, the network parameter adjustment can be realized by the Adam optimizer, the basic learning rate can be set to 1e-4, the weight decay can be set to 1e-5, and the batch size can be set to 128. Of course, the above parameter settings are only one possible implementation in the training process, and do not limit the specific training parameters.

[0106] In one implementation scenario, based on the fine-grained verification and test set with OOV and IV labels, the effect of the text recognition method of the present application and the prior art is shown in Table 1:

[0107] Table 1 Comparison of effects of the text recognition method of the present application and the prior art

[0108]

[0109] As shown in Table 1, the text recognition method of the present application achieves the best performance compared with the prior art, whether in OOV or in OOV+IV. In addition, in order to further verify the effectiveness of each decoding strategy (such as the first decoding operation, the second decoding operation, the mutual decoding and the combination thereof), the ablation verification is further performed on the validation and test sets, and the verification results are shown in Table 2:

[0110] Table 2 Ablation verification results

[0111]

[0112] As shown in Table 2, the baseline is a simple decoder based on formula (3), and the performance thereof can be comparable to that of the comparative scheme RobustScanner. In addition, the first decoding operation, the second decoding operation and the mutual decoding are also proved to be effective.

[0113] Please refer to Figure 5 , Figure 5 is an effect diagram of an embodiment of the text recognition method of the present application. As shown in Figure 5 , Figure 5 contains 5 test images (respectively containing the following texts: "tickets:www.sunsettickets", "http: / / www.ameibo.com / bundle...", "HealthCity", "BURDIGALA", "¥580"), and the recognition results of the text recognition method of the present application (ours) and other three prior arts (SAR, RobustScanner, SATRN) are respectively shown under each test image, wherein if there is an underscore under a character in the recognition result, it means that the character recognition is wrong. As shown in Figure 5 , on the above 5 test images, the text recognition method of the present application has obvious advantages compared with the other three prior arts (SAR, RobustScanner, SATRN). Figure 5

[0114] ​The scheme extracts image features of the image to be recognized, and the image features are used for several decoding operations, and the several decoding operations at least include a first decoding operation. The first decoding operation is performed based on the image features as follows: first visual features of a current decoding time are extracted from the image features based on decoding information of a previous decoding time, language features of the current decoding time are obtained based on the first visual features of the current decoding time and the decoding information of the previous decoding time, and decoding is performed based on the first visual features and the language features to obtain decoded characters of the current decoding time. The decoding information includes at least one of the decoded characters and a decoding state. The decoded characters of each decoding time are combined to obtain candidate recognition text of the first decoding operation. Finally, the target recognition text of the image to be recognized is obtained based on the candidate recognition text of each decoding operation. Since the visual features and the language features are modeled based on the image features at each decoding time and combined for decoding, the features in the visual dimension and the language dimension can be combined for decoding, thereby helping to improve the accuracy of text recognition, especially the accuracy of OOV.

[0115] Please refer to Figure 6 , Figure 6 is a schematic diagram of an embodiment of the text recognition device 60. The text recognition device 60 includes an encoding module 61, a decoding module 62, and a determination module 63. The encoding module 61 is configured to extract image features of an image to be recognized. The image features are used for several decoding operations, and the several decoding operations at least include a first decoding operation. The decoding module 62 is configured to perform the first decoding operation based on the image features as follows: first visual features of a current decoding time are extracted from the image features based on decoding information of a previous decoding time; language features of the current decoding time are obtained based on the first visual features of the current decoding time and the decoding information of the previous decoding time; and decoding is performed based on the first visual features and the language features to obtain decoded characters of the current decoding time. The decoding information includes at least one of the decoded characters and a decoding state. The decoded characters of each decoding time are combined to obtain candidate recognition text of the first decoding operation. The determination module 63 is configured to obtain the target recognition text of the image to be recognized based on the candidate recognition text of each decoding operation.

[0116] In the scheme, the text recognition apparatus 60 extracts image features of the image to be recognized, and the image features are used for a plurality of decoding operations, the plurality of decoding operations at least include a first decoding operation, and the first decoding operation is performed based on the image features as follows: first visual features of a current decoding time are extracted from the image features based on decoding information of a previous decoding time, language features of the current decoding time are obtained based on the first visual features of the current decoding time and the decoding information of the previous decoding time, and decoding is performed based on the first visual features and the language features to obtain decoded characters of the current decoding time, and the decoding information includes at least one of the decoded characters and a decoding state, candidate recognition texts of the first decoding operation are obtained by combining the decoded characters of each decoding time, and finally, target recognition texts of the image to be recognized are obtained based on the candidate recognition texts of the plurality of decoding operations. Since the visual features and the language features are modeled based on the image features at each decoding time and the two features are combined for decoding, the features in the visual dimension and the language dimension can be combined for decoding, thereby helping to improve the accuracy of text recognition, especially the accuracy of OOV.

[0117] In some disclosed embodiments, the decoding module 62 includes a decoding state updating submodule configured to update a decoding state of a previous decoding time based on decoded characters of the previous decoding time and first visual features of a current decoding time to obtain a decoding state of the current decoding time; and the decoding module 62 includes a decoding information fusion submodule configured to fuse to obtain language features of the current decoding time based on the decoded characters of the previous decoding time and the decoding state of the current decoding time.

[0118] In some disclosed embodiments, the decoding module 62 includes a first feature obtaining submodule configured to obtain first query features of an attention mechanism based on decoding information of a previous decoding time, and obtain first key features and first value features of the attention mechanism based on the image features; and the decoding module 62 includes a first attention processing submodule configured to perform attention processing based on the first query features, the first key features, and the first value features to obtain first visual features of a current decoding time.

[0119] In some disclosed embodiments, the decoding module 62 includes a visual feature obtaining submodule configured to extract second visual features of a current decoding time from the image features based on position features of the current decoding time; and the position features of different decoding times are obtained based on corresponding decoding times; and the decoding module 62 includes a decoded character obtaining submodule configured to perform decoding based on the first visual features, the second visual features, and the language features to obtain decoded characters of the current decoding time.

[0120] In some disclosed embodiments, the visual feature obtaining sub-module comprises a second feature obtaining unit configured to obtain a second query feature of the attention mechanism based on the position feature of the current decoding time, and obtain a second key feature and a second value feature of the attention mechanism based on the image feature; the visual feature obtaining sub-module comprises a second attention processing unit configured to perform attention processing based on the second query feature, the second key feature and the second value feature to obtain a second visual feature of the current decoding time.

[0121] In some disclosed embodiments, the decoding character obtaining sub-module comprises a weight information predicting unit configured to predict based on the first visual feature, the second visual feature and the language feature to obtain weight information of the current decoding time; the decoding character obtaining sub-module comprises a feature weighted fusion unit configured to weight the first visual feature, the second visual feature and the language feature based on the weight information to obtain a fusion feature of the current decoding time; the decoding character obtaining sub-module comprises a current decoding obtaining unit configured to decode based on the fusion feature of the current decoding time to obtain a decoding character of the current decoding time.

[0122] In some disclosed embodiments, the first decoding operation comprises a first forward decoding operation and a first backward decoding operation, the first forward decoding operation is a first decoding operation of character-by-character decoding in a forward order, the first backward decoding operation is a first decoding operation of character-by-character decoding in a backward order, the determining module 63 comprises a re-decoding sub-module configured to reverse the candidate recognized text of the first forward decoding operation to obtain decoding characters respectively referenced by each decoding time when the first backward decoding operation is re-executed, and re-execute the first backward decoding operation; and / or, reverse the candidate recognized text of the first backward decoding operation to obtain decoding characters respectively referenced by each decoding time when the first forward decoding operation is re-executed, and re-execute the first forward decoding operation; the determining module 63 comprises a text selection sub-module configured to select at least one candidate recognized text as a target recognized text based on decoding probability distributions of the same candidate recognized text in at least two times of executing the first decoding operation; wherein the two times of executing the first decoding operation have opposite decoding directions.

[0123] In some disclosed embodiments, the text selection sub-module comprises a probability distribution fusion unit configured to fuse decoding probability distributions of the same candidate recognized text in at least two times of executing the first decoding operation to obtain a final probability distribution corresponding to the candidate recognized text; the text selection sub-module comprises a target text determining unit configured to select at least one candidate recognized text as a target recognized text based on the final probability distribution of each candidate recognized text.

[0124] In some disclosed embodiments, the target recognition text is recognized by the text recognition model from the to-be-recognized image, the text recognition apparatus 60 further comprises a sample encoding module configured to extract sample image features of a sample image; the text recognition apparatus 60 further comprises a sample decoding module configured to perform a first forward decoding operation based on the sample image features to obtain a first predicted forward text decoded in a forward order character by character, and perform a first backward decoding operation based on the sample image features to obtain a first predicted backward text decoded in a backward order character by character; and the text recognition apparatus 60 further comprises a parameter adjustment module configured to adjust network parameters of the text recognition model based on at least a difference between the first predicted forward text in reverse order and the first predicted backward text and / or a difference between the first predicted backward text in reverse order and the first predicted forward text.

[0125] In some disclosed embodiments, the sample image has sample text, and the sample image is labeled with a sample forward text in which sample characters in the sample text are arranged in a forward order character by character and a sample backward text in which sample characters in the sample text are arranged in a backward order character by character, the text recognition apparatus 60 further comprises a first loss measurement module configured to obtain a first loss based on a difference between the first predicted forward text and the sample forward text and / or a difference between the first predicted backward text and the sample backward text; the text recognition apparatus 60 further comprises a second loss measurement module configured to obtain a second loss based on a difference between the first predicted forward text in reverse order and the first predicted backward text and / or a difference between the first predicted backward text in reverse order and the first predicted forward text; and the parameter adjustment module is specifically configured to adjust the network parameters based on the first loss and the second loss.

[0126] In some disclosed embodiments, the plurality of decoding operations further comprise a second decoding operation independent of the first decoding operation and having the same decoding direction, the decoding module 62 is specifically configured to perform decoding based on the first visual features and the language features to obtain a first probability of each preset character being a decoding character at a current decoding time; perform the following second decoding operation based on the image features: perform decoding based on the image features and decoding information at a previous decoding time to obtain a second probability of each preset character being a decoding character at the current decoding time; and select at least one preset character as the decoding character at the current decoding time based on the first probability and the second probability of each preset character being the decoding character at the current decoding time; and the decoding characters at each decoding time are combined to obtain a candidate recognition text commonly output by the first decoding operation and the second decoding operation having the same decoding direction.

[0127] In some disclosed embodiments, the second decoding operation includes a second forward decoding operation and a second backward decoding operation, the second forward decoding operation being the second decoding operation that decodes character by character in a forward order, and the second backward decoding operation being the second decoding operation that decodes character by character in a backward order, the determining module 63 further includes a re-decoding sub-module configured to: reverse the candidate recognized text that is output by both the first decoding operation and the second decoding operation that decodes in the forward order, to obtain the decoding characters that are respectively referenced by each decoding time when the first decoding operation and the second decoding operation that decodes in the backward order are re-executed, and re-execute the first decoding operation and the second decoding operation that decodes in the backward order; and / or reverse the candidate recognized text that is output by both the first decoding operation and the second decoding operation that decodes in the backward order, to obtain the decoding characters that are respectively referenced by each decoding time when the first decoding operation and the second decoding operation that decodes in the forward order are re-executed, and re-execute the first decoding operation and the second decoding operation that decodes in the forward order; the determining module 63 further includes a text selection sub-module configured to select at least one candidate recognized text as the target recognized text based on the decoding probability distribution of the same candidate recognized text in at least two times of executing the decoding operation.

[0128] In some disclosed embodiments, the target recognized text is obtained by the text recognition model from the to-be-recognized image, the text recognition apparatus 60 further includes a sample encoding module configured to extract a sample image feature of a sample image; the text recognition apparatus 60 further includes a sample decoding module configured to execute the first decoding operation and the second decoding operation that decodes in the forward order based on the sample image feature to obtain a second predicted forward text, and execute the first decoding operation and the second decoding operation that decodes in the backward order based on the sample image feature to obtain a second predicted backward text; the text recognition apparatus 60 further includes a parameter adjustment module configured to adjust the network parameter of the text recognition model based on at least the difference between the second predicted forward text after being reversed and the second predicted backward text and / or the difference between the second predicted backward text after being reversed and the second predicted forward text.

[0129] Please refer to Figure 7 , Figure 7 is a schematic diagram of a framework of an embodiment of the electronic device 70. The electronic device 70 includes a memory 71 and a processor 72 that are coupled to each other, the memory 71 stores program instructions, and the processor 72 is configured to execute the program instructions to implement the steps in any of the above text recognition method embodiments. Specifically, the electronic device 70 can include but is not limited to a smart phone, a tablet computer, a learning machine, an office book, a server, a computer, etc., which are not limited herein.

[0130] Specifically, the processor 72 is configured to control itself and the memory 71 to implement the steps in any of the above-described text recognition method embodiments. The processor 72 can also be referred to as a CPU (Central Processing Unit). The processor 72 can be an integrated circuit chip having a processing capability of signals. The processor 72 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 72 can be implemented by an integrated circuit chip together.

[0131] In the above solution, the electronic device 70 extracts the image features of the image to be recognized, and the image features are used for a plurality of decoding operations, the plurality of decoding operations at least include a first decoding operation, and the first decoding operation is performed based on the image features as follows: based on the decoding information at the last decoding moment, first visual features at the current decoding moment are extracted from the image features, and based on the first visual features at the current decoding moment and the decoding information at the last decoding moment, language features at the current decoding moment are obtained, and decoding is performed based on the first visual features and the language features to obtain decoded characters at the current decoding moment, and the decoding information includes at least one of the decoded characters and a decoding state, the decoded characters at each decoding moment are combined to obtain a candidate recognized text of the first decoding operation, and finally, based on the candidate recognized texts of the plurality of decoding operations, a target recognized text of the image to be recognized is obtained. Since the visual features and the language features are modeled based on the image features at each decoding moment and combined for decoding, the features in the visual dimension and the language dimension can be combined for decoding, thereby helping to improve the accuracy of text recognition, especially the accuracy of OOV.

[0132] Please refer to Figure 8 , Figure 8 is a schematic diagram of an embodiment of the computer readable storage medium 80 of the present application. The computer readable storage medium 80 stores program instructions 81 capable of being executed by the processor, and the program instructions 81 are used to implement the steps in any of the above-described text recognition method embodiments.

[0133] In the scheme, the computer readable storage medium 80 extracts the image features of the image to be recognized, and the image features are used for several decoding operations, and the several decoding operations at least include a first decoding operation. The first decoding operation is performed based on the image features as follows: first visual features of a current decoding time are extracted from the image features based on decoding information of a previous decoding time, language features of the current decoding time are obtained based on the first visual features of the current decoding time and the decoding information of the previous decoding time, and decoding is performed based on the first visual features and the language features to obtain decoded characters of the current decoding time. The decoding information includes at least one of the decoded characters and a decoding state. The decoded characters of each decoding time are combined to obtain candidate recognition text of the first decoding operation. Finally, the target recognition text of the image to be recognized is obtained based on the candidate recognition text of each decoding operation of the several decoding operations. Since the visual features and the language features are modeled based on the image features at each decoding time and the two features are combined for decoding, the features in the visual dimension and the language dimension can be combined for decoding, thereby helping to improve the accuracy of text recognition, especially the accuracy of OOV.

[0134] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and the specific implementation can be referred to the description of the above method embodiments. For brevity, details are not described here.

[0135] The above description of each embodiment tends to emphasize the differences between the embodiments, and the same or similar parts can be mutually referred to, and for brevity, details are not described here.

[0136] In several embodiments provided in the present application, it should be understood that the disclosed methods and apparatuses can be implemented in other ways. For example, the apparatus implementation described above is only schematic, and for example, the division of the modules or units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed mutual elements can be indirect coupling or communication connection through some interfaces, apparatuses or units, and can be electrical, mechanical or other forms.

[0137] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.

[0138] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0139] If the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the present application, essentially or in the form of a contribution to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods in the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, and various media that can store program codes.

[0140] If the technical solutions of the present application involve personal information, the product applying the technical solutions of the present application has been explicitly informed of the personal information processing rules before processing the personal information, and has obtained the personal independent consent. If the technical solutions of the present application involve sensitive personal information, the product applying the technical solutions of the present application has obtained the personal independent consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at the personal information collection device such as camera, a clear and prominent mark is set to inform that the personal information collection range has been entered, and the personal information will be collected. If the person voluntarily enters the collection range, it is regarded as agreeing to collect the personal information. Or, on the device for processing personal information, the personal information processing rules are informed by using obvious marks / information, and the personal authorization is obtained by means of pop-up information or asking the person to upload his / her personal information. The personal information processing rules can include personal information processor, personal information processing purpose, processing method, and personal information type, etc.< / eos> < / end> < / s> < / start>

Claims

1. A text recognition method, characterized by, The method comprises: extracting image features of a to-be-recognized image; wherein the image features are used for a plurality of decoding operations, and the plurality of decoding operations at least include a first decoding operation; performing the first decoding operation based on the image features as follows: based on decoding information of a previous decoding time, extracting first visual features of a current decoding time from the image features; and based on the first visual features of the current decoding time and the decoding information of the previous decoding time, obtaining language features of the current decoding time; and based on the first visual features and the language features, decoding to obtain decoded characters of the current decoding time; wherein the decoding information includes at least one of decoded characters and a decoding state, and combining the decoded characters of each decoding time obtains candidate recognition text of the first decoding operation; based on the candidate recognition text of each of the plurality of decoding operations, obtaining target recognition text of the to-be-recognized image; wherein the obtaining of the language features of the current decoding time based on the first visual features of the current decoding time and the decoding information of the previous decoding time comprises: updating the decoding state of the previous decoding time based on the decoded characters of the previous decoding time and the first visual features of the current decoding time, to obtain the decoding state of the current decoding time; based on the decoded characters of the previous decoding time and the decoding state of the current decoding time, fusing to obtain the language features of the current decoding time.

2. The method of claim 1, wherein, The extracting of the first visual features of the current decoding time from the image features based on the decoding information of the previous decoding time comprises: based on the decoding information of the previous decoding time, obtaining first query features of an attention mechanism, and based on the image features, obtaining first key features and first value features of the attention mechanism; based on the first query features, the first key features and the first value features, performing attention processing to obtain the first visual features of the current decoding time.

3. The method of claim 1, wherein, Before the decoding based on the first visual features and the language features to obtain the decoded characters of the current decoding time, the method further comprises: based on position features of the current decoding time, extracting second visual features of the current decoding time from the image features; wherein the position features of different decoding times are respectively obtained based on encoding corresponding to the decoding times; the decoding based on the first visual features and the language features to obtain the decoded characters of the current decoding time comprises: based on the first visual features, the second visual features and the language features, decoding to obtain the decoded characters of the current decoding time.

4. The method of claim 3, wherein, The decoding based on the first visual features, the second visual features and the language features to obtain the decoded characters of the current decoding time comprises: based on the first visual features, the second visual features and the language features, predicting to obtain weight information of the current decoding time; based on the weight information, weighting the first visual features, the second visual features and the language features to obtain fused features of the current decoding time; decoding based on the fusion feature of the current decoding moment to obtain a decoded character of the current decoding moment.

5. The method of claim 1, wherein, The first decoding operation includes a first forward decoding operation and a first backward decoding operation, and the obtaining of the target recognition text of the to-be-recognized image based on the candidate recognition texts of the decoding operations respectively includes: reversing the candidate recognition text of the first forward decoding operation to obtain the decoded character referenced by each decoding moment when the first backward decoding operation is re-executed, and re-executing the first backward decoding operation; and / or, reversing the candidate recognition text of the first backward decoding operation to obtain the decoded character referenced by each decoding moment when the first forward decoding operation is re-executed, and re-executing the first forward decoding operation; selecting at least one candidate recognition text as the target recognition text based on the decoding probability distribution of the same candidate recognition text in at least two executions of the first decoding operation; and wherein the two executions of the first decoding operation have opposite decoding directions.

6. The method of claim 5, wherein, The selecting at least one candidate recognition text as the target recognition text based on the decoding probability distribution of the same candidate recognition text in at least two executions of the first decoding operation includes: fusing the decoding probability distribution of the same candidate recognition text in two executions of the first decoding operation to obtain a final probability distribution corresponding to the candidate recognition text; and selecting at least one candidate recognition text as the target recognition text based on the final probability distribution of each candidate recognition text.

7. The method of claim 5, wherein, The target recognition text is obtained by a text recognition model recognizing the to-be-recognized image, and a training step of the text recognition model includes: extracting a sample image feature of a sample image; performing the first forward decoding operation based on the sample image feature to obtain a first predicted forward text decoded character by character in a forward order, and performing the first backward decoding operation based on the sample image feature to obtain a first predicted backward text decoded character by character in a backward order; adjusting a network parameter of the text recognition model based on at least a difference between the first predicted forward text after reverse order and the first predicted backward text and / or a difference between the first predicted backward text after reverse order and the first predicted forward text.

8. The method according to any one of claims 1 to 7, characterized in that, The decoding operations further include a second decoding operation independent of the first decoding operation and having the same decoding direction, and the decoding based on the first visual feature and the language feature to obtain the decoded character of the current decoding moment includes: decoding based on the first visual feature and the language feature to obtain a first probability that each preset character is the decoded character of the current decoding moment; The method further includes: performing the second decoding operation based on the image feature and the decoding information of the previous decoding moment to obtain a second probability that each preset character is the decoded character of the current decoding moment; and performing the second decoding operation based on the image feature and the decoding information of the previous decoding moment to obtain a second probability that each preset character is the decoded character of the current decoding moment. The at least one preset character is selected as the decoding character of the current decoding moment based on first probabilities and second probabilities of each preset character being the decoding character of the current decoding moment; and the decoding characters of each decoding moment are combined to obtain a candidate recognized text commonly output by the first decoding operation and the second decoding operation having the same decoding direction.

9. A text recognition apparatus characterized by comprising: The method comprises: The encoding module is configured to extract image features of the image to be recognized; wherein the image features are used for a plurality of decoding operations, and the plurality of decoding operations at least include a first decoding operation; The decoding module is configured to perform the first decoding operation as follows based on the image features: extract first visual features of a current decoding moment from the image features based on decoding information of a previous decoding moment; obtain language features of the current decoding moment based on the first visual features of the current decoding moment and the decoding information of the previous decoding moment; and perform decoding based on the first visual features and the language features to obtain a decoding character of the current decoding moment; wherein the decoding information includes at least one of a decoding character and a decoding state, and the decoding characters of each decoding moment are combined to obtain a candidate recognized text of the first decoding operation; The determining module is configured to obtain a target recognized text of the image to be recognized based on the candidate recognized texts of the plurality of decoding operations; wherein the obtaining of the language features of the current decoding moment based on the first visual features of the current decoding moment and the decoding information of the previous decoding moment includes: updating a decoding state of the previous decoding moment based on the decoding character of the previous decoding moment and the first visual features of the current decoding moment to obtain a decoding state of the current decoding moment; and fusing the decoding character of the previous decoding moment and the decoding state of the current decoding moment to obtain the language features of the current decoding moment.

10. An electronic device, comprising: The memory and the processor are coupled to each other, the memory stores program instructions, and the processor is configured to execute the program instructions to implement the text recognition method of any one of claims 1 to 8.

11. A computer readable storage medium, characterized in that, The memory stores program instructions executable by the processor, and the program instructions are used to implement the text recognition method of any one of claims 1 to 8.

Citation Information

Patent Citations

  • Semantic enhanced scene text recognition method and device

    CN113591546A

  • Text recognition method and device, electronic equipment and readable storage medium

    CN114298054A