A text image recognition method, system and related device
By combining attention mechanisms with semantic information, this text image recognition method solves the problem of poor recognition performance of OCR technology under environmental interference, achieves automated and highly accurate error correction, and reduces manual intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2026-04-10
AI Technical Summary
Existing OCR technology is easily affected by the shooting environment when recognizing text in images, resulting in poor recognition results, and traditional error correction methods require a lot of manpower and resources.
A text image recognition method that combines attention mechanism with semantic information is adopted. By obtaining the recognition confidence score and semantic confidence score of the initial text, the text to be corrected is identified and corrected, thereby improving the recognition accuracy.
It improves the accuracy and stability of text image recognition and reduces the need for manual error correction.
Smart Images

Figure CN115909381B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image recognition, in particular to a text image recognition method, system and related device. BACKGROUND
[0002] With the wide popularity of automatic office scenarios, the industry has increasingly high requirements for the accuracy of electronic documents, especially in the financial, medical and other fields. Existing electronic documents are mainly stored in text and picture formats, and the two document formats are often converted to each other in different scenarios. Among them, the conversion from picture format to text format is usually realized by using optical character recognition (OCR) technology.
[0003] However, due to the interference of factors such as shooting environment, the effect of OCR recognition is prone to be poor. The traditional method for solving text misrecognition is manual correction, which requires a lot of manpower and material resources. Therefore, how to improve the recognition accuracy of picture text and correct the recognized text becomes the key of image text recognition technology. SUMMARY
[0004] The technical problem solved by the present application is to provide a text image recognition method, system and related device, which can improve the accuracy of text image recognition.
[0005] To solve the above technical problems, one technical solution adopted by the present application is to provide a text image recognition method, comprising: obtaining a text image including to-be-recognized text, obtaining initial characters corresponding to the to-be-recognized text and recognition confidence scores corresponding to the initial characters based on the text image; obtaining semantic confidence scores of each initial character based on semantic information of each initial character; determining at least part of to-be-corrected characters from all the initial characters based on the recognition confidence scores and the semantic confidence scores corresponding to each initial character; correcting the to-be-corrected characters to obtain target text corresponding to the text image.
[0006] To solve the above technical problems, another technical solution adopted by the present application is to provide a text image recognition system, comprising: an identification module, configured to acquire a text image comprising to-be-identified text, obtain initial text corresponding to the to-be-identified text based on the text image, and obtain an identification confidence score corresponding to the initial text; a semantic analysis module, configured to obtain a semantic confidence score of each initial text based on semantic information of each initial text; a processing module, configured to determine at least part of to-be-corrected text from all the initial texts based on the identification confidence score and the semantic confidence score corresponding to each initial text; and a correction module, configured to correct the to-be-corrected text to obtain target text corresponding to the text image.
[0007] To solve the above technical problems, another technical solution adopted by the present application is to provide a text image recognition system, comprising: an identification module, configured to acquire a text image comprising to-be-identified text, obtain initial text corresponding to the to-be-identified text based on the text image, and obtain an identification confidence score corresponding to the initial text; a semantic analysis module, configured to obtain a semantic confidence score of each initial text based on semantic information of each initial text; a processing module, configured to determine at least part of to-be-corrected text from all the initial texts based on the identification confidence score and the semantic confidence score corresponding to each initial text; and a correction module, configured to correct the to-be-corrected text to obtain target text corresponding to the text image.
[0008] To solve the above technical problems, another technical solution adopted by the present application is to provide a text image recognition system, comprising: an identification module, configured to acquire a text image comprising to-be-identified text, obtain initial text corresponding to the to-be-identified text based on the text image, and obtain an identification confidence score corresponding to the initial text; a semantic analysis module, configured to obtain a semantic confidence score of each initial text based on semantic information of each initial text; a processing module, configured to determine at least part of to-be-corrected text from all the initial texts based on the identification confidence score and the semantic confidence score corresponding to each initial text; and a correction module, configured to correct the to-be-corrected text to obtain target text corresponding to the text image.
[0009] The present application has the following beneficial effects: Different from the prior art, the text image recognition method proposed by the present application uses the attention mechanism to combine the semantic information of to-be-identified text in the to-be-identified text to perform identification, so as to obtain the initial text corresponding to the to-be-identified text and the corresponding identification confidence score. After obtaining the initial text, the semantic confidence score of each initial text is obtained based on the semantic information of the text composed of each initial text. By combining the identification confidence score and the semantic confidence score, it can be accurately judged whether each initial text needs to be corrected. By correcting the to-be-corrected text, the accuracy and stability of text image recognition are improved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor. Among them:
[0011] Figure 1 is a flowchart of an embodiment of the text image recognition method of the present application;
[0012] Figure 2is a flowchart corresponding to an embodiment of step S101 for step B;
[0013] Figure 3 is a flowchart corresponding to an embodiment of step S202;
[0014] Figure 4 is a flowchart corresponding to an embodiment of step S103;
[0015] Figure 5 is a flowchart corresponding to another embodiment of step S104;
[0016] Figure 6 is a flowchart corresponding to an embodiment of step S302;
[0017] Figure 7 is a structural diagram of an embodiment of the text image recognition system of the present application;
[0018] Figure 8 is a structural diagram of an embodiment of the electronic device of the present application;
[0019] Figure 9 is a structural diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0021] Please refer to Figure 1 , Figure 1 is a flowchart of an embodiment of the text image recognition method of the present application, which comprises:
[0022] S101: obtaining a text image comprising to-be-recognized text, and obtaining an initial character corresponding to the to-be-recognized text and a recognition confidence score corresponding to the initial character based on the text image.
[0023] In the present embodiment, step S101 specifically comprises the following steps:
[0024] A: in response to the need to recognize and extract to-be-recognized text on a paper file or the surface of an object, scanning through a corresponding recognition device to obtain a text image comprising the to-be-recognized text. The recognition device can be a camera, a scanner or other devices with a shooting function, etc. Alternatively, in response to the need to recognize a file in a picture format, directly taking the file in the picture format as the text image.
[0025] In an embodiment, after obtaining the text image, the text image is preprocessed to remove noise in the text image and enhance information of the to-be-recognized character part in the text image. Thus, it is helpful to separate the to-be-recognized character and the background in the text image, and improve the accuracy of text recognition.
[0026] Specifically, in response to the to-be-recognized character in the text image containing handwriting, the text image is first subjected to erosion processing and then to inflation processing, so as to remove stroke connection caused by continuous writing and remove small particle noise in the text image, and fill small holes in the text image. In response to the difference in brightness of different regions in the text image, the text image is adjusted by a local adaptive threshold method, so as to make the text image smoother. In addition, finally, the text image is subjected to edge enhancement processing, which can effectively improve the fullness of the boundary of the to-be-recognized text and reduce the difficulty of extracting the to-be-recognized character.
[0027] It should be noted that in other embodiments, in the process of preprocessing the text image, inflation processing can be performed first and then erosion processing can be performed according to different text images.
[0028] B: Feature extraction is performed on the obtained text image to obtain a feature vector corresponding to each to-be-recognized character in the text image, and the feature vector is processed by using an attention mechanism to obtain an initial character corresponding to the to-be-recognized character and a recognition confidence score corresponding to the initial character.
[0029] In an embodiment, please refer to Figure 2 , Figure 2 is a flowchart of an embodiment corresponding to step B in step S101. Specifically, this step includes:
[0030] S201: Feature extraction is performed on the text image to obtain a feature sequence corresponding to all to-be-recognized characters in the text image.
[0031] In this embodiment, step S201 includes: inputting the preprocessed text image into a feature extraction model to obtain a feature sequence corresponding to all to-be-recognized characters in the text image. The feature extraction model can be a ConvNet network.
[0032] It should be noted that after the feature extraction model obtains the feature sequence corresponding to all to-be-recognized characters, a 1*1 convolution layer is used to perform convolution processing on the feature sequence corresponding to all to-be-recognized characters, so as to transform the feature sequence corresponding to each to-be-recognized character to a target dimension. By transforming the feature sequence to the target dimension, it is helpful to encode the feature sequence corresponding to each to-be-recognized character subsequently.
[0033] S202: encode the feature sequence to obtain a first feature vector corresponding to each to-be-recognized character.
[0034] In the embodiment, step S202 includes: arranging each to-be-recognized character as a to-be-recognized text according to the position information of each to-be-recognized character in the text image, and inputting the feature sequence corresponding to each to-be-recognized character into the encoding model in sequence according to the order in the to-be-recognized text. The encoding model is a BiLSTM network, which includes a forward encoding sub-network and a reverse encoding sub-network.
[0035] It should be noted that the dimension of the vector input into the BiLSTM network is (Batchsize, Hide Layer, Max Length), and the feature extraction network in step S201 extracts features from the text image and converts the obtained feature sequence corresponding to each to-be-recognized character into the same dimension as the input vector of the BiLSTM network. Hide Layer represents the number of hidden layers in the BiLSTM network, and Max Length represents the maximum text length that the BiLSTM network can process.
[0036] Further, the process of encoding the feature sequence by using the BiLSTM network includes: first processing the feature sequence of each to-be-recognized character by using the forward encoding sub-network to obtain a first encoding vector corresponding to each to-be-recognized character, and then processing the first encoding vector of each to-be-recognized character by using the reverse encoding sub-network to obtain a second encoding vector corresponding to each to-be-recognized character. The first encoding vector contains the semantic information of the corresponding to-be-recognized character and all to-be-recognized characters before the to-be-recognized character; the second encoding vector contains the semantic information of the corresponding to-be-recognized character in the entire to-be-recognized text.
[0037] Specifically, the forward encoding sub-network and the reverse encoding sub-network in the BiLSTM network each have nodes consistent with the number of to-be-recognized characters. First, according to the position information of all to-be-recognized characters in the text image, the feature sequence corresponding to each to-be-recognized character is input into the node of the corresponding forward encoding sub-network in sequence. The current node encodes the input feature sequence to obtain the first encoding vector corresponding to the to-be-recognized character, and extracts the character semantic information of the corresponding to-be-recognized character as the input of the next node; the next node processes the input feature sequence and the character semantic information output by the previous node to obtain the first encoding vector of the corresponding to-be-recognized character. The first encoding vector output by the current node is obtained by combining the character semantic information output by the previous node and the feature sequence of the input to-be-recognized character, so that the first encoding vector output by the current node contains the character semantic information of all to-be-recognized characters corresponding to the previous nodes.
[0038] Further, the first encoding vector corresponding to each to-be-recognized character is input into the corresponding node in the reverse encoding sub-network in an order opposite to the arrangement order of the to-be-recognized characters. That is, the first encoding vector corresponding to the last to-be-recognized character is input into the reverse encoding sub-network first, and the first encoding vector corresponding to the first to-be-recognized character in the to-be-recognized text is input into the reverse encoding sub-network last. The current node in the reverse encoding sub-network encodes the input first encoding vector to obtain a second encoding vector corresponding to the to-be-recognized character, and extracts the character semantic information of the corresponding to-be-recognized character as the input of the next node; the next node processes the input second encoding vector and the character semantic information output by the previous node to obtain the second encoding vector corresponding to the to-be-recognized character. The second encoding vector output by the current node contains not only the forward semantic information of the corresponding to-be-recognized character in the to-be-recognized text, but also the reverse semantic information of the corresponding to-be-recognized character in the to-be-recognized text, thereby improving the accuracy of recognizing the to-be-recognized character.
[0039] Further, each second encoding vector output by the reverse encoding sub-network is taken as the first feature vector corresponding to each to-be-recognized character, and the first feature vector contains the phonetic information of the corresponding to-be-recognized character in the to-be-recognized text.
[0040] In a specific embodiment, please refer to Figure 3 , Figure 3 is a schematic diagram corresponding to an embodiment of step S202. Specifically, in response to the text image containing four to-be-recognized characters, the feature sequence corresponding to each to-be-recognized character is obtained, that is, “X1, X2, X3, X4”. Each feature sequence is input into the forward encoding sub-network 20 in the BiLSTM network in turn, that is, the feature sequence X1 corresponding to the first to-be-recognized character is input into the first node A1 in the forward encoding sub-network 20 to obtain the first encoding vector corresponding to the first to-be-recognized character and the character semantic information of the feature sequence X1; the character semantic information is input into the node A2 in the forward encoding sub-network 20, and the node A2 obtains the first encoding vector corresponding to the second to-be-recognized character according to the character semantic information output by the node A1 and the corresponding feature sequence X2, and the first encoding vector output by the node A2 contains the character semantic information of the to-be-recognized characters corresponding to the feature sequence X1 and the feature sequence X2. Similarly, the first encoding vectors corresponding to the feature sequence X3 and the feature sequence X4 are obtained according to the above method, and the specific process is not described in detail.
[0041] Further, the first encoding vectors output by the nodes in the forward encoding subnetwork 20 are input into the reverse encoding subnetwork 30 in an order opposite to the arrangement order of the to-be-recognized characters, i.e., the first encoding vector output by the node A4 is input into the node B4 in the reverse encoding subnetwork 30, to obtain a second encoding vector corresponding to the first encoding vector and extract the semantic information of the character in the first encoding vector to input the node B3. The node B3 outputs a second encoding vector according to the input character semantic information and the corresponding first encoding vector. The second encoding vector output by the node B3 contains the semantic information of the corresponding to-be-recognized character in the entire to-be-recognized text. Similarly, the second encoding vectors corresponding to the feature sequences X2 and X1 are obtained by the nodes B2 and B3 in the reverse encoding subnetwork 30 according to the above method.
[0042] Further, the second encoding vectors output by the nodes in the reverse encoding subnetwork 30 are taken as the first feature vectors of the corresponding to-be-recognized characters.
[0043] S203: performing feature mapping on the first feature vectors to obtain second feature vectors corresponding to the first feature vectors.
[0044] In the embodiment, the implementation process of step S203 includes: performing feature mapping on the first feature vectors corresponding to the to-be-recognized characters obtained in step S202 to obtain second feature vectors corresponding to the first feature vectors. The second feature vectors are used to represent the weights of the semantic information contained in the first feature vectors in the to-be-recognized text.
[0045] S204: obtaining initial characters and recognition confidence scores corresponding to the initial characters based on the first feature vectors and the second feature vectors.
[0046] In an embodiment, the implementation process of step S204 includes: in response to obtaining the first feature vectors and the second feature vectors corresponding to the to-be-recognized characters based on the above steps S201 to S203, decoding the first feature vectors corresponding to the to-be-recognized characters to obtain decoding vectors corresponding to the first feature vectors.
[0047] Specifically, the first feature vectors are input into the nodes in the decoding model in sequence to obtain decoding vectors corresponding to the to-be-recognized characters. In the embodiment, the decoding model is a Gate Recurrent Unit (GRU) structure.
[0048] Further, the product of the first feature vector and the corresponding second feature vector is normalized to obtain a weight matrix. The weight matrix is used to represent the weight of the semantic information contained in the first feature vector in the to-be-identified text, so as to help enhance the connection between the corresponding to-be-identified character and the context. In addition, in the embodiment, the product of the first feature vector and the corresponding second feature vector can be normalized by using a SoftMax algorithm.
[0049] Further, based on the weight matrix and the corresponding decoding vector, an initial character corresponding to the to-be-identified character and a recognition confidence score corresponding to the initial character are obtained. The recognition confidence score is obtained by normalizing the processing result of the attention mechanism.
[0050] Specifically, based on the attention mechanism, the weight matrix is multiplied by the corresponding decoding vector and normalized to obtain an alignment score between the to-be-identified character corresponding to the decoding vector and each alignment character in the character library. The alignment character corresponding to the highest alignment score is taken as the initial character corresponding to the to-be-identified character, and the corresponding alignment score is taken as the recognition confidence score of the initial character.
[0051] In the embodiment, the weight matrix corresponding to the to-be-identified character is obtained by using the attention mechanism, and based on the weight matrix and the corresponding decoding vector, the initial character corresponding to the to-be-identified character is obtained according to the semantic information of the to-be-identified character in the to-be-identified text, so as to improve the accuracy of recognizing the to-be-identified character.
[0052] Alternatively, in other embodiments, step S204 can also first perform feature splicing on the first feature vectors corresponding to all to-be-identified characters to obtain a text semantic vector containing the semantic information of the entire to-be-identified text, multiply the second feature vectors corresponding to each to-be-identified character with the text semantic vector to obtain the weight matrix corresponding to each to-be-identified character. The encoding vector corresponding to the to-be-identified character is multiplied by the corresponding weight matrix and normalized to obtain the alignment score between the corresponding to-be-identified character and each alignment character in the character library, and based on the alignment score, the initial character corresponding to the to-be-identified character and the recognition confidence score corresponding to the initial character are obtained.
[0053] S102: Based on the semantic information of each initial character, a semantic confidence score of the initial character is obtained.
[0054] In an embodiment, the implementation process of step S102 includes: based on the position information of each to-be-identified character in the text image, arranging all initial characters to obtain a recognized text corresponding to the text image.
[0055] Specifically, the obtained initial characters are arranged in a correct semantic order to obtain the recognized text corresponding to the text image. The correct semantic order is consistent with the arrangement order of the characters to be recognized in the text image.
[0056] Further, based on the text semantic information of the recognized text, a semantic confidence score of each initial character in the recognized text is obtained.
[0057] Specifically, a semantic analysis model is constructed, and each sentence in the recognized text is input into the constructed semantic analysis model in turn in units of periods. The semantic analysis model processes each input sentence to obtain a semantic confidence score of each initial character according to the semantic information of the initial character and the sentence in which the initial character is located. In this embodiment, the semantic analysis model is a BERT model. Obtaining the semantic confidence score helps to determine whether the initial character obtained by recognition is accurate according to the corresponding semantic information, so as to correct the initial character that is not accurately recognized to improve the recognition accuracy.
[0058] Alternatively, in other embodiments, step S102 can also be directly inputting the complete recognized text into the constructed semantic analysis model to obtain the semantic confidence score of each initial character according to the semantic information of each initial character and the recognized text. By obtaining the semantic confidence score of each initial character based on the complete recognized text, the accuracy of the semantic confidence score of the corresponding initial character is improved.
[0059] S103: determining at least part of the characters to be corrected from all the initial characters based on the recognition confidence score and the semantic confidence score of each initial character.
[0060] In an embodiment, the implementation process of step S103 includes: for each initial character, obtaining the mean value of the recognition confidence score and the semantic confidence score corresponding to the initial character, and taking the mean value as the comprehensive confidence score of the corresponding initial character. The comprehensive confidence score is compared with a second threshold value, and if the comprehensive confidence score is less than the second threshold value, the initial character corresponding to the comprehensive confidence score is taken as the character to be corrected. The second threshold value can be estimated or obtained by relevant technical personnel through trial and error.
[0061] In another embodiment, please refer to Figure 4 , Figure 4Fig. 1 is a schematic diagram of an embodiment corresponding to step S103. Specifically, step S103 comprises obtaining a first Gaussian distribution model based on the recognition confidence scores of the initial characters and obtaining a second Gaussian distribution model based on the semantic confidence scores of the initial characters. The calculation formulas of the first Gaussian distribution model and the second Gaussian distribution model are as follows:
[0062]
[0063]
[0064] wherein, represents the first Gaussian distribution model, μ1 represents the mean of the first Gaussian distribution model, represents the variance of the first Gaussian distribution model; represents the second Gaussian distribution model, μ2 represents the mean of the first Gaussian distribution model, represents the variance of the first Gaussian distribution model.
[0065] Further, please refer to Figure 4 , the mixed Gaussian distribution model is obtained based on the first Gaussian distribution model and the second Gaussian distribution model.
[0066] Specifically, in the present embodiment, the respective weights are set for the first Gaussian distribution model and the second Gaussian distribution model, and the mixed Gaussian distribution model is obtained based on the first Gaussian distribution model and the corresponding weight and the second Gaussian distribution model and the corresponding weight. That is, the first Gaussian distribution model is multiplied by the corresponding first weight to obtain a first product, the second Gaussian distribution model is multiplied by the corresponding second weight to obtain a second product, and the sum of the first product and the second product is taken as the mixed Gaussian distribution model. The sum of the first weight and the second weight is 1. The specific calculation formula of the mixed Gaussian distribution model is as follows:
[0067]
[0068] wherein, C represents the first Gaussian distribution model, H is 2; when h is 1, α h represents the first weight, when h is 2, α h represents the second weight.
[0069] Further, for the mixed Gaussian distribution model, the initial character corresponding to the value less than the threshold value is taken as the to-be-corrected character. The threshold value is the difference between the mean and the standard deviation in the mixed Gaussian distribution model.
[0070] Specifically, a difference between the mean and the standard deviation in the Gaussian mixture distribution model is obtained, and the difference is taken as a threshold value; a region in which the value is less than the threshold value in the Gaussian mixture distribution model is obtained, and an initial character corresponding to the region is taken as the to-be-corrected character. In this embodiment, the Gaussian mixture distribution model is constructed to combine the semantic confidence score and the recognition confidence score to determine whether the initial character is accurate, avoiding separate processing of the first Gaussian distribution model and / or the second Gaussian distribution model, improving the efficiency of detecting errors of the initial character, and saving the calculation cost.
[0071] S104: correcting the to-be-corrected character to obtain a target text corresponding to the text image.
[0072] In an embodiment, the specific implementation process of step S104 includes: obtaining a to-be-corrected sentence in which the to-be-corrected character is located, that is, taking the sentence in which the to-be-corrected character is located as the to-be-corrected sentence. The position of the to-be-corrected character in the to-be-corrected sentence is marked.
[0073] Further, the marked to-be-corrected sentence is input into a semantic analysis model to obtain sentence semantic information of the to-be-corrected sentence, and at least part of candidate characters at the marked position and a candidate score corresponding to each candidate character are obtained based on the sentence semantic information. The candidate character corresponding to the highest candidate score is taken as a target character at the marked position in the to-be-corrected sentence, and the target character is used to replace the corresponding to-be-corrected character to obtain a target text corresponding to the text image. The semantic analysis model includes a BERT model.
[0074] Specifically, the above process includes: deleting the to-be-corrected character at the marked position, inputting the to-be-corrected sentence after deleting the to-be-corrected character into the semantic analysis model, and the semantic analysis model obtaining at least part of candidate characters that meet the semantic information at the marked position from a corpus according to the sentence semantic information of the to-be-corrected sentence, and obtaining a corresponding candidate score according to the semantic information of each candidate character in the to-be-corrected sentence. The higher the candidate score is, the higher the matching degree of the corresponding candidate character with the corresponding to-be-corrected sentence is. Therefore, the candidate character corresponding to the highest candidate score is taken as a target character, and the target character is used to replace the to-be-corrected character, which can improve the accuracy of correcting the recognized text.
[0075] In addition, in this embodiment, before at least part of the candidate characters at the marked position are obtained by using the semantic analysis model, the following is included: constructing a semantic analysis model, and training the constructed semantic analysis model by using a training database to obtain a trained semantic analysis model.
[0076] Specifically, the training database contains a plurality of corpus data, the corpus data after the mask processing is input into the semantic analysis model, that is, part of the words in the corpus data are randomly deleted, and the corpus data after the deletion of the part of the words is input into the semantic analysis model. The semantic analysis model predicts the words at the blank positions in the corpus data based on semantic information, and compares the predicted words with the deleted words to adjust the parameters in the semantic analysis model, thereby obtaining the trained semantic analysis model.
[0077] In addition, it should be noted that in other embodiments, the sentence in which the to-be-corrected character is located and the sentences before and after the sentence in which the to-be-corrected character is located can all be regarded as to-be-corrected sentences in step S104, and all the to-be-corrected sentences are input into the semantic analysis model. By combining the sentence in which the to-be-corrected character is located and the sentences before and after the sentence, it is helpful to make the semantic analysis model predict at least part of the candidate characters based on more abundant semantic information, thereby improving the accuracy of the candidate characters.
[0078] The text image recognition method provided in the present application uses the attention mechanism to combine the semantic information of the to-be-recognized character in the to-be-recognized text to perform recognition, so as to obtain the initial character corresponding to the to-be-recognized character and the corresponding recognition confidence score. After obtaining the initial character, the semantic confidence score of each initial character is obtained based on the semantic information of the text composed of each initial character. By combining the recognition confidence score and the semantic confidence score, it can be accurately judged whether each initial character needs to be corrected. By correcting the to-be-corrected character, the accuracy and stability of the text image recognition are improved.
[0079] In another embodiment, please refer to Figure 5 , Figure 5 is a flowchart diagram corresponding to another embodiment of step S104. In the present embodiment, step S104 specifically includes:
[0080] S301: obtaining at least part of the candidate characters corresponding to the to-be-corrected character based on the semantic information of the sentence in which the to-be-corrected character is located.
[0081] In the present embodiment, step S301 includes: obtaining the to-be-corrected sentence in which the to-be-corrected character is located, and marking the position of the to-be-corrected character in the to-be-corrected sentence.
[0082] Further, the marked to-be-corrected sentence is input into the semantic analysis model to obtain the sentence semantic information of the to-be-corrected sentence, and at least part of the candidate characters at the marked position are obtained based on the sentence semantic information. The specific structure and construction method of the above-mentioned semantic analysis model can be referred to step S104, and will not be described in detail here.
[0083] S302: Determine the target text for replacing the text to be corrected based on the similarity between the text to be corrected and each candidate text, and obtain the target text corresponding to the text image.
[0084] In this embodiment, the implementation process of step S302 includes: in response to obtaining at least some candidate texts corresponding to the text to be corrected, splitting the text to be corrected and each candidate text to respectively obtain the stroke sequences of the text to be corrected and the candidate texts.
[0085] Specifically, according to the structures of the text to be corrected and the candidate texts, they can be split into tree structures. For example, in response to the text to be corrected or a candidate text being a left-right structure, first split it into a left part and a right part, and then separately split the left part and the right part; in response to the text to be corrected being an upper-middle-lower structure, first split it into an upper part, a middle part, and a lower part, and then separately split the upper part, the middle part, and the lower part individually. Finally, split the text to be corrected and each candidate text into corresponding stroke sequences.
[0086] In a specific embodiment, please refer to Figure 6 , Figure 6 which is a schematic diagram of a corresponding embodiment of step S302. As Figure 6 shown, for the character "贫 (pín)", first split it into "分 (fēn)" and "贝 (bèi)" according to the upper-lower structure; then split "分 (fēn)" into "八 (bā)" and "刀 (dāo)", and split "贝 (bèi)" into "冂 (jiōng)" and "人 (rén)"; finally, split "八 (bā)", "刀 (dāo)", "冂 (jiōng)", and "人 (rén)" into corresponding component strokes respectively, so as to obtain the stroke sequence corresponding to "贫 (pín)".
[0087] Furthermore, based on the stroke sequences, obtain the edit distances between the text to be corrected and each candidate text.
[0088] Specifically, compare the stroke sequence corresponding to the text to be corrected with the stroke sequences corresponding to each candidate text to calculate the minimum number of edit operations required for the text to be corrected to be converted into the corresponding candidate text by changing strokes. For example, for the Chinese characters "方 (fāng)" and "万 (wàn)", the stroke sequence of "方 (fāng)" is "丶一丿", and the stroke sequence of "万 (wàn)" is "一丿", so converting "方 (fāng)" into "万 (wàn)" only requires deleting the "丶" in the feature sequence, that is, the edit distance between "方 (fāng)" and "万 (wàn)" is considered to be 1.
[0089] Furthermore, obtain the similarity between the text to be corrected and each candidate text based on the edit distance, take the candidate text corresponding to the maximum similarity value as the target text, and use the target text to replace the text to be corrected.
[0090] In an embodiment, based on the edit distance between the to-be-corrected character and each candidate character, the candidate character corresponding to the smallest numerical value of the edit distance is taken as the target character, and the target character is used to replace the to-be-corrected character in the recognized text, so as to obtain the target text after correction.
[0091] Alternatively, in another embodiment, the similarity between the to-be-corrected character and each candidate character is calculated based on the edit distance between the to-be-corrected character and each candidate character. The candidate character corresponding to the largest numerical value of the similarity is taken as the target character. Wherein, the greater the edit distance, the smaller the corresponding similarity.
[0092] The embodiment determines the target character corresponding to the to-be-corrected character according to the similarity between the candidate character and the to-be-corrected character, and replaces the to-be-corrected character with the target character, so as to further improve the accuracy of the target text obtained by recognition.
[0093] In yet another embodiment, the text image recognition method provided by the present application can further include: a reference library is set in advance, and the reference library includes a plurality of reference words. Wherein, each reference word in the reference library can be classified according to the field. For example, in the reference library, the financial field includes reference words such as "capital structure, trend tracking, sovereign fund, private equity fund" and the like.
[0094] Further, after obtaining the recognized text corresponding to the text image and the to-be-corrected character through steps S101 to S103 in the above embodiment, the field of the recognized text is determined, and the recognized text is processed by word segmentation. If the to-be-corrected character belongs to a word with high similarity to the reference word in the reference library corresponding to the field, the to-be-corrected character is directly replaced by the word in the reference library. Or, the character corresponding to the to-be-corrected character in the reference word is taken as a candidate character, and based on the similarity between the to-be-corrected character and each candidate character, the target character for replacing the to-be-corrected character is determined. For specific process, please refer to step S302 in the above embodiment.
[0095] Please refer to Figure 7 , Figure 7 is a structural schematic diagram of an embodiment of the text image recognition system of the present application. The text image recognition system includes an identification module 40, a semantic analysis module 50, a processing module 60, and a correction module 70 which are coupled to each other.
[0096] Specifically, the identification module 40 is configured to obtain a text image including to-be-recognized characters, obtain initial characters corresponding to the to-be-recognized characters based on the text image, and obtain a recognition confidence score corresponding to the initial characters.
[0097] In an implementation scenario, the recognition module 40 performs feature extraction on the text image to obtain a feature sequence corresponding to all the to-be-recognized characters in the text image; encodes the feature sequence to obtain a first feature vector corresponding to each to-be-recognized character; performs feature mapping on the first feature vector to obtain a second feature vector corresponding to the first feature vector; and obtains an initial character and a recognition confidence score corresponding to the initial character based on the first feature vector and the second feature vector.
[0098] In the implementation scenario, the recognition module 40 obtains the initial character and the recognition confidence score corresponding to the initial character based on the first feature vector and the second feature vector, including: decoding the first feature vector to obtain a decoding vector; normalizing the product of the first feature vector and the second feature vector corresponding to the first feature vector to obtain a weight matrix; and multiplying the weight matrix and the decoding vector corresponding to the weight matrix based on an attention mechanism and normalizing the product to obtain the initial character and the recognition confidence score corresponding to the initial character.
[0099] The semantic analysis module 50 is configured to obtain a semantic confidence score of each initial character based on semantic information of the initial character.
[0100] In an implementation scenario, the semantic analysis module 50 arranges all the initial characters based on position information of each to-be-recognized character in the text image to obtain a recognized text corresponding to the text image; and obtains a semantic confidence score of each initial character in the recognized text based on text semantic information of the recognized text.
[0101] The processing module 60 is configured to determine at least part of the to-be-corrected characters from all the initial characters based on the recognition confidence score and the semantic confidence score corresponding to each initial character.
[0102] In an implementation scenario, the processing module 60 is configured to obtain a first Gaussian distribution model based on the recognition confidence score corresponding to each initial character; obtain a second Gaussian distribution model based on the semantic confidence score corresponding to each initial character; set a corresponding weight for the first Gaussian distribution model and the second Gaussian distribution model respectively, and obtain a mixed Gaussian distribution model based on the first Gaussian distribution model and the corresponding weight and the second Gaussian distribution model and the corresponding weight; and determine, for the mixed Gaussian distribution model, an initial character corresponding to a value less than a threshold value as a to-be-corrected character; wherein the threshold value is a difference between a mean value and a standard deviation in the mixed Gaussian distribution model.
[0103] The error correction module 70 is configured to correct the to-be-corrected characters to obtain a target text corresponding to the text image.
[0104] In an implementation scenario, the error correction module 70 obtains at least part of the candidate characters corresponding to the to-be-corrected character based on the semantic information of the sentence in which the to-be-corrected character is located; and determines the target character for replacing the to-be-corrected character based on the similarity between the to-be-corrected character and each candidate character, to obtain the target text corresponding to the text image.
[0105] In an implementation scenario, the error correction module 70 obtains the to-be-corrected sentence in which the to-be-corrected character is located, and marks the position of the to-be-corrected character in the to-be-corrected sentence; inputs the marked to-be-corrected sentence into a semantic analysis model to obtain the sentence semantic information of the to-be-corrected sentence, and obtains at least part of the candidate characters at the marked position based on the sentence semantic information; wherein the semantic analysis model includes a BERT network.
[0106] In an implementation scenario, the error correction module 70 is further configured to split the to-be-corrected character and each candidate character to obtain the stroke sequence corresponding to the to-be-corrected character and the candidate character, respectively; obtain the edit distance between the to-be-corrected character and each candidate character based on the stroke sequence; obtain the similarity between the to-be-corrected character and each candidate character based on the edit distance, take the candidate character corresponding to the maximum similarity value as the target character, and replace the to-be-corrected character with the target character.
[0107] Please refer to Figure 8 , Figure 8 The structure of an embodiment of the electronic device is shown in FIG. 1. The electronic device includes a memory 80 and a processor 90 coupled to each other. The memory 80 stores program instructions. The processor 90 is configured to execute the program instructions to implement the steps of the text image recognition method in the above embodiments. Specifically, the electronic device includes, but is not limited to, a desktop computer, a notebook computer, a tablet computer, a server, and the like, which are not limited herein. In addition, the processor 90 can also be referred to as a CPU (Center Processing Unit). The processor 90 can be an integrated circuit chip with signal processing capability. The processor 90 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. In addition, the processor 90 can be implemented by an integrated circuit chip together.
[0108] Please refer to Figure 9 , Figure 9A structural diagram of an embodiment of a computer readable storage medium proposed in the present application is shown in the figure. The computer readable storage medium 100 stores program instructions 110 capable of being executed by a processor, and the program instructions 110 are used to implement the text image recognition method in any of the above embodiments.
[0109] It should be noted that the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.
[0110] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0111] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0112] The above is only an embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation based on the content of the specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A method of text image recognition, characterized by, The method comprises: obtaining a text image comprising to-be-recognized characters, obtaining initial characters corresponding to the to-be-recognized characters based on the text image, and obtaining recognition confidence scores corresponding to the initial characters; obtaining semantic confidence scores of the initial characters based on semantic information of each initial character; determining at least part of to-be-corrected characters from all the initial characters based on the recognition confidence scores and the semantic confidence scores corresponding to each initial character; correcting the to-be-corrected characters to obtain target text corresponding to the text image; wherein the obtaining process of the initial characters comprises: obtaining a weight matrix corresponding to the to-be-recognized characters by using an attention mechanism, and obtaining the corresponding initial characters based on the weight matrix and a corresponding decoding vector according to semantic information of the to-be-recognized characters in the to-be-recognized text; wherein the decoding vector is obtained by decoding a first feature vector, and the first feature vector is obtained by encoding a feature sequence corresponding to the to-be-recognized characters; The determining step of at least part of the to-be-corrected characters comprises: obtaining a first Gaussian distribution model based on the recognition confidence scores corresponding to each initial character, and obtaining a second Gaussian distribution model based on the semantic confidence scores corresponding to each initial character; setting corresponding weights for the first Gaussian distribution model and the second Gaussian distribution model respectively, and obtaining a mixed Gaussian distribution model based on the first Gaussian distribution model, the corresponding weight of the first Gaussian distribution model, the second Gaussian distribution model and the corresponding weight of the second Gaussian distribution model; for the mixed Gaussian distribution model, the initial characters corresponding to values less than a threshold value are regarded as the to-be-corrected characters; wherein the threshold value is the difference between the mean value and the standard deviation in the mixed Gaussian distribution model.
2. The method of claim 1, wherein, The method comprises: performing feature extraction on the text image to obtain feature sequences corresponding to all the to-be-recognized characters in the text image; encoding the feature sequences to obtain first feature vectors corresponding to each to-be-recognized character; performing feature mapping on the first feature vectors to obtain second feature vectors corresponding to the first feature vectors; obtaining the initial characters and the recognition confidence scores corresponding to the initial characters based on the first feature vectors and the second feature vectors.
3. The method of claim 2, wherein, The method comprises: decoding the first feature vectors to obtain decoding vectors; normalizing the product of the first feature vectors and the corresponding second feature vectors to obtain a weight matrix; based on the attention mechanism, multiplying the weight matrix and the corresponding decoding vectors and normalizing to obtain the initial characters corresponding to the to-be-recognized characters and the recognition confidence scores of the initial characters.
4. The method of claim 1, wherein, The method comprises: arrange all the initial characters based on position information of each of the to-be-recognized characters in the text image, to obtain the recognized text corresponding to the text image; obtain a semantic confidence score of each of the initial characters in the recognized text based on text semantic information of the recognized text.
5. The method of claim 1, wherein, The correcting the to-be-corrected character to obtain the target text corresponding to the text image comprises: obtaining at least part of the candidate character corresponding to the to-be-corrected character based on semantic information of a sentence in which the to-be-corrected character is located; determining a target character for replacing the to-be-corrected character based on similarity between the to-be-corrected character and each of the candidate characters, to obtain the target text corresponding to the text image.
6. The method of claim 5, wherein, The obtaining at least part of the candidate character corresponding to the to-be-corrected character based on semantic information of a sentence in which the to-be-corrected character is located comprises: obtaining a to-be-corrected sentence in which the to-be-corrected character is located, and marking a position of the to-be-corrected character in the to-be-corrected sentence; inputting the marked to-be-corrected sentence into a semantic analysis model to obtain sentence semantic information of the to-be-corrected sentence, and obtaining at least part of the candidate character at the marked position based on the sentence semantic information; wherein the semantic analysis model comprises a BERT network.
7. The method of claim 5, wherein, The determining a target character for replacing the to-be-corrected character based on similarity between the to-be-corrected character and each of the candidate characters, to obtain the target text corresponding to the text image, comprises: splitting the to-be-corrected character and each of the candidate characters to obtain stroke sequences corresponding to the to-be-corrected character and the candidate characters respectively; obtaining an edit distance between the to-be-corrected character and each of the candidate characters based on the stroke sequences; obtaining the similarity between the to-be-corrected character and each of the candidate characters based on the edit distance, taking the candidate character corresponding to the maximum similarity value as the target character, and replacing the to-be-corrected character with the target character.
8. A text image recognition system, characterized by, comprise: an identification module, configured to obtain a text image comprising to-be-recognized characters, and obtain initial characters corresponding to the to-be-recognized characters and a recognition confidence score corresponding to the initial characters based on the text image; a semantic analysis module, configured to obtain a semantic confidence score of each of the initial characters based on semantic information of each of the initial characters; a processing module, configured to determine at least part of to-be-corrected characters from all the initial characters based on the recognition confidence score and the semantic confidence score corresponding to each of the initial characters; a correction module, configured to correct the to-be-corrected character to obtain a target text corresponding to the text image; The initial character obtaining process comprises: obtaining a weight matrix corresponding to the to-be-recognized character by using an attention mechanism, and obtaining the initial character according to semantic information of the to-be-recognized character in the to-be-recognized text based on the weight matrix and a corresponding decoding vector; the decoding vector is obtained by decoding a first feature vector, and the first feature vector is obtained by encoding a feature sequence corresponding to the to-be-recognized character. The determining step of the at least part of the to-be-corrected character comprises: obtaining a first Gaussian distribution model based on the recognition confidence score corresponding to each initial character; obtaining a second Gaussian distribution model based on the semantic confidence score corresponding to each initial character; setting corresponding weights for the first Gaussian distribution model and the second Gaussian distribution model respectively, and obtaining a mixed Gaussian distribution model based on the first Gaussian distribution model and the corresponding weight and the second Gaussian distribution model and the corresponding weight; and regarding the initial character corresponding to a value less than a threshold value in the mixed Gaussian distribution model as the to-be-corrected character; wherein the threshold value is a difference between a mean value and a standard deviation in the mixed Gaussian distribution model.
9. An electronic device, comprising: The device comprises a memory and a processor coupled to each other, the memory stores program instructions, and the processor is configured to execute the program instructions to implement the text image recognition method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The device stores program instructions capable of being executed by the processor, and the program instructions are configured to implement the text image recognition method according to any one of claims 1-7.
Citation Information
Patent Citations
Text error correction method, system and device and readable storage medium
CN112016310A
Character recognition method and device, electronic equipment and storage medium
CN112686263A