End-to-end text recognition model training method, text recognition method and device
Through the end-to-end text recognition model training method, the combination of feature extraction module, feature encoder and feature decoder is used to solve the problem of high time complexity of text recognition in the prior art, and efficient text recognition is achieved.
Patent Information
- Application Number
- CN202210704167.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-06-21
AI Technical Summary
The existing end-to-end Pix2seq method based on the Transformer model has a high time complexity in text recognition and is difficult to apply in actual scenarios.
An end-to-end text recognition model training method is proposed. Through the combination of feature extraction module, feature encoder and feature decoder, multiple iterative training is performed using the input feature vector, character position vector and output feature vector of the target text line image to reduce the time complexity.
It realizes the simultaneously obtaining the position and content of each character in the text line within a decoding step, reducing the complexity of text recognition and improving the efficiency of text recognition.
Smart Images

Figure CN115082937B_ABST
Abstract
Claims
1. An end-to-end text recognition model training method, It is characterized in that The method comprises: Input the target text line image into the feature extraction module to obtain the target input feature vector; Obtaining a target character position vector corresponding to the target text line image, and inputting the target input feature vector and the target character position vector into a feature encoder to obtain a first feature vector; Based on the label corresponding to the target text line image, a target output feature vector is obtained; the label corresponding to the target text line image includes the real character position and real text content of each character in the target text line image; the target output feature vector is composed of the real position vector and the real text content vector corresponding to each character in the target text line image; Repeat the operation on the first feature vector to obtain a second feature vector; the dimension of the second feature vector is the same as the dimension of the target output feature vector; Input the second feature vector, the target output feature vector and the target character position vector into a feature decoder to obtain a prediction probability distribution result corresponding to the target text line image; the prediction probability distribution result includes a prediction position probability distribution and a prediction text content probability distribution of each character in the target text line image; Obtaining a loss value according to a label corresponding to the target text line image and a predicted probability distribution result corresponding to the target text line image; The feature extraction module, the feature encoder and the feature decoder are trained based on the loss value, and the steps of inputting the target text line image into the feature extraction module, obtaining the target input feature vector and subsequent steps are repeatedly performed until a preset condition is met.
2. The method according to claim 1, It is characterized in that The step of obtaining a target output feature vector based on a label corresponding to the target text line image includes: Constructing a target dictionary, converting the label corresponding to the target text line image into a corresponding dictionary value in the target dictionary based on the target dictionary, and obtaining a numerical vector corresponding to the label; The numerical vector corresponding to the label is converted into a vector in a high-dimensional space to obtain a target output feature vector.
3. The method according to claim 1, It is characterized in that The repeatedly operating the first feature vector to obtain a second feature vector includes: A repetition number parameter in a repetition operation function is set, and the first feature vector is input into the set repetition operation function to obtain a second feature vector.
4. The method according to claim 1, It is characterized in that The feature encoder includes a first multi-head attention module and a first feedforward network module; the step of inputting the target input feature vector and the target character position vector into the feature encoder to obtain a first feature vector includes: Concatenate the target input feature vector and the target character position vector, input the concatenated vector into the first multi-head attention module, and obtain an output vector of the first multi-head attention module; The output vector of the first multi-head attention module is input into the first feedforward network module to obtain the first feature vector output by the first feedforward network module.
5. The method according to claim 1, It is characterized in that The feature decoder includes a second multi-head attention module, a third multi-head attention module and a second feedforward network module, and the end-to-end text recognition model also includes a regression module; the second feature vector, the target output feature vector and the target character position vector are input into the feature decoder to obtain the predicted probability distribution result corresponding to the target text line image, including: Concatenate the target output feature vector and the target character position vector, input the concatenated vector into the second multi-head attention module, and obtain a third feature vector output by the second multi-head attention module; Input the second eigenvector and the third eigenvector into the third multi-head attention module to obtain a fourth eigenvector output by the third multi-head attention module; Inputting the fourth eigenvector into the second feedforward network module to obtain a fifth eigenvector output by the second feedforward network module; The fifth feature vector is input into the regression module to obtain the predicted probability distribution result corresponding to the target text line image output by the regression module.
6. The method according to any one of claims 1 to 5, It is characterized in that The obtaining of the loss value according to the label corresponding to the target text line image and the predicted probability distribution result corresponding to the target text line image includes: Based on the label corresponding to the target text line image, obtaining a true probability distribution result of the target text line image; Based on the true probability distribution result of the target text line image and the predicted probability distribution result corresponding to the target text line image, a cross entropy loss is obtained.
7. The method according to any one of claims 1 to 5, It is characterized in that The step of inputting the target text line image into the feature extraction module to obtain the target input feature vector comprises: Acquire a target text line image, perform a scaling operation and / or a padding operation on the target text line image, and obtain a preprocessed target text line image; The preprocessed target text line image is input into the feature extraction module to obtain the target input feature vector.
8. An end-to-end text recognition method, It is characterized in that The method comprises: Get the character position vector corresponding to the image to be recognized; Inputting the image to be recognized and the character position vector corresponding to the image to be recognized into an end-to-end text recognition model, and obtaining a probability distribution result of each character in the image to be recognized output by the end-to-end text recognition model; the probability distribution result includes the position probability distribution of the character and the text content probability distribution; Obtaining a character detection result and a character recognition result for each character in the image to be recognized according to a probability distribution result of each character in the image to be recognized; The end-to-end text recognition model is trained according to the end-to-end text recognition model training method according to any one of claims 1-7.
9. An end-to-end text recognition model training device, It is characterized in that The device comprises: A first acquisition unit, used for inputting a target text line image into a feature extraction module to acquire a target input feature vector; A second acquisition unit is used to acquire a target character position vector corresponding to the target text line image, and input the target input feature vector and the target character position vector into a feature encoder to obtain a first feature vector; A third acquisition unit is used to acquire a target output feature vector based on the label corresponding to the target text line image; the label corresponding to the target text line image includes the real character position and real text content of each character in the target text line image; the target output feature vector is composed of the real position vector and the real text content vector corresponding to each character in the target text line image; a fourth acquisition unit, configured to repeatedly operate the first feature vector to acquire a second feature vector; the dimension of the second feature vector is the same as the dimension of the target output feature vector; An input unit, used to input the second feature vector, the target output feature vector and the target character position vector into a feature decoder to obtain a prediction probability distribution result corresponding to the target text line image; the prediction probability distribution result includes a prediction position probability distribution and a prediction text content probability distribution of each character in the target text line image; A fifth acquisition unit, configured to acquire a loss value according to a label corresponding to the target text line image and a predicted probability distribution result corresponding to the target text line image; A training unit is used to train the feature extraction module, the feature encoder and the feature decoder based on the loss value, and repeatedly execute the steps of inputting the target text line image into the feature extraction module, obtaining the target input feature vector and subsequent steps until a preset condition is met.
10. An end-to-end text recognition device, It is characterized in that The device comprises: A first acquisition unit, used to acquire a character position vector corresponding to the image to be recognized; A second acquisition unit is used to input the image to be recognized and the character position vector corresponding to the image to be recognized into an end-to-end text recognition model to obtain a probability distribution result of each character in the image to be recognized output by the end-to-end text recognition model; the probability distribution result includes the position probability distribution of the character and the text content probability distribution; A third acquisition unit, configured to acquire a character detection result and a character recognition result of each character in the image to be recognized according to a probability distribution result of each character in the image to be recognized; The end-to-end text recognition model is trained according to the end-to-end text recognition model training method according to any one of claims 1-7.
11. An electronic device, It is characterized in that include: one or more processors; a storage device having one or more programs stored thereon, When the one or more programs are executed by the one or more processors, the one or more processors implement the end-to-end text recognition model training method as described in any one of claims 1-7, or the end-to-end text recognition method as described in claim 8.
12. A computer readable medium, It is characterized in that A computer program is stored thereon, wherein when the program is executed by a processor, the end-to-end text recognition model training method as described in any one of claims 1 to 7, or the end-to-end text recognition method as described in claim 8 is implemented.
13. A computer program product, It is characterized in that The computer program product includes a computer program / instructions, which, when executed by a processor, implements the end-to-end text recognition model training method as described in any one of claims 1 to 7, or the end-to-end text recognition method as described in claim 8.
Citation Information
Patent Citations
Text recognition model training method and device, text recognition method and device, equipment and medium
CN114022882A
KR20210040851A