A self-supervised text recognition method based on a character movement task
By introducing character movement tasks and self-supervised learning methods in the text recognition model, the problem of relying on labeled data and prior knowledge in the prior art is solved, and efficient text recognition feature learning and recognition accuracy are achieved.
Patent Information
- Application Number
- CN202211017001.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-23
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-08-23
AI Technical Summary
The existing fully supervised text recognition model relies on a large amount of labeled data, and the self-supervised learning method for unlabeled data has oversegment and undersegment in the serialization process, and has failed to make full use of the unique prior knowledge of handwritten text images.
A self-supervised text recognition method based on character movement task is proposed. Through character positioning, character selection and character movement, neural network is constructed for pre-training, combined with contrast learning and classification tasks, the feature representation of text images is learned, and data augmentation and momentum update technology is used to improve the robustness and accuracy of the model.
It realizes the rapid learning of feature representation of text images without manual annotation of data, improves the convergence speed and recognition accuracy of downstream text recognition tasks, and has high application value.
Smart Images

Figure CN115439859B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pattern recognition and artificial intelligence, and particularly relates to a self-supervised text recognition method based on a character movement task. Background Art
[0002] Text recognition is of great significance for the digitization of various paper documents. Currently, most text recognition models are based on fully supervised training methods, which rely on a large amount of labeled data, and the labeling of data requires a lot of manpower and material resources. At the same time, with the development of Internet technology, the acquisition of data is easier, and the scale of data can even reach the order of trillions. For these unlabeled data, it is unrealistic to perform manual labeling. Therefore, it is very necessary to explore a self-supervised training method that does not require the use of manual labeling.
[0003] In recent years, with the development of various deep learning technologies, self-supervised learning methods based on contrastive learning have shown great potential in the detection and recognition of general targets. It learns the feature representation of general targets by performing contrastive learning on images with different data augmentation methods of the same image, which can accelerate the convergence speed of downstream tasks and also achieve good task effects with a small amount of training data.
[0004] Currently, for handwritten text, the self-supervised learning method SeqCLR method (Aberdam A, Litman R, Tsiper S, et al. Sequence-to-sequence contrastive learning for text recognition[C]. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 15302-15312.) utilizes the characteristics of the sequential composition of text, adds an instance mapping function after feature extraction to serialize the feature vectors of the text, and then performs contrastive learning to learn the feature representation of the text. However, the SeqCLR method will have the phenomena of over-segmentation and under-segmentation of the original text during the serialization process, so this serialization method is actually not accurate enough. In addition, the SeqCLR method does not make good use of the unique prior knowledge of the original text image. Summary of the Invention
[0005] The objective of the present invention is to effectively utilize the feature representation ability of the deep network model and the unique attributes of the handwritten text image, learn the distribution of the text image data samples, and thus implement a self-supervised text recognition method. This solution has characteristics such as accelerating the convergence of downstream tasks and improving the recognition accuracy, and has high application value.
[0006] The present invention is implemented by at least one of the following technical solutions.
[0007] A self-supervised text recognition method based on the character movement task includes the following steps:
[0008] (1) Obtain the image of the handwritten word through an electronic device;
[0009] (2) Perform data preprocessing on the word image;
[0010] (3) Perform character positioning, character selection, and character movement on the word image, and then determine the label of the character movement;
[0011] (4) Construct a neural network for pre-training, which is to perform contrastive learning on different data-augmented images of the same original image and classify the images after character movement;
[0012] (5) Read the encoder parameters of the neural network pre-trained in step (4) into the encoder of the text recognition model, and then use the text recognition model to adjust the handwritten word image and label.
[0013] Furthermore, use an electronic device capable of handwritten input to obtain the grayscale image of the handwritten word.
[0014] Furthermore, the preprocessing in step (2) is to perform data augmentation T(·) on the original image I, including affine transformation, stroke jitter, stroke coverage, and stroke thickness change, where the relevant parameters of each augmentation method are randomly selected within the set range each time; randomly select a set of parameters t 1 , to obtain the first image Randomly select a set of parameters t 2 , to obtain the second image Then perform image size adjustment on the first image I k and the second image I q , adjust them to H×W, where H is the image height and W is the image width; then normalize the first image I k to [0, 1].
[0015] Furthermore, the character positioning in step (3) includes the following steps:
[0016] (311) For the second image I q, the vertical projection distribution Sta is obtained through vertical projection;
[0017] First, perform adaptive binarization on the second image I q , then normalize it to [0, 1]. At this time, the value of the area where the text is located is 1, and then perform row summation to obtain the vertical projection distribution Sta;
[0018] (312) Set the numbers less than the value m in the vertical projection distribution Sta to zero, where m takes the second smallest projection value in Sta, and then obtain the character block region set U = {u 1 , u 2 ,..., u i ..., u l} from the vertical projection distribution Sta, where u i is defined as the character block region, that is, the continuous region with non-zero projection value; l represents the number of character block regions.
[0019] Furthermore, the character selection in step (3) includes the following steps:
[0020] (321) Randomly select two positions loc b and loc a from the character block region set U as the position where the character is located before movement and the target position of the character movement respectively. The selection of loc b and loc a is divided into the following three cases:
[0021] If |U| = 0, it means there is no character block region. Let m be the smallest projection value in Sta, and return to step (312) to continue obtaining the character block region set U;
[0022] If |U| = 1, it means there is only one character block region, that is, U = u 1 . At this time, select one position from the first 40% position h 1 and the last 40% position h 1 of u 2 , and then randomly use these two positions as loc b and loc a ;
[0023] If |U| ≥ 2, it means there are two or more character block regions. At this time, randomly select two character block regions u b and u a from U as the initial character block region where the character is located and the target character block region for movement respectively; then randomly select a position from u b as loc b , and randomly select a position from u a as loc a;
[0024] (322) Determine the width of the characters to be moved, and finally select the character images to be moved; the initial half-width of the character images to be moved is set as:
[0025]
[0026] where W is the width of the second image I q ; set the target position loc a of the character movement and the minimum distance from the image boundary as border a , and the minimum distance from the position loc b where the character is located before movement to the image boundary is border b , and the half-width of the character image to be moved is:
[0027] w move = min(w ini , border a , border b ) (1)
[0028] Select the character images to be moved as:
[0029] img b = I q [0:H, loc b - w move : loc b + w move
[0030] where H is the height of the second image I q , and w move is the half-width of the character image to be moved.
[0031] Furthermore, the character movement in step (3) includes the following steps:
[0032] The original picture of the target position of the character movement is:
[0033] img a = I q [0:H, loc a - w move : loc a + w move
[0034] Superimpose the character image img b to be moved onto the img q of the second image I a at a ratio of 1 - λ, and keep the other positions of the second image I q unchanged, then obtain the moving image MI, that is
[0035] img a = λimg a +(1 - λ)img b (2)
[0036] Where λ represents the superposition ratio, 0 < λ < 1.
[0037] Furthermore, the specific label for determining the character movement is:
[0038] The pixel value pixel of the character movement move = loc a - loc b , when pixel move < 0, it means the character moves to the left; when pixel move > 0, it means the character moves to the right; define the character movement task as a classification task, and let the classification label label = pixel move + W, where W is the width of the second image I q .
[0039] Furthermore, the neural network includes an encoding mapping module Q, a momentum encoding mapping module K, and a multi-layer perceptron;
[0040] The encoding mapping module Q includes an encoder E and a mapper, and the encoding mapping module Q is trained according to the stochastic gradient descent optimizer; input the output features of the encoder E in the encoding mapping module Q into the multi-layer perceptron, and then classify the output feature vectors to predict the pixel value of the character movement in the image;
[0041] The momentum encoding mapping module K has the same network structure as the encoding mapping module Q, and uses the parameters of the encoding mapping module Q for momentum update; let the parameters of the encoder E and the mapper in the encoding mapping module Q be θ q , and the parameters of the encoder and the mapper in the momentum encoding mapping module K be θ k , and the formula for momentum update is:
[0042] nθ k +(1 - n)θ q → θ k (3)
[0043] Where n represents the momentum magnitude, 0 < n < 1.
[0044] Furthermore, the pre-training of the neural network includes: the first image I obtained after data augmentation and the first image I obtained after data augmentation k and the first image I obtained after data augmentation The original image and the moving image MI obtained after character movement respectively pass through the momentum encoding mapping module K and the encoding mapping module Q, and then the loss value is calculated. The formula of the loss function is:
[0045]
[0046] where C is the length of the negative sample; τ is a hyperparameter; MI q is the feature vector after passing through the encoding mapping module Q; k + is the feature vector after passing through the momentum encoding mapping module K, which is the positive sample of MI q and comes from the same original image as MI q ; k i is the feature vector after passing through the momentum encoding mapping module K, which is the negative sample of MI q i.e., it does not come from the same original image as MI q , and i = 1...C;
[0047] For the negative sample, a size of the negative sample is preset, and then the feature vectors after each passing through the momentum encoding mapping module K are stored. After reaching the preset amount of negative samples, the earliest stored feature vector is deleted, and then the new feature vector is stored;
[0048] For the moving image MI obtained after data augmentation and character movement, in addition to using the output vector of its passing through the encoding mapping module Q to participate in the calculation of formula (4), the output features of the encoder E in the encoding mapping module Q are also input into a multi-layer perceptron, and then the feature vectors output by the multi-layer perceptron are classified to predict the pixel value of character movement in the image. The classification formula is:
[0049]
[0050] where N is the batch size; y i is the one-hot vector of the character movement label corresponding to the moving image MI; p i is the probability vector predicted by the multi-layer perceptron, and the calculation formula is:
[0051]
[0052] where F(MI i ) is the output feature vector of the encoder E of the encoding mapping module Q and the multi-layer perceptron for the i-th moving image MI in a batch; MI i is the i-th moving image MI in a batch; MI j is the j-th moving image MI in a batch; The final total loss function is where α is a hyperparameter.
[0053] Furthermore, the text recognition model adopts an encoder-decoder structure, and the structure of the encoder of the text recognition model is the same as that of the encoder E of the coding mapping module Q;
[0054] In the training process, it is necessary to first read the encoder parameters of the neural network pre-trained in step (4) into the encoder of the text recognition model, and the parameters of the decoder are randomly initialized, and then the entire text recognition model is fine-tuned and trained according to the input handwritten word image and the corresponding label.
[0055] Compared with the existing technologies, the beneficial effects of the present invention are as follows:
[0056] (1) In view of the unique attributes of handwritten text images, the present invention proposes a character movement task. By moving the characters in the text image and then enabling the network to predict the pixel values of the character movement, character-level feature learning is achieved.
[0057] (2) In the pre-training stage of the present invention, the feature representations of handwritten text images are jointly learned through two levels, namely the character level and the whole-word level, so as to learn effective text image representations.
[0058] (3) In the pre-training stage, there is no need to use manually labeled data, thus saving a large amount of manpower and material resources, and thousands of unlabeled data can be utilized, which has great application value.
[0059] (4) The encoder parameters obtained in the pre-training stage of the present invention can accelerate the convergence speed of the downstream text recognition task and achieve better recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a schematic flowchart of a self-supervised text recognition method based on a character movement task in an embodiment;
[0061] Figure 2 is a schematic diagram of a deep model in an embodiment;
[0062] Figure 3 is an example diagram of character movement in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] The present invention will be further described below in conjunction with the embodiments and the drawings, but the embodiments of the present invention are not limited thereto.
[0064] Embodiment 1
[0065] A self-supervised text recognition method based on a character movement task in this embodiment is as Figure 1 shown, and includes the following steps:
[0066] (1) Data acquisition: Use electronic devices such as mobile phones and tablets that support handwritten input to obtain grayscale images of handwritten words. Since subsequent character positioning is performed through vertical projection, languages in which letters are horizontally serialized into words, such as English, German, French, etc., are obtained here.
[0067] (2) Data processing, including the following steps:
[0068] (2-1) Perform data augmentation on the original image I twice and resize the image to H×W, where H is the image height and W is the image width, and then obtain the first image I k and the second image I q . Data augmentation includes affine transformation, stroke jitter, stroke coverage, and stroke thickness variation, etc. The relevant parameters of each augmentation method are randomly selected within a specific range each time. For example, the scaling range of affine transformation is [0.5, 1.05], the jitter range of stroke jitter is [0.2, 0.5] of the image width, the rotation superposition angle of stroke coverage is [-8, 8], and the multiple range of the original thickness for stroke thickness variation is [0.2, 3].
[0069] (2-2) For the second image I q , obtain the vertical projection distribution Sta through vertical projection.
[0070] First, perform adaptive binarization on the second image I q , and then normalize it to [0, 1]. At this time, the value of the area where the text is located is 1, and then perform row summation to obtain the vertical projection distribution Sta, which can reflect the approximate positions of each character in the word image.
[0071] (2-3) To roughly eliminate the problem of handwritten stroke adhesion, set to zero the numbers in the vertical projection distribution Sta that are less than the value m, where m can take the second smallest projection value in Sta. Then, from the vertical projection distribution Sta, the character block region set U = {u 1 , u 2 ,..., u i ..., u l} can be obtained, where u i is defined as the character block region, that is, the continuous region with non-zero projection values; l represents the number of character block regions.
[0072] (2-4) Randomly select two positions loc b and loc a from the character block region set U as the position where the character is located before movement and the target position of the character movement, respectively. The selection of loc b and loc a is divided into the following three cases:
[0073] If |U| = 0, it indicates that there is no character block region. Let m be the minimum projection value in Sta, and return to step (2-3) to continue obtaining the character block region set U.
[0074] If |U| = 1, it means there is only one character block region, i.e., U = u 1 . At this time, from the first 40% position h 1 and the last 40% position h 1 of u 2 , select one position each, and then randomly use these two positions as loc b and loc a .
[0075] If |U| ≥ 2, it means there are two or more character block regions. At this time, randomly select two character block regions u b and u a from U as the character block regions before and after character movement respectively. Then randomly select a position from u b as loc b , and randomly select a position from u a as loc a .
[0076] (2-5) Determine the width of character movement, and then superimpose the characters to be moved on the target position of movement.
[0077] The initial half-width of character movement is where W is the width of the image I q . Set the minimum distance between the position loc a and the image boundary as border a , and the minimum distance between the position loc b and the image boundary as border b . The final half-width of character movement is
[0078] w move = min(w ini , border a , border b ) (1)
[0079] At this time, the character picture to be moved is:
[0080] img b = I q [0:H, loc b - w move :loc b + w move
[0081] where H is the height of the second image I q The height, w move is the half-width of the character image to be moved, obtained from formula (1).
[0082] The original picture of the target position where the character moves is:
[0083] img a = I q [0:H, loc a - w move : loc a + w move
[0084] Finally, the character image img to be moved b is superimposed on the second image I at a ratio of 1 - λ q of img a on it. Other positions of the second image I q remain unchanged, and then the moved image MI is obtained, that is
[0085] img a = λimg a + (1 - λ)img b (2)
[0086] where λ represents the superposition ratio (0 < λ < 1).
[0087] (2 - 6) Determine the label of character movement.
[0088] The pixel value pixel of character movement move = loc a - loc b . When pixel move < 0, it means the character moves to the left; when pixel move > 0, it means the character moves to the right. Here, the character movement task is defined as a classification task, and let the classification label label = pixel move + W, where W is the width of the image I q . Since the width of the image I q is adjusted to W before character movement, the maximum value of the left and right movement pixels is W, and the number of classification categories is 2W + 1.
[0089] (3) Network pre-training, including the following steps:
[0090] (3 - 1) Construct a neural network, including an encoder, a mapper, and a multi-layer perceptron. The encoder is shown in Table 1. The mapper includes fully connected layers with 512 and 128 nodes, shown in Table 2. The multi-layer perceptron structure is shown in Table 3, including fully connected layers with 512 and 201 nodes.
[0091] Table 1 Encoder Structure
[0092]
[0093]
[0094] Table 2 Mapper Structure
[0095] Network layer Specific settings Feature map size Fully connected layer Number of nodes: 512 512×512 Fully connected layer Number of nodes: 128 512×128
[0096] Table 3 Multilayer Perceptron Structure
[0097] Network layer Specific settings Feature map size Fully connected layer Number of nodes: 512 512×512 Fully connected layer Number of nodes: 201 512×201
[0098] First, the encoder E and the mapper are combined into an encoding and mapping module Q, which is trained according to the stochastic gradient descent optimizer. The momentum encoding and mapping module K with the same network structure as module Q uses the parameters of module Q for momentum update. Let the parameters of module Q be θ q , and the parameters of module K be θ k , and the update formula is
[0099] nθ k +(1 - n)θ q →θ k (3)
[0100] where n represents the momentum magnitude (0 < n < 1).
[0101] (3 - 2) is pre-trained. Image I k and image MI pass through module K and module Q respectively, and then the loss value is calculated. The formula of the loss function is
[0102]
[0103] where C is the length of the negative samples, τ is a hyperparameter; MI q is the feature vector after passing through the encoding and mapping module Q; k + is the feature vector after passing through the momentum encoding and mapping module K, and it is the positive sample of MI q , that is, it comes from the same original image as MI q ; k i (i = 1...C) is the feature vector after passing through the momentum encoding and mapping module K, and it is the negative sample of MI q , that is, it does not come from the same original image as MI q .
[0104] For the negative samples, a size of the negative samples is preset, and then the feature vectors after each passing through module K are stored. After reaching the preset amount of negative samples, the earliest batch of feature vectors will be deleted, and then new feature vectors will be stored.
[0105] For the image MI, the output features of the encoder E in the encoding mapping module Q are input into a multi-layer perceptron, and then the output feature vectors are classified to predict the pixel values of character movement in the image. The classification formula is
[0106]
[0107] where N is the batch size; y i is the one-hot vector of the character movement label corresponding to the moving image MI; p i is the probability vector predicted by the multi-layer perceptron, and the calculation formula is:
[0108]
[0109] where F(MI i ) is the output feature vector of the i-th moving image MI in a batch passing through the encoder E of the encoding mapping module and the multi-layer perceptron; MI i is the i-th moving image MI in a batch; MI j is the j-th moving image MI in a batch.
[0110] Finally, the total loss function is where α is a hyperparameter. Then, the network is pre-trained according to the above settings.
[0111] (4) Read the encoder parameters of the pre-trained neural network into the encoder of the text recognition model, and then use the text recognition model to fine-tune the handwritten word image and label.
[0112] The text recognition model adopts an "encoder-decoder" structure. The structure of the encoder is the same as that of the encoder model pre-trained in step (3). The decoder can adopt a sequence decoder based on CTC, or based on Attention, or based on Transformer. For example, the Attention-based decoder is an Attention model with 256 hidden layer nodes.
[0113] The fine-tuning training process needs to first read the encoder parameters of the neural network pre-trained in step (3) into the encoder of the text recognition model, and the parameters of the decoder are randomly initialized. Then, the entire text recognition model is fine-tuned according to the input handwritten word image and the corresponding label. In Figure 2 the shown example, the schematic diagram of the model of the present invention is displayed.
[0114] The present invention first improves this over-segmentation and under-segmentation phenomenon, that is, without performing the serialization process, directly learning the overall representation in the text image from the whole-word level contrast learning. At the same time, based on the unique prior attributes of handwritten text images, the present invention also proposes a character-level character movement task, that is, moving the characters in the image and then predicting the moved pixel values. The present invention uses this character-level movement task to assist the whole-word level contrast learning to effectively learn the feature representation in the text image, thereby accelerating the convergence speed and recognition accuracy of the downstream text recognition task, and has high application value.
[0115] Embodiment 2
[0116] The difference between the self-supervised text recognition method based on the character movement task in this embodiment and that in Embodiment 1 lies in the different data acquisition in the pre-trained model in step (3-2). Other steps are the same as those in Embodiment 1.
[0117] The pre-training data acquisition in this embodiment is that for an original image I, four different data augmentations are continuously performed to obtain four images I k , and four different data augmentations and character movements are continuously performed to obtain four images MI. Then in one batch, every time the four images I k from the same original image are arranged adjacent to each other and input into the momentum encoding mapping module K for feature extraction; every time the four images MI from the same original image are arranged adjacent to each other and input into the encoding mapping module Q for feature extraction. Therefore, the batch size in this embodiment is four times that in Embodiment 1.
[0118] Embodiment 3
[0119] The difference between the self-supervised text recognition method based on the character movement task in this embodiment and that in Embodiment 1 lies in the different data acquisition and the input of the contrast learning loss function in the pre-trained model in step (3-2). Other steps are the same as those in Embodiment 1.
[0120] The pre-training data acquisition in this embodiment is that for the original image I, two data augmentations are respectively performed to obtain the first image I k and the second image I q , and then the second image I q is subjected to character movement to obtain the image MI.
[0121] Then the first image I k and the second image I q respectively pass through the momentum encoding mapping module K and the encoding mapping module Q, and then the loss value is calculated. The formula of the loss function is
[0122]
[0123] Where C is the length of the negative sample, τ is a hyperparameter; q is the feature vector of the image I q after passing through the encoding and mapping module Q; k + is the feature vector after passing through the momentum encoding and mapping module K, which is the positive sample of q, that is, it comes from the same original image as q; k i (i = 1...C) is the feature vector after passing through the momentum encoding and mapping module K, which is the negative sample of q, that is, it does not come from the same original image as q.
[0124] Meanwhile, the image MI passes through the encoder E and the multi-layer perceptron in the encoding and mapping module Q. The structures of the encoder E and the multi-layer perceptron are the same as those in Embodiment 1. Then classification is performed to predict the pixel value of the character movement in the image. The classification formula is as shown in formula (5) of Embodiment 1.
[0125] The implementation manners of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement manners and are all included in the protection scope of the present invention.
Claims
1. A self-supervised text recognition method based on a character movement task, characterized in that, it includes the following steps: (1) Obtain an image of a handwritten word through an electronic device; (2) Perform data preprocessing on the word image; (3) Perform character localization, character selection, and character movement on the word image, and then determine the label of the character movement; The character localization in step (3) includes the following steps: For the second image , a vertical projection distribution is obtained through vertical projection ; First, perform adaptive binarization on the second image , and then normalize it to . At this time, the value of the area where the text is located is 1, and then perform row summation to obtain the vertical projection distribution ; For the vertical projection distribution Set to zero the numbers less than the value in which Take the second smallest projection value in, and then obtain the character block region set from the vertical projection distribution where is defined as the character block region, i.e., a continuous region with non-zero projection values; represents the number of character block regions; The character selection in step (3) includes the following steps: (321) Randomly select two positions from the character block area set and respectively as the position where the character is located before movement and the target position of the character movement. Regarding and and The selection of is divided into the following three cases: If , it indicates that there is no character block area. Let be the minimum projection value in ; return to step (312) to continue obtaining the character block area set If , it means there is only one character block area, that is . At this time, select a position from the first 40% position and the last 40% position respectively, and then randomly use these two positions as and ; If , it indicates that there are two or more character block regions. At this time, randomly select two character block regions from and respectively as the initial character block region where the character is located and the target character block region to be moved; then randomly select a position from as , and randomly select a position from as ; (322) Determine the width of the character to be moved, and finally select the character image to be moved; the initial half-width of the character image to be moved is set to: Among them is the width of the second image Set the target position for character movement The minimum distance from the image boundary is The position where the character is located before movement The minimum distance from the image boundary is The half-width of the character image to be moved is: (1) The selected character image to be moved is: wherein is the height of the second image , and is the half-width of the character image to be moved; (4) Construct a neural network for pre-training, which is to perform contrastive learning on different data-augmented images of the same original image and classify the images after character movement; (5) Read the encoder parameters of the neural network pre-trained in step (4) into the encoder of the text recognition model, and then use the text recognition model to train the handwritten word image and label.
2. A self-supervised text recognition method based on a character movement task according to claim 1, characterized in that, a grayscale image of a handwritten word is obtained using an electronic device capable of handwritten input.
3. A self-supervised text recognition method based on a character movement task according to claim 1, characterized in that, The preprocessing in step (2) is performed on the original image for data augmentation , including affine transformation, stroke jitter, stroke coverage, and stroke thickness variation, where the relevant parameters of each augmentation method are randomly selected within the set range each time; randomly select a set of parameters within the set range to obtain the first image ; randomly select a set of parameters within the set range to obtain the second image ; then resize the first image and the second image to , where is the image height and is the image width; then normalize the first image to .
4. A self-supervised text recognition method based on a character movement task according to claim 1, characterized in that, The character movement in step (3) includes the following steps: The original picture of the target position of the character movement is: The character image to be moved is superimposed on the second image at a ratio of . Other positions of the second image remain unchanged, and then the moved image is obtained, that is (2) wherein represents the superposition ratio .
5. A self-supervised text recognition method based on a character movement task according to claim 4, characterized in that, Determining the label of the character movement is specifically: The pixel value of character movement , when , it means the character moves to the left; when , it means the character moves to the right; define the character movement task as a classification task, and let the classification label , where is the width of the second image .
6. A self-supervised text recognition method based on a character movement task according to claim 1, characterized in that, The neural network includes an encoding and mapping module , a momentum encoding and mapping module and a multi-layer perceptron; The encoding mapping module includes an encoder and a mapper, and the encoding mapping module is trained according to the stochastic gradient descent optimizer; Input the output features of the encoder in the encoding mapping module into a multi-layer perceptron, and then classify the output feature vectors to predict the pixel values of character movement in the image; The momentum encoding mapping module has the same network structure as the encoding mapping module and uses the parameters of the encoding mapping module for momentum update; Let the parameters of the encoder E and the mapper in the encoding mapping module be , and the parameters of the encoder and the mapper in the momentum encoding mapping module be . The formula for momentum update is: (3) Among them represents the magnitude of momentum, .
7. A self-supervised text recognition method based on a character movement task according to claim 6, characterized in that, The pre-training of the neural network includes: the first image obtained after data augmentation and the moving image obtained after data augmentation and character movement respectively pass through the momentum encoding mapping module and the encoding mapping module , and then calculate the loss value. The formula of the loss function is: (4) Among them, is the length of the negative sample; is a hyperparameter; is the feature vector after passing through the encoding mapping module ; is the feature vector after passing through the momentum encoding mapping module , and it is the positive sample of , coming from the same original image as ; is the feature vector after passing through the momentum encoding mapping module , and it is the negative sample of , that is, it does not come from the same original image as , ; For negative samples, a size of the negative sample is preset, and then the feature vectors after each pass through the momentum encoding mapping module are stored. After reaching the preset negative sample quantity, the feature vectors stored first are deleted, and then new feature vectors are stored; For the augmented data and the shifted image obtained after character shifting , in addition to using the output vector of its encoding mapping module to participate in the calculation of formula (4), the output features of the encoder in the encoding mapping module will also be input into a multi-layer perceptron, and then the feature vectors output by the multi-layer perceptron will be classified to predict the pixel values of character shifting in the image. The classification formula is: (5) where is the batch size; is the moving image corresponding one-hot vector of the character movement label; is the probability vector predicted by the multi-layer perceptron, and the calculation formula is: (6) Among them is the th moving image in a batch after passing through the encoder of the coding mapping module and the output feature vector of the multi-layer perceptron; is the th moving image in a batch ; is the th moving image in a batch ; The final overall loss function is where is a hyperparameter.
8. A self-supervised text recognition method based on a character movement task according to any one of claims 1 to 7, characterized in that, The text recognition model adopts an encoder-decoder structure, and the structure of the encoder of the text recognition model is the same as that of the encoder E of the coding mapping module ; During the training process, it is necessary to first read the encoder parameters of the neural network pre-trained in step (4) into the encoder of the text recognition model, and the parameters of the decoder are randomly initialized, and then fine-tune and train the entire text recognition model according to the input handwritten word image and the corresponding label.
Citation Information
Patent Citations
Text input on an interactive display
CN106233240A
Method for detecting and identifying continuous segmented texts in image
CN110399845A