Feature reconstruction and consistency CTC-based semantic enhancement text recognition method
Through the semantic enhancement method of feature reconstruction and consistent CTC, the alignment problem of traditional CTC models in text recognition in different font styles is solved, the accuracy and robustness of handwritten text recognition is improved, and the recognition performance of the model in complex scenarios is enhanced.
Patent Information
- Application Number
- CN202510371688.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The traditional connection attention-based time classification (CTC) model has alignment problems when dealing with texts of different font styles and lacks semantic supervision, resulting in limited recognition performance in complex scenarios.
The semantic enhanced text recognition method based on feature reconstruction and consistency CTC is adopted, and different views are generated through RandAugment data augmentation technology, and consistency regularization loss is introduced. The feature reconstruction and semantic enhancement are carried out, and a multi-layer perceptron and self-attention mechanism is used for deep processing, and a semantic supervision loss and consistency regularization loss optimization model is introduced.
It improves the accuracy and robustness of handwritten text recognition, effectively solves the alignment problem, and enhances the recognition performance of the model in complex scenarios.
Smart Images

Figure CN120299053A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and deep learning, and particularly relates to a semantic-enhanced text recognition method for handwritten text recognition tasks based on feature reconstruction and consistent CTC. Background Art
[0002] In handwritten text recognition tasks, traditional models based on connectionist temporal classification (CTC) have alignment problems, especially when dealing with texts of different font styles. In addition, existing models lack semantic supervision, resulting in limited recognition performance in complex scenarios. Summary of the Invention
[0003] The main objective of the present invention is to propose a semantic-enhanced text recognition method for handwritten text recognition tasks based on feature reconstruction and consistent CTC, aiming to improve the accuracy and robustness of text recognition through feature reconstruction and semantic enhancement technologies, and effectively solve the alignment problems existing in the prior art.
[0004] To achieve the above objective, the present invention provides a semantic-enhanced text recognition method based on feature reconstruction and consistent CTC, and the method includes the following steps:
[0005] Step S00, pre-establish a CTC model;
[0006] Step S10, obtain a text image Image, and preprocess the text image based on the height and aspect ratio of the text image and a preset maximum aspect ratio;
[0007] Step S20, generate two different enhanced views Image1 and Image2 of the preprocessed text image through the RandAugment data augmentation technique;
[0008] Step S30, input the two different enhanced views Image1 and Image2 into a pre-trained CTC model for processing;
[0009] Step S40, output the processing result as the text recognition result;
[0010] The step S30, the step of inputting the two different enhanced views Image1 and Image2 into a pre-trained CTC model for processing includes:
[0011] Process the two different enhanced views Image1 and Image2 through a shared visual encoder f to obtain corresponding feature distributions;
[0012] x1 = f(Image1) (1-1)
[0013] x2 = f(Image2) (1-2);
[0014] Introduce the consistency regularization loss LCR to calculate the bidirectional Kullback-Leibler divergence of the two feature distributions, and use it as a regularization term:
[0015]
[0016] where sg represents the stop-gradient operation to ensure that the consistency regularization does not backpropagate to the target distribution.
[0017] A further technical solution of the present invention is that the step S00 includes:
[0018] Step S001, preprocess the text image based on the height and aspect ratio of the original text image, and a preset maximum aspect ratio:
[0019] Step S002, combine the convolutional neural network and the Transformer architecture, extract the local features of the preprocessed image through convolutional operations, imitate the structure of the Transformer block to enhance the feature interaction ability, and perform deep processing on the input features through the self-attention mechanism and the multi-layer perceptron, so as to effectively capture the global information and obtain two-dimensional features:
[0020] Step S003, reconstruct the two-dimensional features and convert them into a feature sequence that conforms to the reading order of the text image;
[0021] Step S004, input the feature sequence that conforms to the reading order of the text image into the classifier, calculate the predicted character sequence, and align the predicted character sequence with the label sequence according to the CTC rule.
[0022] A further technical solution of the present invention is that the step S001 includes:
[0023] Keep the height H of the original text image fixed, calculate the aspect ratio r of the original text image, and its calculation formula is:
[0024]
[0025] where W orig and H orig respectively represent the width and height of the original text image;
[0026] Round the aspect ratio r of the original text image to the closest integer r round , and its calculation formula is:
[0027] r round = round(r) (1-5);
[0028] According to the rounded aspect ratio r round , calculate the new width W of the image new , and its calculation formula is:
[0029] W new = H × r round (1-6);
[0030] The step S001 further includes setting the maximum aspect ratio r max , if the calculated aspect ratio r round exceeds the maximum aspect ratio r max , then limit the width to the corresponding maximum width W max = H × r max , and the final image width takes the smaller value of W new and W max , and its calculation formula is:
[0031] W new = min(H × r round , H × r max ) (1-7);
[0032] The finally obtained image size is (W new , H), and the calculation formula is:
[0033] (W new , H) = (min(H × r round , H × r max ), H) (1-8).
[0034] A further technical solution of the present invention is that the step S002 includes:
[0035] Use PatchEmbedding to convert the input image into several local regions, and represent the global features of the image through the embedding vectors of these regions, specifically:
[0036] Map the number of channels of the input image to a set number of channels embed_dim through two convolutional layers. Each convolutional layer includes a convolutional operation, batch normalization, and a non-linear activation function. The size of the convolutional kernel is 3×3, the stride is 2, the padding is 1, and the GELU activation function is used. Among them, the calculation formula is:
[0037] x1 = ConvBNLayer(x) (1-9)
[0038] x2 = ConvBNLayer(x1) (1-10);
[0039] Among them, x1 and x2 are the outputs after two convolutional operations respectively, and x is the image input;
[0040] For an image with an input size of H×W, after PatchEmbedding, the output feature map size is Each spatial position corresponds to a feature vector of embed_dim dimension, and the finally output feature representation is
[0041] A further technical solution of the present invention is that after the step S002, the following is further included:
[0042] After the convolution operation, a multi-layer perceptron is used to perform non-linear transformation on the extracted features to enhance the learning ability of global information. Among them, layer normalization and DropPath regularization mechanisms are introduced in each stage of the non-linear transformation.
[0043] A further technical solution of the present invention is that the step S003 includes:
[0044] Convert the two-dimensional feature into a feature sequence that conforms to the reading order of the text image That is, map the relevant features from to where the value ranges of j and m are The feature mapping process can be formally expressed with a matrix That is, through matrix operations obtain the reconstructed feature sequence
[0045] Among them, is a tensor in the real number field, and its shape Among them, F is the feature extracted, H is the height of the original image input, W is the width of the original image input, and D2 is the number of channels that can be set.
[0046] A further technical solution of the present invention is that the learning matrix M is divided into horizontal direction reconstruction and vertical direction mapping:
[0047] In the horizontal direction reconstruction, for each row of the feature map perform unfolding processing to learn a horizontal rearrangement matrix The matrix element reflects the feature after horizontal direction reconstruction corresponding to the probability of the original feature F i,m ; among them, a variant of the self-attention mechanism is used to calculate the horizontal rearrangement matrix Specifically, perform linear transformation on the feature row F i to obtain query, key, and value vectors. Let Among them is a learnable weight matrix, the horizontal reconstruction matrix The calculation process is as follows:
[0048] First, calculate the attention scores:
[0049]
[0050] To make the scores more stable, perform a scaling operation on them, and then convert them into a probability distribution through the Softmax function, that is
[0051]
[0052] Based on the learned Reconstruct the features of each row horizontally. The reconstructed features are processed through residual connections and a multi-layer perceptron. The specific calculation is as follows:
[0053]
[0054] Among them, LayerNorm represents the layer normalization operation, and MLP is a two-layer feed-forward neural network composed of two linear layers and an activation function;
[0055] In the vertical mapping step, introduce a learnable global context vector Let this global context vector interact with the features of each column in F h to learn the vertical reconstruction matrix
[0056] Similarly, use the attention mechanism to calculate the vertical reconstruction matrix First, use the global context vector T as the query vector, and perform a linear transformation on the column features to obtain the key vector Among them is a learnable weight matrix;
[0057] The attention scores are calculated as After the scaling operation obtain the vertical reconstruction matrix through the Softmax function
[0058]
[0059] Based on the vertical reconstruction matrix obtain the vertically reconstructed features
[0060]
[0061] All column features share the same global context vector T;
[0062] Finally, all the features after vertical reconstruction are combined to obtain Denoted as the feature sequence after rearrangement
[0063] A further technical solution of the present invention is that the step S003 further includes:
[0064] Rewrite the mapping relationship between the feature sequence and the original feature F:
[0065]
[0066] where M j is a matrix that combines horizontal and vertical reconstruction information, and multiplying it with the original feature F gives the feature after vertical reconstruction
[0067] Combine all the M j to obtain Verify again
[0068] A further technical solution of the present invention is that the step S00 further includes:
[0069] Introduce context information, and combine the multi-head attention mechanism and context embedding to optimize the semantic understanding ability of the model, specifically including:
[0070] For each character c L with a character label Y = {c1, c2,..., c i} in the text image, its context is composed of the left string and the right string where l s is the context window length;
[0071] First, map the characters in to string embeddings Convert the characters into numerical features for subsequent model processing; then, calculate the context representation of the left string through the multi-head attention mechanism Assume the number of heads is h, and perform linear transformations on and the predefined token respectively to obtain queries, keys, and values for multiple heads:
[0072]
[0073] where is a learnable weight matrix. The attention scores and outputs for the k-th head are as follows:
[0074]
[0075]
[0076] The final output of the multi-head attention is:
[0077]
[0078] where is a learnable weight matrix that normalizes the multi-head attention output to obtain the context representation:
[0079]
[0080] Next, calculate the attention map through the multi-head attention mechanism Assume the number of heads is h. Linearly transform and the visual feature F respectively to obtain the queries, keys, and values for multiple heads:
[0081]
[0082] where is a learnable weight matrix. The attention scores and outputs for the k-th head are as follows:
[0083]
[0084] The final output of the multi-head attention is:
[0085]
[0086] where is a learnable weight matrix. Normalize the multi-head attention output to obtain the attention map:
[0087]
[0088] Use the attention map to weight the visual feature F to obtain the feature corresponding to the character c i :
[0089]
[0090] Input into a fully connected layer for classification prediction:
[0091]
[0092] where is the weight matrix is the bias vector, N c is the character set size; finally, use the prediction result and the true label c i to calculate the cross-entropy loss to train the model;
[0093] The right-side string is processed symmetrically to the left-side string, and finally an attention map visual features and the prediction result are involved in the loss calculation.
[0094] A further technical solution of the present invention is that the step S003 further includes:
[0095] constraining and optimizing the model based on the connection attention time classification loss, semantic supervision loss, and consistency regularization loss, where
[0096] The calculation formula of the connection attention time classification loss is:
[0097]
[0098] where is the prediction of the i-th character by the model, F v is the feature sequence rearranged by the feature reconstruction module, Y is the true label sequence, represents the probability of predicting the i-th character under the condition of the given rearranged feature F v and the true label Y;
[0099] The semantic supervision loss is measured by calculating the cross-entropy between the character category probability based on the context prediction information and the true label, and the specific calculation formula is:
[0100]
[0101] where and are the predicted category probabilities obtained by using the left-side and right-side string information of the target character respectively, c i is the true label of the i-th character, and ce represents the cross-entropy loss function;
[0102] The consistency regularization loss is calculated by calculating the consistency between the CTC predictions of the two branches, and its calculation formula is:
[0103] L cr = L CR (x1, x2) (1 - 38);
[0104] where represents the Kullback-Leibler divergence, where predicts1 and predicts2 are the CTC prediction distributions of the two branches respectively;
[0105] By integrating the CTC loss, semantic supervision loss, and consistency regularization loss, the total loss function of the text recognition algorithm based on feature reconstruction and semantic enhancement is obtained:
[0106] L total = αL ctc + βL gtc + γL cr (1 - 39);
[0107] It is added to the model training process, and the gradient descent algorithm is used for iteration until the maximum number of iterations is reached or the model converges.
[0108] Through the above technical solutions, the present invention performs sequence learning on image information in order to better fuse image information with speech and text, establishes a temporal order model for extracting semantic information, and can improve the accuracy and robustness of text recognition through feature reconstruction and semantic enhancement technologies, effectively solving the alignment problem existing in the prior art. Brief Description of the Drawings
[0109] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on the structures shown in these drawings without creative efforts.
[0110] Figure 1 is a schematic flowchart of a preferred embodiment of the text recognition method based on feature reconstruction and semantic enhancement of the present invention;
[0111] Figure 2 is the overall framework diagram of the text recognition method based on feature reconstruction and semantic enhancement of the present invention;
[0112] Figure 3 is a schematic diagram of the feature extraction module;
[0113] Figure 4 is a schematic diagram of the global feature extraction module.
[0114] The realization, functional characteristics, and advantages of the object of the present invention will be further described in combination with the embodiments with reference to the drawings. Detailed Embodiments
[0115] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0116] To effectively solve the alignment problem existing in the prior art and improve the text recognition accuracy through semantic supervision, the present invention proposes a semantic-enhanced text recognition method based on feature reconstruction and consistent CTC. The implementation of the present invention involves the following key modules: a multi-scale adaptive adjustment module, a local feature and global feature extraction module, a feature reconstruction module, a consistent regularization CTC module, and a semantic supervision module.
[0117] Please refer to Figure 1 and Figure 2 , a preferred embodiment of the text recognition method based on feature reconstruction and semantic enhancement of the present invention includes the following steps:
[0118] Step S00, a CTC model is established in advance;
[0119] Step S10, obtain a text image Image, and preprocess the text image based on the height and aspect ratio of the text image and a preset maximum aspect ratio;
[0120] Step S20, generate two different enhanced views Image1 and Image2 of the preprocessed text image through the RandAugment data augmentation technique;
[0121] Step S30, input the two different enhanced views Image1 and Image2 into the pre-trained CTC model for processing;
[0122] Step S40, output the processing result as the text recognition result;
[0123] The step S30, the step of inputting the two different enhanced views Image1 and Image2 into the pre-trained CTC model for processing includes:
[0124] Process the two different enhanced views Image1 and Image2 through a shared visual encoder f to obtain corresponding feature distributions;
[0125] x1 = f(Image1) (1-1)
[0126] x2 = f(Image2) (1-2);
[0127] The consistency regularization loss \(L_{CR}\) is introduced to calculate the bidirectional Kullback-Leibler divergence of two feature distributions, which is used as a regularization term:
[0128]
[0129] where \(sg\) represents the stop-gradient operation to ensure that the consistency regularization does not backpropagate to the target distribution.
[0130] In this embodiment, the execution subject of step S30 is the consistency regularization CTC module. For a given input image Image1, in this embodiment, two different augmented views Image1 and Image2 are generated through the RandAugment data augmentation technique. These two views are respectively processed by a shared visual encoder \(f\) to obtain corresponding feature distributions \(x1\) and \(x2\).
[0131] To ensure the consistency of the CTC distributions of these two augmented views, this embodiment introduces the consistency regularization loss \(L_{CR}\). Specifically, the present invention calculates the bidirectional Kullback-Leibler divergence between the two distributions and uses it as a regularization term.
[0132] Furthermore, in this embodiment, step S00 includes:
[0133] Step S001, preprocess the text image based on the height and aspect ratio of the original text image and a preset maximum aspect ratio;
[0134] Step S002, combining a convolutional neural network and a Transformer architecture, extract local features of the preprocessed image through convolutional operations, imitate the structure of the Transformer block to enhance the feature interaction ability, and perform in-depth processing on the input features through the self-attention mechanism and multi-layer perceptron, so as to effectively capture global information and obtain two-dimensional features;
[0135] Step S003, reconstruct the two-dimensional features into a feature sequence that conforms to the reading order of the text image;
[0136] Step S004, input the feature sequence that conforms to the reading order of the text image into a classifier, calculate the predicted character sequence, and align the predicted character sequence with the label sequence according to the CTC rule.
[0137] Step S001 includes:
[0138] Keep the height \(H\) of the original text image fixed, calculate the aspect ratio \(r\) of the original text image, and its calculation formula is:
[0139]
[0140] Among them, W orig and H orig respectively represent the width and height of the original text image;
[0141] Round the aspect ratio r of the original text image to the nearest integer r round , and its calculation formula is:
[0142] r round = round(r) (1-5);
[0143] According to the rounded aspect ratio r round , calculate the new width W new of the image, and its calculation formula is:
[0144] W new = H × r round (1-6);
[0145] The step S001 further includes setting a maximum aspect ratio r max . If the calculated aspect ratio r round exceeds the maximum aspect ratio r max , then limit the width to the corresponding maximum width W max = H × r max . The final image width takes the smaller value of W new and W max , and its calculation formula is:
[0146] W new = min(H × r round , H × r max ) (1-7);
[0147] The finally obtained image size is (W new , H), and the calculation formula is:
[0148] (W new , H) = (min(H × r round , H × r max ), H) (1-8).
[0149] During the image preprocessing process, the traditional unified size adjustment method may destroy the aspect ratio of the image, resulting in the loss of key information. Especially in tasks such as scene text recognition, the spatial structure of the image is very important. To solve this problem, this embodiment adopts a multi-scale adaptive adjustment module and proposes a preprocessing strategy based on the aspect ratio of the image, in which the height of the image remains fixed, the width is adjusted according to the aspect ratio of the original image, and the maximum aspect ratio is set to avoid generating an overly large image size.
[0150] Assume that the height H of the image is fixed at 32, and the width W is adjusted according to the original aspect ratio of the image. For each input image, first calculate its original aspect ratio r, that is, the ratio of the width to the height of the image. The specific calculation steps are shown in formula (1-4). Then, round the calculated aspect ratio r to the nearest integer r round , as shown in formula (1-5). According to the rounded aspect ratio r round , the new width W of the image can be calculated new , and its calculation formula is (1-6):
[0151] For example, H = 32 is the fixed height, and W new is the adjusted width. round() is the rounding function.
[0152] To avoid the image width from being too large, set a maximum aspect ratio r max . If the calculated aspect ratio r round exceeds r max , then limit the width to the corresponding maximum width W max = H × r max . Therefore, the final width W of the image new will be the smaller of the following two values, as shown in formula (1-7). The final obtained image size is (W new , H), and the calculation formula is as (1-8). Among them, r max is the set maximum aspect ratio to ensure that the image width does not exceed this value. For example, if H = 32 and r max = 4, then the maximum width of the image is 32 × 4 = 128. If the aspect ratio r of the original image is greater than 4, the rounded width will still be limited to 128.
[0153] Through this method, the aspect ratio of the image is reasonably controlled, which not only retains the proportional characteristics of the original image but also avoids the problem of too large image size. Setting the maximum aspect ratio r max can prevent the model from generating too large inputs when processing images with extreme aspect ratios, and maintain the stability and efficiency of the training process.
[0154] Furthermore, in this embodiment, the step S002 includes:
[0155] Use PatchEmbedding to convert the input image into several local regions, and represent the global features of the image through the embedding vectors of these regions. Specifically:[[]]
[0156] The number of channels of the input image is mapped to a set number of channels embed_dim through two convolutional layers. Each convolutional layer includes a convolution operation, batch normalization, and a non-linear activation function. The size of the convolutional kernel is 3×3, the stride is 2, and the padding is 1. And the GELU activation function is used, where the calculation formula is:
[0157] x1 = ConvBNLayer(x) (1-9)
[0158] x2 = ConvBNLayer(x1) (1-10);
[0159] where x1 and x2 are the outputs after two layers of convolution operations respectively, and x is the image input;
[0160] For an image with an input size of H×W, after PatchEmbedding, the size of the output feature map is Each spatial position corresponds to a feature vector of embed_dim dimensions, and the finally output feature representation is
[0161] In this embodiment, the local feature and global feature extraction module combines a convolutional neural network (CNN) and a Transformer architecture, extracts local features through convolution operations, and enhances global feature interaction using the self-attention mechanism. The specific implementation includes PatchEmbedding, positional encoding, and a multi-layer perceptron (MLP).
[0162] The image first goes through PatchEmbedding. PatchEmbedding is a technique that converts the input image into a series of fixed-size image patches and maps them to a high-dimensional feature space. This method is a common step in deep learning architectures such as Vision Transformers (ViT). Its core idea is to convert the input image into several local regions through convolution operations, and represent the global features of the image through the embedding vectors of these regions.
[0163] In the specific implementation process, two convolutional layers are used to process the input image, and the specific process is as follows:
[0164] First, the number of channels of the input image, which is 3, is mapped to embed_dim through two convolutional layers. embed_dim is a set number of channels, usually 128. Each convolutional layer generally includes a convolution operation, batch normalization (BatchNormalization), and a non-linear activation function (such as GELU). The size of the convolutional kernel is generally 3×3, and the stride is 2, which reduces the spatial size of the image by half. The formulas are expressed as 1-9 and 1-10, where x1 and x2 are the outputs after two layers of convolution operations respectively.
[0165] x is the image input. ConvBNLayer is a commonly used module that can simplify the network construction process and improve the training efficiency and stability of the network by combining a convolutional layer, a batch normalization layer, and an activation function. In this patent, ConvBNLayer uses a 3×3 convolutional kernel, a stride of 2, a padding of 1, and the GELU activation function.
[0166] For an image with an input size of H×W, after PatchEmbedding, the output feature map size is Each spatial position corresponds to an embed_dim-dimensional feature vector. The final output feature representation is These feature vectors are then input into the local feature extraction module for further processing.
[0167] Furthermore, in this embodiment, after the step S002, it further includes:
[0168] After the convolution operation, a multi-layer perceptron is used to perform non-linear transformation on the extracted features to enhance the global information learning ability. Among them, layer normalization and DropPath regularization mechanisms are introduced at each stage of the non-linear transformation.
[0169] As Figure 3 shown, the local feature extraction module combines the ideas of a convolutional neural network (CNN) and a Transformer architecture, aiming to efficiently extract local features in visual images through convolution operations and imitate the structure of the Transformer block to enhance the feature interaction ability. The core of the local feature extraction module is a convolution mixer composed of multiple convolutional layers, where each layer uses group convolution to enhance the flow of feature information between different local regions. Group convolution reduces the computational amount by grouping channels, and at the same time effectively captures the local spatial relationship in the image, thereby improving the flexibility and expressiveness of feature learning.
[0170] After the convolution operation, the module further performs non-linear transformation on the extracted features through a multi-layer perceptron (MLP), thereby enhancing the global information learning ability. Different from the standard convolutional network, this module introduces regularization mechanisms such as layer normalization (LayerNorm) and DropPath at each stage, which helps to stabilize the training process and reduce overfitting. Specifically, layer normalization is used to standardize the input of each layer, thereby improving the gradient flow; while DropPath further enhances the generalization ability of the model by randomly discarding some paths.
[0171] As Figure 4As shown in the figure, the global feature extraction module draws on the design concept of Transformer. It mainly processes the input features deeply through the self-attention mechanism and the multi-layer perceptron (MLP), so as to effectively capture global information. The module structure consists of two main parts: self-attention and MLP. Among them, self-attention is used to model global dependencies, and MLP is used to further enhance the non-linear expression ability of features.
[0172] As Figure 4 shown, this global feature extraction module processes the mutual dependencies of global features through the self-attention mechanism and enhances the non-linear ability of feature representation using the multi-layer perceptron. Through the application of layer normalization and DropPath, not only the stability of the model is improved, but also overfitting is prevented. The entire structure is similar to the basic block of Transformer and can perform deep interaction and transformation on the input features globally, which is beneficial to capturing global context information.
[0173] After the above local feature and global feature extraction modules extract features, a two-dimensional feature F ∈ Next, the feature reconstruction module will be used to reconstruct the features.
[0174] Furthermore, in this embodiment, the step S003 includes:
[0175] Convert the two-dimensional feature into a feature sequence that conforms to the reading order of the text image That is, map the relevant features from to where the value ranges of j and m are The feature mapping process can be formally expressed with a matrix That is, through matrix operations to obtain the reconstructed feature sequence
[0176] where, is a tensor over the real number field, and its shape where F is the feature extracted, H is the height of the original image input, W is the width of the original image input, and D2 is the number of channels that can be set.
[0177] In the handwritten character recognition task, different font styles will pose alignment problems for models based on connectionist temporal classification (CTC). Therefore, this embodiment realizes feature dimensionality reduction through the feature reconstruction module. Different from directly using some simple operations, this feature reconstruction module uses the Attention mechanism and retains more text features.
[0178] The main objective of the feature reconstruction module is to transform two-dimensional features into a feature sequence that conforms to the reading order of the text image From a conceptual level, this transformation process can be regarded as a mapping operation.
[0179] Among them, the learning matrix M is divided into two steps: horizontal direction reconstruction and vertical direction mapping.
[0180] The first step is horizontal direction reconstruction. In horizontal direction reconstruction, each row of the feature map is unfolded and processed to learn a horizontal rearrangement matrix The matrix elements reflect the features after horizontal direction reconstruction and the probability corresponding to the original feature F i,m
[0181] In this embodiment, a variant of the self-attention mechanism is adopted to calculate the horizontal rearrangement matrix Specifically, a linear transformation is performed on the feature row F i to obtain query, key, and value vectors. Let where is a learnable weight matrix. The calculation process of the horizontal direction reconstruction matrix is as follows:
[0182] First, calculate the attention scores:
[0183]
[0184] To make the scores more stable, a scaling operation is performed on them, and then they are converted into a probability distribution through the Softmax function, that is
[0185]
[0186] Based on the learned In this embodiment, the features of each row are reconstructed in the horizontal direction. The reconstructed features are processed through residual connection and a multi-layer perceptron (MLP). The specific calculation is as follows:
[0187]
[0188] Among them, LayerNorm represents the layer normalization operation, which helps to stabilize the training process of the model.
[0189] MLP is a two-layer feedforward neural network composed of two linear layers and an activation function. After the above processing, we obtain the horizontally rearranged features The horizontally rearranged features are denoted as Fh , ∈ is a mathematical symbol, belonging to, This represents the matrix size.
[0190] After completing the horizontal reconstruction, vertical mapping is performed. During the vertical mapping step, a learnable global context vector is introduced It can be regarded as a prior representation of the entire text context. Let this global context vector interact with the h features of each column in F to learn the vertical reconstruction matrix
[0191] The attention mechanism is also used to calculate the vertical reconstruction matrix First, take the global context vector T as the query vector, and perform a linear transformation on the column features to obtain the key vector where is a learnable weight matrix;
[0192] The attention score is calculated as After the scaling operation and then, the vertical reconstruction matrix is obtained through the Softmax function
[0193]
[0194] Based on the vertical reconstruction matrix we obtain the features after vertical reconstruction
[0195]
[0196] All column features share the same global context vector T; this design enables the model to better capture the overall context information of the text, enhancing the model's generalization ability for long text sequences. Even when the number of column features encountered during testing exceeds the number during training, it can effectively recognize long texts.
[0197] Finally, combine all the features after vertical reconstruction to obtain denoted as the rearranged feature sequence
[0198] To more clearly understand the mapping relationship of the features in the feature mapping module, in this embodiment, the mapping relationship between the feature sequence and the original feature F is rewritten as:
[0199]
[0200] where, M jis a matrix that combines horizontal and vertical reconstruction information and multiplies it by the original feature F to obtain the vertically reconstructed feature
[0201] Combine all the M j to obtain Verify again
[0202] Different from existing image feature reconstruction models, the feature mapping module operates in the feature space. Through horizontal and vertical reconstruction, it maps the visual features in the original feature F that are irregularly organized but contain complete text information to be recognized into a more regular feature sequence This structured rearrangement method enables the feature reconstruction module to better extract text features and obtain features that are more matched to CTC classification, thus effectively solving the CTC alignment problem caused by feature mapping.
[0203] The feature sequence after feature mapping is input into the classifier, and through the formula to obtain the predicted character sequence where are the learnable weights of the classifier, and N c is the size of the character set. The predicted character sequence is then aligned with the label sequence Y according to the CTC rule.
[0204] Furthermore, in this embodiment, the step S00 further includes:
[0205] Introduce context information and combine the multi-head attention mechanism and context embedding to optimize the semantic understanding ability of the model.
[0206] In this embodiment, this step is mainly implemented through the semantic supervision module.
[0207] The semantic supervision module optimizes the semantic understanding ability of the model by introducing context information, and the specific implementation includes the multi-head attention mechanism and context embedding.
[0208] In handwritten text recognition, models based on Connectionist Temporal Classification (CTC) mainly rely on direct classification of visual features to obtain recognition results. However, this method lacks semantic supervision, while the semantic supervision module optimizes the performance by introducing context information.
[0209] The steps of introducing context information and combining the multi-head attention mechanism and context embedding to optimize the semantic understanding ability of the model specifically include:
[0210] For each character c with the character label Y = {c1, c2,..., c L} in the text image i, whose context consists of the left string and the right string where l s is the context window length. The core purpose of the semantic supervision module is to incorporate the context information of the left and right strings into the visual features.
[0211] Taking the left string as an example (the right side is processed symmetrically), first map the characters in to string embeddings Convert the characters into numerical features for subsequent model processing. Then, calculate the context representation of the left string through the multi - head attention mechanism Assume the number of heads is h, and linearly transform and the predefined token respectively to obtain queries, keys, and values for multiple heads:
[0212]
[0213] where is the learnable weight matrix, and the attention scores and outputs for the k - th head are respectively:
[0214]
[0215] The final output of the multi - head attention is:
[0216]
[0217] where is the learnable weight matrix, which normalizes the multi - head attention output to obtain the context representation:
[0218]
[0219] Next, calculate the attention map through the multi - head attention mechanism Assume the number of heads is h, and linearly transform and the visual feature F respectively to obtain queries, keys, and values for multiple heads:
[0220]
[0221] V k = FW v,k (1 - 29);
[0222] where is the learnable weight matrix, and the attention scores and outputs for the k - th head are respectively:
[0223]
[0224] The final output of the multi-head attention is:
[0225]
[0226] where is a learnable weight matrix. The output of the multi-head attention is normalized to obtain the attention map:
[0227]
[0228] Using the attention map weight the visual feature F to obtain the feature corresponding to the character c i :
[0229]
[0230] Input into a fully connected layer for classification prediction:
[0231]
[0232] where is the weight matrix, is the bias vector, and N c is the size of the character set; finally, use the prediction result and the true label c i to calculate the cross-entropy loss to train the model.
[0233] The processing of the right-side string is symmetric to that of the left-side string, and finally the attention map the visual feature and the prediction result are obtained and participate in the loss calculation. During training, the semantic supervision module can effectively guide the visual model to integrate the language context information into the visual features. Even when the semantic supervision module is not used during inference, this context information will be retained in the visual features, thereby improving the recognition accuracy of the CTC model.
[0234] Furthermore, in this embodiment, the step S003 further includes:
[0235] Constraining and optimizing the model based on the connection attention time classification loss, the semantic supervision loss, and the consistency regularization loss.
[0236] In the handwritten text recognition task, to further improve the performance of the text recognition algorithm based on feature reconstruction and semantic enhancement, we designed an optimization objective with a two-branch structure. This objective consists of three parts: the Connectionist Temporal Classification (CTC) loss, the Semantic Supervision (GTC) loss, and the Consistency Regularization Loss (CR-Loss). These three parts of the loss constrain and optimize the model from different perspectives, thereby improving the accuracy and robustness of the model in text recognition.
[0237] Among them, the Connectionist Temporal Classification (CTC) loss is the basis of the optimization objective, mainly used to solve the alignment problem between the character sequence predicted by the model and the true label sequence. In handwritten text recognition, the correspondence between the input visual features and the output character sequence is complex and uncertain, and the CTC loss provides an effective solution for this.
[0238] The calculation formula of the Connectionist Attention Temporal Classification loss is as follows:
[0239]
[0240] Where, is the prediction of the i-th character by the model, F v is the feature sequence rearranged by the feature reconstruction module, Y is the true label sequence, represents the probability of predicting the i-th character under the condition of the given rearranged feature F v and the true label Y.
[0241] The introduction of the semantic supervision loss is to enable the model to better utilize the language context information. The semantic supervision loss is measured by calculating the cross-entropy between the character category probability based on the context prediction information and the true label. The specific calculation formula is as follows:
[0242]
[0243] Where, and are the predicted category probabilities obtained by using the string information on the left and right sides of the target character respectively, c i is the true label of the i-th character, and ce represents the cross-entropy loss function. By minimizing L gtc , the model can pay more attention to the context information. Especially when dealing with easily confused characters, the context information can help the model make more accurate judgments and improve the recognition performance of the model in complex scenarios.
[0244] The introduction of the Consistency Regularization Loss (CR-Loss) is to further improve the generalization ability and robustness of the model. By calculating the consistency between the CTC predictions of two branches, CR-Loss can ensure the consistency of the model on different branches, thus reducing overfitting.
[0245] The consistency regularization loss calculates the consistency between the CTC predictions of two branches, and its calculation formula is:
[0246] L cr = L CR (x1, x2) (1-38);
[0247] where represents the Kullback-Leibler divergence, and predicts1 and predicts2 are the CTC prediction distributions of the two branches respectively. By minimizing L cr , the model can learn a smoother and more consistent prediction distribution, thereby improving the performance on unseen data.
[0248] Combining the CTC loss, semantic supervision loss, and consistency regularization loss, the total loss function of the text recognition algorithm based on feature reconstruction and semantic enhancement is obtained:
[0249] L total = αL ctc + βL gtc + γL cr (1-39);
[0250] Add it to the model training process and use the gradient descent algorithm for iteration until the maximum number of iterations is reached or the model converges.
[0251] The implementation process of the semantic enhancement text recognition method based on feature reconstruction and consistency CTC of the present invention is as follows:
[0252] 1. In the step of data preprocessing, the text line data is used as the input.
[0253] 2. Add L total to the model training process and use the gradient descent algorithm for iteration until the maximum number of iterations is reached or the model converges. The overall technical framework diagram is as Figure 1 shown.
[0254] The experimental results (Accuracy) on the multi-modal public dataset CASIA-HWDB are shown in Table 1.
[0255] Table 1
[0256] Acc(%) NED(%) RCNN[2] 90.4 98.4 SVTR[3] 93.5 99.1 Ours 96.7 99.3
[0257] Accuracy (Acc) is a general metric for measuring the proportion of correct predictions by a model and is applicable to various classification tasks. Its calculation formula is:
[0258]
[0259] where: N correct represents the number of correctly predicted samples, and N total represents the total number of samples.
[0260] Normalized Edit Distance (Norm Edit Dis, NED) is a metric used to measure the similarity between the recognition result and the ground truth annotation, especially suitable for text recognition tasks. Its calculation formula is:
[0261]
[0262] where Edit Distance represents the minimum number of operations (insertion, deletion, substitution) required to convert the recognition result into the ground truth annotation. Length of Ground Truth represents the length of the ground truth string.
[0263] In summary, through the above technical solutions, the present invention conducts sequential learning on image information in order to better integrate image information with speech and text, establishes a time sequence model for the extraction of semantic information, and can improve the accuracy and robustness of text recognition and effectively solve the alignment problem existing in the prior art through feature reconstruction and semantic enhancement technologies.
[0264] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural transformation made under the concept of the present invention by using the content of the specification and drawings of the present invention, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present invention.
Claims
1. A semantic enhancement text recognition method based on feature reconstruction and consistent CTC, characterized in that The method includes the following steps: Step S00, pre-establish a CTC model; Step S10, obtain a text image Image, and preprocess the text image based on the height and aspect ratio of the text image and a preset maximum aspect ratio: Step S20, generate two different augmented views Image1 and Image2 of the preprocessed text image through the RandAugment data augmentation technique; Step S30, input the two different augmented views Image1 and Image2 into a pre-trained CTC model for processing; Step S40, output the processing result as the text recognition result; In step S30, the step of inputting the two different augmented views Image1 and Image2 into a pre-trained CTC model for processing includes: Process the two different augmented views Image1 and Image2 through a shared visual encoder f to obtain corresponding feature distributions; x1 = f(Image1)(1-1) x2 = f(Image2)(1-2); Introduce a consistency regularization loss LCR to calculate the bidirectional Kullback-Leibler divergence of the two feature distributions and use it as a regularization term: where sg represents the stop gradient operation to ensure that the consistency regularization does not backpropagate to the target distribution.
2. The semantic enhancement text recognition method based on feature reconstruction and consistency CTC according to claim 1, wherein Step S00 includes: Step S001, preprocess the text image based on the height and aspect ratio of the original text image and a preset maximum aspect ratio; Step S002, combine a convolutional neural network and a Transformer architecture, extract local features of the preprocessed image through convolutional operations, imitate the structure of the Transformer block to enhance the feature interaction ability, and perform deep processing on the input features through the self-attention mechanism and a multi-layer perceptron to effectively capture global information and obtain two-dimensional features; Step S003, perform feature reconstruction on the two-dimensional features to convert them into a feature sequence that conforms to the reading order of the text image; Step S004, input the feature sequence that conforms to the reading order of the text image into a classifier, calculate the predicted character sequence, and align the predicted character sequence with the label sequence according to the CTC rule.
3. The semantic enhancement text recognition method based on feature reconstruction and consistent CTC according to claim 2, wherein Step S001 includes: Keep the height H of the original text image fixed, calculate the aspect ratio r of the original text image, and its calculation formula is: Among them, W orig and H orig respectively represent the width and height of the original text image; Round the aspect ratio r of the original text image to the nearest integer r round , and its calculation formula is: r round = round(r) (1 - 5); According to the rounded aspect ratio r round , calculate the new width W of the image new , and its calculation formula is: W new = H × r round (1 - 6); The step S001 further includes setting a maximum aspect ratio r max , if the calculated aspect ratio r round exceeds the maximum aspect ratio r max , then the width is limited to the corresponding maximum width W max = H × r max , and the final image width takes the smaller value of W new and W max , and its calculation formula is: W new = min(H × r round , H × r max ) (1 - 7); The final obtained image size is (W new , H), and the calculation formula is: (W new , H) = (min(H × r round , H × r max ), H) (1 - 8).
4. The semantic enhancement text recognition method based on feature reconstruction and consistency CTC according to claim 3, characterized in that Step S002 includes: Use PatchEmbedding to convert the input image into several local regions and represent the global features of the image through the embedding vectors of these regions. Specifically: Map the number of channels of the input image to a set number of channels embed_dim through two convolutional layers. Each convolutional layer includes a convolutional operation, batch normalization, and a non-linear activation function. The size of the convolutional kernel is 3×3, the stride is 2, the padding is 1, and the GELU activation function is used. The calculation formula is: x1 = ConvBNLayer(x)(1-9) x2 = ConvBNLayer(x1)(1-10); Among them, x1 and x2 are the outputs after two layers of convolution operations respectively, and x is the image input; For an image with an input size of H×W, after PatchEmbedding, the output feature map size is Each spatial position corresponds to a feature vector of embed_dim dimensions, and the final output feature representation is 5. The semantic enhancement text recognition method based on feature reconstruction and consistency CTC according to claim 4, wherein After the step S002, the following is further included: After the convolution operation, a multi-layer perceptron is used to perform non-linear transformation on the extracted features to enhance the learning ability of global information. Among them, a layer normalization and DropPath regularization mechanism are introduced in each stage of the non-linear transformation.
6. The semantic enhancement text recognition method based on feature reconstruction and consistent CTC according to claim 5, characterized in that The step S003 includes: Convert two-dimensional features into a feature sequence that conforms to the reading order of the text image That is, the relevant features are transferred from mapped to where the value ranges of j and m are The feature mapping process can be formally expressed with the help of a matrix That is, through matrix operations obtain the reconstructed feature sequence Among them, is a tensor over the real number field, and its shape where F is the feature extracted by feature extraction, H is the height of the original image input, W is the width of the original image input, and D2 is the number of channels that can be set.
7. The semantic enhancement text recognition method based on feature reconstruction and consistency CTC according to claim 6, characterized in that The learning matrix M is divided into horizontal direction reconstruction and vertical direction mapping: In the horizontal reconstruction, for each row of the feature map perform an unfolding operation to learn a horizontal rearrangement matrix The matrix elements reflect the features after horizontal reconstruction and the probabilities corresponding to the original feature F i,m Specifically, a variant of the self-attention mechanism is used to calculate the horizontal rearrangement matrix Specifically, perform a linear transformation on the feature row F i to obtain query, key, and value vectors. Let where is a learnable weight matrix, and the calculation process of the horizontal reconstruction matrix is as follows: First, calculate the attention score: To make the score more stable, a scaling operation is performed on it, and then it is converted into a probability distribution through the Softmax function, that is Based on learning Reconstruct the features of each row in the horizontal direction. The reconstructed features are processed through residual connection and multi-layer perceptron, and the specific calculation is as follows: Among them, LayerNorm represents the layer normalization operation, and MLP is a two-layer feed-forward neural network composed of two linear layers and an activation function; In the vertical direction mapping step, a learnable global context vector is introduced Let this global context vector interact with each column feature in F h to learn the vertical reconstruction matrix through such interaction The attention mechanism is also used to calculate the vertical reconstruction matrix First, the global context vector T is used as the query vector for the column features to perform a linear transformation to obtain the key vector where is a learnable weight matrix; The attention score is calculated as After a scaling operation a vertical reconstruction matrix is obtained through the Softmax function Based on the vertical reconstruction matrix Obtain the features after vertical reconstruction All column features share the same global context vector T; Finally, all the features after vertical reconstruction are combined to obtain denoted as the feature sequence after rearrangement 8. The semantic enhancement text recognition method based on feature reconstruction and consistency CTC according to claim 7, characterized in that The step S003 also includes: Rewrite the mapping relationship between the feature sequence and the original feature F: Among them, M j is a matrix that combines horizontal and vertical reconstruction information and multiplies with the original feature F to obtain the vertically reconstructed feature Combine all M j to obtain Verify again 9. The semantic enhancement text recognition method based on feature reconstruction and consistency CTC according to claim 8, characterized in that The step S00 also includes: Introduce context information, and combine the multi-head attention mechanism and context embedding to optimize the semantic understanding ability of the model, specifically including: For each character c in the text image with character label Y = {c1, c2,..., c L}, its context is composed of the left string i and the right string , where l is the context window length; s is the context window length; First, map the characters in to string embeddings by converting the characters into numerical features for subsequent model processing. Then, calculate the context representation of the left string through the multi-head attention mechanism Assume the number of heads is h. Linearly transform and the predefined token respectively to obtain queries, keys, and values for multiple heads: where is a learnable weight matrix, and the attention scores and outputs of the k-th head are respectively: The final output of the multi-head attention is: Among them is a learnable weight matrix that normalizes the output of the multi-head attention to obtain the context representation: Next, calculate the attention map through the multi-head attention mechanism Assume the number of heads is h, and linearly transform and the visual feature F respectively to obtain queries, keys, and values for multiple heads: K k = FW k,k (1 - 28); V k = FW v,k (1-29); Among them is a learnable weight matrix. The attention scores and outputs of the k-th head are respectively The final output of the multi-head attention is: where is a learnable weight matrix. The output of the multi-head attention is normalized to obtain the attention map: Using the attention map weight the visual feature F to obtain the feature corresponding to the character c i : Input into a fully connected layer for classification prediction: Among them is the weight matrix, is the bias vector, and Nc is the size of the character set; finally, use the prediction result and the true label c i to calculate the cross-entropy loss to train the model; The processing of the right-side string is symmetrical to that of the left-side string, and finally an attention map will be obtained visual features and prediction results and participate in the loss calculation 10. The semantic enhancement text recognition method based on feature reconstruction and consistency CTC according to claim 9, characterized in that The step S003 also includes: The model is constrained and optimized based on the connection attention time classification loss, semantic supervision loss, and consistency regularization loss. Among them, The calculation formula of the connection attention time classification loss is: Among them, is the prediction of the model for the i-th character, F v is the feature sequence rearranged by the feature reconstruction module, Y is the true label sequence, represents the probability of predicting the i-th character under the condition of the given rearranged feature F v and the true label Y; The semantic supervision loss is measured by calculating the cross-entropy between the character category probability based on the context prediction information and the true label. The specific calculation formula is: Among them, and are the predicted class probabilities obtained by using the string information on the left and right sides of the target character, respectively. c i is the true label of the i-th character, and ce represents the cross-entropy loss function; The consistency regularization loss is calculated by calculating the consistency between the CTC predictions of two branches. Its calculation formula is: L cr = L CR (x1, x2) (1 - 38); Among them, represents the Kullback-Leibler divergence, and predicts1 and predicts2 are the CTC prediction distributions of the two branches respectively; Integrate the CTC loss, semantic supervision loss, and consistency regularization loss to obtain the total loss function of the text recognition algorithm based on feature reconstruction and semantic enhancement: L total = αL ctc + βL gtc + γL cr (1 - 39); Add it to the model training process, and use the gradient descent algorithm for iteration until the maximum iteration number is reached or the model converges.
Citation Information
Cited By
Electric power communication cross-medium data identification method and device and storage medium
CN122200675A
Power communication cross-medium data identification method and device and storage medium
CN122200675B