Self-adaptive length handwritten formula recognition method based on global attention mechanism
By introducing global attention mechanism and adaptive length processing in formula recognition technology, the problems of limited recognition accuracy and large preprocessing requirements of traditional methods when dealing with complex formulas are solved, and higher recognition accuracy and applicability are achieved.
Patent Information
- Application Number
- CN202510267900.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
The existing formula recognition technology relies on traditional convolutional neural networks, and it is difficult to effectively process complex formulas and requires a lot of preprocessing, which affects the recognition speed and accuracy.
Adaptive length handwriting formula recognition method based on global attention mechanism is adopted, and handwriting formula images are processed through image enhancement, feature extraction and multi-head attention mechanisms, and vocabulary distribution is generated and formulas are predicted.
This method can show excellent performance in complex image-to-text recognition tasks, adapt to diverse usage scenarios, improve the analytical fault tolerance of multi-line formulas and nested structures, and improve the applicability of processing complex formulas.
Smart Images

Figure CN120220166A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of formula recognition, and particularly to an adaptive-length handwritten formula recognition method based on a global attention mechanism. Background Art
[0002] With the popularization of digital education and scientific research, the demand for mathematical expression recognition has been increasing. Mathematical Expression Recognition (MER) is a key task in document analysis, aiming to convert image-based mathematical expressions into corresponding markup languages. MER is important in applications such as scientific document extraction, and a powerful MER model helps maintain the logical consistency and structural integrity of the document. Most existing formula recognition technologies rely on traditional Convolutional Neural Network (CNN) structures, which have problems in effectively handling complex formula recognition and limited recognition accuracy. In addition, traditional models often require a large amount of preprocessing when dealing with mathematical formulas, affecting the recognition speed and accuracy. Summary of the Invention
[0003] Aiming at the above problems, the purpose of the present invention is to provide an Adaptive Length Handwritten Formula Recognition System based on a global attention mechanism.
[0004] The specific technical solution to achieve the purpose of the present invention is as follows:
[0005] An adaptive-length handwritten formula recognition method based on a global attention mechanism includes the following steps:
[0006] Step 1: Enhance the input handwritten formula image to obtain the enhanced image I enhanced ;
[0007] Step 2: Extract features from the enhanced handwritten formula image;
[0008] Step 3: Process the features of the extracted handwritten formula image based on the multi-head attention mechanism to generate a vocabulary distribution and obtain the predicted formula.
[0009] Compared with the prior art, the beneficial effect of the present invention is as follows:
[0010] This solution proposes an Adaptive Length Handwritten Formula Recognition System (ALHFRS) based on a global self-attention mechanism, including image enhancement, extraction of local and global features of the image, and formula generation. Among them, rich image enhancement methods, such as image dilation, erosion, weather noise, etc., improve the model performance in low-light, blurred and other scenarios, enabling the model to adapt to more diverse usage scenarios;
[0011] The global information coding integration enables the system to learn the overall information of the formula and adjust its prediction, showing excellent performance in complex image-to-text recognition tasks such as mathematical formula recognition.
[0012] Finally, when generating the formula, by introducing hierarchical features and global representations, and introducing context information in the decoding stage, it can predict the termination position of the formula symbol sequence in real time by enhancing semantic associations. Compared with the traditional fixed-length decoding strategy, while avoiding redundant calculations and saving resources, it shows stronger parsing fault tolerance for complex mathematical expressions such as multi-line formulas and nested structures, and can also improve the applicability to a certain extent when processing complex formulas.
[0013] The following further describes the present invention in conjunction with specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic flowchart of the adaptive-length handwritten formula recognition method based on the global attention mechanism for this solution. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] EXAMPLE
[0016] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] As shown in this application and the claims, unless the context clearly indicates an exception, the words "a", "an", "one" and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list, and the method or device may also include other steps or elements.
[0018] Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present application. At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the authorization specification. In all the examples shown and discussed here, any specific value should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that: similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0019] Combined with Figure 1 , an adaptive-length handwritten formula recognition method based on a global attention mechanism, comprising the following steps:
[0020] Step 1. Enhance the input handwritten formula image, and process the real-world test image using the image dilation method, image erosion method, and adding weather noise method of morphological operations to obtain the enhanced image I enhanced :
[0021] Step 1-1. Perform image dilation processing on the input handwritten formula image:
[0022] Specifically, for each pixel in the image, the image dilation method views the neighborhood determined by the structuring element around the pixel and assigns the maximum value in the neighborhood to the pixel to "expand" the highlighted area:
[0023] Take the structuring element S as a sliding window, traverse each pixel of the original image, and achieve geometric constraints on the target through extreme value calculation:
[0024]
[0025] where I dilated (x, y) is the dilated image, S represents the defined structuring element, (x, y) represents the pixel coordinates of the image, and (x ′ , y ′ ) represents the pixel coordinates offset based on the structuring element S.
[0026] Step 1-2. Perform image erosion processing on the input handwritten formula image. The image erosion method views the neighborhood determined by the structuring element around each pixel in the image and assigns the minimum value in the neighborhood to the pixel, reducing and refining the highlighted area or white part in the image. Specifically:
[0027] Taking the structural element S as a sliding window, traverse each pixel of the original image, and achieve geometric constraints on the target through extreme value calculation:
[0028]
[0029] Among them, I eroded (x, y) is the eroded image, S represents the defined structural element, (x, y) represents the pixel coordinates of the image, and (x ′ , y ′ ) represents the pixel coordinates offset based on the structural element S.
[0030] Step 1-3: Add weather noise to the input handwritten formula image. The method of adding weather noise first generates Gaussian noise to simulate the movement of raindrops, then uses a mask matrix to control the distribution of raindrops, and finally superimposes it on the original image:
[0031] First generate Gaussian noise to simulate the movement of raindrops, then use a mask matrix to control the distribution of raindrops, and finally superimpose it on the original image:
[0032] R(x, y) = α × G(x, y) × M(x, y)
[0033] I rain (x, y) = I(x, y) + R(x, y)
[0034] G(x, y) is Gaussian noise used to simulate the movement trajectory of raindrops, M(x, y) is the mask matrix for controlling the distribution of raindrops, randomly set to 0 or 1, indicating whether there is a raindrop at this point, and α is the noise intensity parameter.
[0035] Step 1-4: Merge the image data processed in Step 1-1 to Step 1-3 to obtain the enhanced image set I enhanced .
[0036] Step 2: Extract features from the enhanced handwritten formula image. The input of this step is the enhanced image set I obtained in Step 1 enhanced ;
[0037] Step 2-1: Extract the hierarchical features Z enhanced of the enhanced image I l+1 :
[0038] For the enhanced image I enhancedIt is divided into non - overlapping image patches of a fixed size, and mapped to a vector representation of a fixed dimension through a linear projection layer to obtain the embedding matrix Z of all image patches. First, the input image Img in the training batch is divided into non - overlapping image patches of a fixed size. The size of the input image is H×W×C, where H and W are the height and width of the image respectively, and C is the number of channels of the image. The image is divided into P×P small patches Patch i , and each patch is mapped to a vector representation of a fixed dimension through a linear projection layer:
[0039] z i =Linear(Flatten(Patch i ))
[0040] To obtain the embedding matrix Z of all image patches, whose size is Calculate the self - attention mechanism for each image patch z i :
[0041]
[0042] Among them, Q = z i W Q , K = z i W K , V = z i W V , W Q , W K , W V are learnable projection matrices, and z i represents the i - th image patch;
[0043] After the attention calculation of each layer, the embedding matrix Z of the image patches is further processed through a multi - layer perceptron:
[0044] Z ′ =MLP(Z)
[0045] Then, hierarchical features are constructed through downsampling to obtain the hierarchical features of the final enhanced image I enhanced :
[0046] Emb I =LayerNorm(Z l+1 +Z ′ )
[0047] Z l+1 =Downsample(Z ′ )
[0048] Step 2 - 2, Extract the global representation information Emb enhancee of the enhanced image I global, which can capture global features while ensuring computational efficiency;
[0049] Input the enhanced image I enhanced 's hierarchical features into a four-layer multi-head attention (MHSA), and concatenate the outputs of all heads:
[0050] MHSA(Z) = Concat(head1, …, head h )W i
[0051] where head i = Attention(Q i , K i , V i ), and W i is the linear transformation weight matrix of the i-th layer. Here, h is taken as 16, and the output is Z4:
[0052] Z4 = MHSA4(MHSA3(MHSA2(MHSA1(Emb I ))))
[0053] And transform the global features after global pooling through a final fully connected layer to obtain the final global representation vector:
[0054] Emb global = FC(GAP(ReLU(Z4)))
[0055] Step 3: Process the features of the handwritten formula image extracted based on the multi-head attention mechanism to generate a vocabulary distribution and obtain the predicted formula:
[0056] Take the hierarchical features Z l+1 and the global representation Emb global together as the context and input them into the multi-head attention mechanism. Map the output of the multi-head attention mechanism to the dimension of the vocabulary size through a linear layer to generate a vocabulary distribution for predicting the next word;
[0057] Q ′ = Emb global W Q′ , K ′ = Z l+1 W K′ , V ′ = Z l+1 W V′
[0058]
[0059] H out= ReLU(Decoder - Attention(Q ′ , K ′ , V ′ ))
[0060] The output H of the decoder out is mapped to the dimension of the vocabulary size through a linear layer to generate a vocabulary distribution for predicting the next word:
[0061] P(y t |y <t , Z l+1 , Emb global ) = Softmax(H out W vocab + b vocab )
[0062] where W vocab is the weight matrix of the output layer, with size D×V, and b vocab is the bias of the output layer;
[0063] Finally, the next word in the target sequence is generated by taking the word with the highest probability, and the final prediction is completed by the formula generation:
[0064] y t+1 = argmax P(y t |y <t , Z l+1 , Emb global )
[0065] As shown in Table 1, for the dataset selection, the present invention chose the im2latex dataset, which contains 70,000 training data and 30,000 test data. Each picture consists of handwritten formulas with black characters on a white background. The Texify method was used for comparison. The evaluation metrics are respectively the BLEU score, which quantifies the n - gram matching degree between the candidate sentence and the reference sentence; the edit distance score, which measures the minimum number of character changes required to convert the prediction result to the true value; and the expression recognition rate, which is the percentage of the predicted mathematical expression that perfectly matches the actual result.
[0066] Table 1 Schematic table of experimental results
[0067]
[0068]
[0069] As shown in Table 2, this experiment compared the performance differences between the original model and two feature ablation models. The evaluation metrics include the BLEU score, the edit distance score (the lower the better), and the mathematical expression recognition rate. The ablation experiments include the following two: deleting the global information feature Emb globalWith the deletion of the hierarchical information feature Z l+1 in the context;
[0070] Among them: The ablation effect of the global information feature: The BLEU score drops by 0.1 percentage point, the edit distance increases significantly by 17.2%, and the expression recognition rate decreases by 6.3 percentage points, indicating that the global feature has an obvious effect on optimizing sentence coherence.
[0071] The ablation effect of the hierarchical information feature: The BLEU score drops by 2.4 percentage points, which is the largest decline. The decline rate of the expression recognition rate reaches 8.1 percentage points (67.40% → 59.30%), reflecting the key role of the hierarchical feature in modeling complex structures.
[0072] Advantages of the original model: Achieve the best balance of the three indicators while retaining all features, proving the complementarity of the global and hierarchical features. Especially in the expression recognition task, the accuracy of the complete model is improved by 4.3 - 8.1 percentage points compared with the ablation model.
[0073] Table 2 Schematic table of ablation experiment results
[0074]
[0075] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for adaptive length handwritten formula recognition based on global attention mechanism, characterized in that: The following steps are involved: Step 1: Enhance the input handwritten formula image to obtain the enhanced image I enhanced ; Step 2: Extract features from the enhanced handwritten formula image; Step 3: Process the extracted handwritten formula image features based on the multi-head attention mechanism, generate vocabulary distribution, and obtain the predicted formula.
2. The method for adaptive length handwritten formula recognition based on global attention mechanism according to claim 1, characterized in that: The step 1 of enhancing the input handwritten formula image is specifically as follows: Step 1-1, performing image dilation processing on the input handwritten formula image; Step 1-2, performing image corrosion processing on the input handwritten formula image; Step 1-3, adding weather noise to the input handwritten formula image; Step 1-4: Merge the image data processed from step 1-1 to step 1-3 to obtain an enhanced image set I enjanced .
3. The method for adaptive length handwritten formula recognition based on global attention mechanism according to claim 2, characterized in that: The image expansion processing in step 1-1 is specifically as follows: The structural element S is used as a sliding window to traverse each pixel of the original image and realize the geometric constraints of the target through extreme value calculation: Among them, I dilated (x, y) is the expanded image, S represents the defined structural element, (x, y) represents the pixel coordinates of the image, and (x′, y′) represents the pixel coordinates after offset based on the structural element S.
4. The method for adaptive length handwritten formula recognition based on global attention mechanism according to claim 2, characterized in that: The image corrosion processing in step 1-2 is specifically as follows: The structural element S is used as a sliding window to traverse each pixel of the original image and realize the geometric constraints of the target through extreme value calculation: Among them, I eroded (x, y) is the image after erosion, S represents the defined structural element, (x, y) represents the pixel coordinates of the image, and (x′, y′) represents the pixel coordinates after offset based on the structural element S.
5. The method for adaptive length handwritten formula recognition based on global attention mechanism according to claim 2, characterized in that: The step 1-3 of adding weather noise to the input handwritten formula image is specifically as follows: First, generate Gaussian noise to simulate the movement of raindrops, then use the mask matrix to control the distribution of raindrops, and finally superimpose it on the original image: R(x,y)=α×G(x,y)×M(x,y) I rain (x,y)=I(x,y)+R(x,y) G(x,y) is Gaussian noise used to simulate the motion trajectory of raindrops, M(x,y) is the mask matrix that controls the distribution of raindrops, which is randomly set to 0 or 1 to indicate whether there is a raindrop at the point, and α is the noise intensity parameter.
6. The method for adaptive length handwritten formula recognition based on global attention mechanism according to claim 1, characterized in that: The feature extraction of the enhanced handwritten formula image in step 2 is specifically as follows: Step 2-1: Extract enhanced image I enhanced The hierarchical features Z l+1 ; Step 2-2: Extract the enhanced image I enhanced Global representation information Emb global .
7. The method for adaptive length handwritten formula recognition based on global attention mechanism according to claim 6, characterized in that: The image I after extraction and enhancement in step 2-1 enhanced The hierarchical features are: The enhanced image I enhanced The image is divided into non-overlapping fixed-size image blocks, which are mapped to a fixed-dimensional vector representation through a linear projection layer to obtain the embedding matrix Z of all image blocks: Where Q = z i W Q ,K=z i W K ,V=z i W V , W Q ,W K ,W V is the learnable projection matrix, z i represents the i-th image block; The embedding matrix Z of the image block is further processed by a multi-layer perceptron: Z′=MLP(Z) Then, hierarchical features are constructed by downsampling to obtain the final enhanced image I enhanced Hierarchical features: Emb I =LayerNorm(Z l+1 +Z′) WITH l+1 =Downsample(Z′)。 8. The method for adaptive length handwritten formula recognition based on global attention mechanism according to claim 7, characterized in that: The image I after extraction and enhancement in step 2-2 enhanced The global representation information is as follows: The enhanced image I enhanced The hierarchical features of are input into the four-layer multi-head attention, and the outputs of all heads are spliced, and the global features after global pooling are transformed through a final fully connected layer to obtain the final global representation vector: MHSA(Z)=Concat(head1,…,head h )W i Z4=MHSA4(MHSA3(MHSA2(MHSA1(Emb I )))) Emb global =FC(GAP(ReLU(Z4))) Among them, head i =Attention(Q i ,K i ,V i ), W i is the linear transformation weight matrix of the i-th layer.
9. The method for adaptive length handwritten formula recognition based on global attention mechanism according to claim 1, characterized in that: The formula for obtaining the prediction in step 3 is specifically: The hierarchical feature Z l+1 and global representation Emb global The output of the multi-head attention mechanism is mapped to the dimension of the vocabulary size through a linear layer to generate a vocabulary distribution for predicting the next word. Q′=Emb global W Q′ ,K′=Z l+1 W K′ ,V′=Z l+1 W V′ H out =ReLU(Dec-Attention(Q′,K′,V′)) P(y t |y <t ,Z l+1 ,Emb global )=Softmax(H out W vocab +b vocab ) Finally, the next word in the target sequence is generated by taking the word with the highest probability, and the final prediction is completed by formula generation: and t+1 =argmaxP(y t |and <t ,Z l+1 ,Emb global ) Among them, W vocab is the weight matrix of the output layer, size D×V, b vocab is the bias of the output layer.