Handwritten mathematical formula identification method based on attention coverage and position awareness

By employing a position awareness module and multi-scale cover attention mechanism, the method effectively addresses character position extraction and attention drift in hand-written mathematical formula recognition, improving accuracy and robustness in complex structures.

CN120318833APending Publication Date: 2025-07-15BEIJING UNIV OF TECH

Patent Information

Application Number
CN202510396462.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

When using complex structures, existing handwritten mathematical formula recognition methods are difficult to effectively capture character position information and optimize attention mechanisms, resulting in insufficient recognition accuracy.

Method used

The position perception module and multi-scale coverage attention mechanism are adopted to generate a position perception map through multi-scale feature extraction and channel attention enhancement, and fuse it with the original feature map, and combine it with the Transformer decoder for LaTeX sequence generation, and use count loss and sequence loss weighting to optimize model parameters.

Benefits of technology

Significantly improves the recognition accuracy of handwritten mathematical formulas, especially when dealing with complex nested structures, reduces training costs and maintains high robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318833A_ABST
    Figure CN120318833A_ABST
Patent Text Reader

Abstract

The invention discloses a handwritten mathematical formula recognition method based on attention coverage and position perception, which comprises the following steps: constructing a position perception module, and generating a character position map through weak supervised learning; a multi-scale attention coverage module is constructed, and attention calculation in a Transform decoding architecture is optimized; and fusing the feature map and the position sensing map, and inputting the fused feature map and the fused position sensing map into a Transform decoder to generate a LaTeX sequence. Character position information is explicitly extracted through the position sensing module, and the understanding of the model on a complex mathematical structure is enhanced; the attention drifting is reduced by a multi-scale attention coverage mechanism; the training cost is reduced and the model performance is improved by adopting a counting label weak supervised learning mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of handwritten mathematical expression recognition, and particularly to a method for recognizing handwritten mathematical formulas based on coverage attention and position perception. Background Art

[0002] Handwritten Mathematical Expression Recognition (HMER) is a core technology in multiple application fields such as automatic grading, error correction systems, office automation, and mathematical formula image retrieval. According to the different forms of input data, HMER can be divided into two categories: offline recognition and online recognition. The main task of offline recognition is to extract and generate the corresponding LaTeX sequence from a static handwritten mathematical formula image.

[0003] In practical applications, the main challenges faced by HMER include: firstly, the significant differences in writing styles result in multiple representations of the same symbol; secondly, the complex two-dimensional layout of mathematical formulas makes it difficult to accurately capture the spatial relationships between symbols. These factors make it particularly difficult to accurately parse the LaTeX sequence from a static image, especially in an offline environment.

[0004] Traditional HMER methods usually adopt a three-step process: character segmentation, character recognition, and structure analysis. However, this method is prone to error accumulation during the processing, affecting the final recognition accuracy. For example, Li et al. proposed a fine-grained segmentation method based on YOLOv5s, which improved the symbol segmentation effect through object detection technology. Tang et al. introduced a graph neural network on this basis, attempting to break through the limitations of single detection through spatial topological association. This method dynamically learns the potential correlations between nodes through a graph neural network for spatial information aggregation, but still fails to completely eliminate the errors caused by symbol segmentation. At the same time, due to the complexity of handwritten strokes and the structure of mathematical expressions, the prediction accuracy of the relationships between nodes is insufficient.

[0005] With the development of deep learning technology, the end-to-end encoder-decoder architecture has been widely applied to HMER tasks, regarding it as an image-to-sequence prediction task. These methods are mainly divided into two ideas: one is to directly use a sequence decoder to convert the image into a LaTeX sequence; the other is to use a tree decoder to convert the image into a tree structure, and then convert it into a LaTeX sequence through post-processing. The tree decoder usually relies on a set of custom grammar rules, aiming to structure the relationship between the picture and the LaTeX sequence. Yuan et al. proposed a new set of grammar rules, effectively decomposing the syntax tree into different components and reducing the ambiguity in the tree structure. Li et al. proposed a branch parallel decoding model from the perspective of decoding efficiency, significantly improving the decoding speed. However, these methods rely on predefined tree rules, restricting their generality and flexibility.

[0006] Meanwhile, the rapid development of sequence decoders in the field of natural language processing has provided new research ideas for HMER. Zhao et al. proposed using the Transformer Decoder to replace the traditional GRU decoder, opening up a new research direction. The Transformer effectively solves the long-term dependence problem by introducing the scaled dot-product attention mechanism and generates an attention weight matrix by calculating the correlation between each word. However, the parallel computing feature of the Transformer leads to the abandonment of the coverage attention mechanism, thus affecting the accuracy of attention. To make up for this deficiency, Zhao et al. introduced the coverage attention mechanism based on the Transformer Decoder, proposed the CoMER model, and significantly improved the model performance on the basis of ensuring parallel training by optimizing the attention refinement module, achieving obvious performance improvements on multiple datasets.

[0007] However, there are still two key problems in the existing methods: one is the insufficient extraction of character position information in the image, and the other is the drift phenomenon of the attention mechanism when dealing with complex structures. To address these problems, Li et al. started from the character density map and proposed a weakly supervised learning network CAN (Counting Awareness Network). By processing the original LaTeX string, a counting vector for each character is obtained, and through such weak labels, an attempt is made to enable the counting awareness module to perceive the global information in the picture. However, there are still obvious errors in this weakly supervised learning, which limits the prediction accuracy of the decoder.

[0008] In summary, there is an urgent need for a new method that can effectively capture character position information and optimize the attention mechanism to improve the accuracy of handwritten mathematical formula recognition. Summary of the Invention

[0009] The purpose of the present invention is to solve the defects existing in the prior art and provide a method for recognizing handwritten mathematical formulas based on coverage attention and position perception. This method effectively improves the accuracy of handwritten mathematical formula recognition through a position perception module and a multi-scale coverage attention mechanism.

[0010] To achieve the above purpose, the technical solution of the present invention is: a method for recognizing handwritten mathematical formulas based on coverage attention and position perception, characterized by including: S1. Construct a position perception module to generate a character position map through counting labels and weakly supervised learning, enhancing the feature extraction ability; S2. Construct a multi-scale coverage attention mechanism to optimize attention calculation in the Transformer decoding architecture, reducing the attention drift problem; S3. Transmit the feature map into the position perception module to generate a position perception map and fuse it with the original feature map; S4. Input the fused feature map into the Transformer decoder to generate a LaTeX sequence; S5. Optimize the training of model parameters through weighted optimization of counting loss and sequence loss; S6. Preprocess the image to be recognized and input it into the trained model to obtain the recognition result.

[0011] Further, the step S1 constructs a position-aware module, which specifically includes: Use a multi-scale feature extraction network to extract features using convolutional kernels of different sizes (3×3 and 5×5); introduce a channel attention mechanism (SE block) after the convolutional layer to further enhance the feature extraction ability; use a fully connected layer to map the features to the dimensional space of the number of symbol classes; introduce counting loss to assist in training, and calculate the counting vector of each symbol class through global pooling operation; finally, use layer normalization (LayerNorm) to fuse the position-aware features with the original feature map.

[0012] Further, the step S2 constructs a multi-scale coverage attention mechanism, which specifically includes: On the basis of the traditional coverage attention mechanism, design a multi-scale coverage attention refinement module (MARM); use convolutional kernels of multiple sizes (3×3 and 5×5) to jointly calculate the coverage attention refinement value; fuse the attention refinement results of multiple scales to form the final attention correction; combine the original attention matrix with the refinement matrix through subtraction operation to generate the corrected attention weight.

[0013] Further, the step S3 inputs the feature map into the position-aware module to generate a position-aware map and fuse it with the original feature map, which specifically includes: Input the feature map extracted by the DenseNet encoder into the convolutional branches of 3×3 and 5×5 respectively; apply the SE block to the output features of each branch to enhance the channel attention; map the enhanced features to the symbol class space through a fully connected layer to generate a position-aware map; fuse the position-aware map with the original feature map according to the weight α: P = LN(α·M + X), where LN represents layer normalization (LayerNorm), and the default value of α is 0.5.

[0014] Further, the step S4 inputs the fused feature map into the Transformer decoder to generate a LaTeX sequence, which specifically includes: Use the fused feature map as the input of the Transformer decoder; the decoder internally uses the multi-head self-attention mechanism to process the sequence information; introduce the multi-scale coverage attention mechanism to process the attention allocation during the decoding process; map the decoding result to the symbol space through a linear classification head to generate a LaTeX sequence.

[0015] Furthermore, in step S5, the model parameters are trained through weighted optimization of the counting loss and the sequence loss, which specifically includes: Calculating the cross-entropy loss L of the sequence prediction rec , expressed as: Calculating the loss L of the symbol count count , expressed as: Adjusting the ratio of the two losses through weights: L total = λ1·L rec + λ2·L count , where λ1 and λ2 are weight coefficients, and the default ratio is λ1:λ2 = 2:1; Using the Adam optimizer for gradient descent to update the model parameters.

[0016] Furthermore, in step S6, the image to be recognized is preprocessed and input into the trained model to obtain the recognition result, which specifically includes: preprocessing the input image such as scaling and normalization; inputting the preprocessed image into the trained model; obtaining the LaTeX sequence output by the model as the recognition result; and post-processing the recognition result to ensure compliance with the LaTeX syntax specification.

[0017] The present invention adopts the following technical means: (1) Position perception module: Through multi-scale feature extraction and channel attention mechanism, explicit perception of character positions is achieved, enhancing the model's ability to understand complex structural relationships in mathematical formulas; (2) Multi-scale coverage attention mechanism: By jointly calculating the coverage attention refinement values with convolution kernels of different scales, the attention calculation process is optimized, reducing the attention drift problem; (3) Weak supervision learning strategy: The position perception module is trained through counting labels, without additional annotation, reducing the training cost; (4) Feature fusion mechanism: By weighted fusion of the position perception map and the original feature map, the information richness of the original features is retained, while enhancing the expression ability of spatial positions; (5) Multi-task joint optimization: By simultaneously optimizing the sequence prediction task and the position perception task, the overall performance of the model is improved.

[0018] The beneficial effects of the present invention: 1. The present invention explicitly extracts character position information through the position perception module, effectively enhancing the model's ability to understand complex structural relationships in mathematical formulas and improving the recognition accuracy; 2. The multi-scale coverage attention mechanism optimizes the attention calculation process, significantly reduces the attention drift problem, and enhances the model's ability to parse complex formulas; 3. The weakly supervised learning method for counting tags requires no additional annotation, reduces the training cost, and improves the model performance at the same time; 4. On multiple benchmark datasets such as CROHME2014, CROHME2016, and CROHME2019, the method of the present invention has achieved significant performance improvements, reaching accuracies of 64.06%, 63.21%, and 65.22% respectively, exceeding the existing optimal methods. 5. The present invention is particularly outstanding for mathematical formulas with complex nested structures (such as fractions, exponents, multiple integrals, etc.), and the improvement in recognition accuracy is more significant, with strong practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is the overall architecture diagram of the handwritten mathematical formula recognition method based on coverage attention and position perception of the present invention;

[0020] Figure 2 It is the structural diagram of the position perception module of the present invention;

[0021] Figure 3 It is the structural diagram of the multi-scale coverage attention refinement module of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0022] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the technical solution and working principle of the present invention. It should be noted that for those of ordinary skill in the art, several adjustments and improvements can be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0023] As Figure 1 shown, the handwritten mathematical formula recognition method based on coverage attention and position perception proposed by the present invention mainly includes four core components: a DenseNet backbone network, a position perception module (PAM), a multi-scale coverage attention refinement module (MARM), and a Transformer decoder.

[0024] Given a handwritten mathematical formula image, the goal of the present invention is to generate the corresponding LaTeX sequence, so that the rendered formula can accurately restore the mathematical formula in the image. First, the DenseNet backbone network extracts two-dimensional visual features from the input image. These features are then fed into the position-aware module, which generates a position-aware map through depth feature extraction. After fusing the visual features with the position features, they are input into the Transformer decoder for sequence decoding, so as to obtain a discriminative symbol representation. Finally, these features are mapped to the LaTeX expression output through a linear classification head.

[0025] As Figure 2 shown, the core design of the position-aware module (PAM) includes multi-scale feature extraction, channel attention enhancement, and weakly supervised counting learning. Considering the diversity of character sizes in handwritten images, PAM uses convolutional kernels of different sizes (3×3 and 5×5) to extract multi-scale features. After the convolutional layer, the SE (Squeeze-and-Excitation) block is applied for channel attention enhancement. For the feature map extracted by the 3×3 branch The enhanced feature S can be expressed as: Q = σ(W1G(H)+b1) where G is global average pooling, and σ and g(·) represent the ReLU function and the sigmoid function respectively, represents channel multiplication, and W1, W2, b1, b2 are trainable weights.

[0026] After obtaining the enhanced feature S, a fully connected layer is used to reduce the number of channels to C, where C is the number of symbol classes. The feature values are mapped to the range (0, 1) through the sigmoid function to generate the character position map For each M i ∈R H×W , it reflects the position distribution of the i-th symbol class in the image. Through the sum pooling operation, the count vector of each symbol class can be obtained: The count vector V i is used to calculate the counting loss during training to achieve the independent learning of the PAM module.

[0027] During model training, we first need to construct a count encoding based on the label count of the sequence task, screening out some invisible characters, such as the structural symbols "{", "}", "^" and "_", and focusing the model task on the position perception of visible characters in the picture. Our purpose is to enable the model to indirectly perceive the character position according to the different numbers of corresponding characters through the count label. Although the PAM module learns through the count loss, after a large number of picture learnings, it can still generate a character position map with relatively high accuracy.

[0028] Finally, fuse the character position perception map with the original feature map to obtain the position perception feature map: P = LN(α·M + X) where α is the weight parameter with a default value of 0.5, and LN(·) represents the LayerNorm normalization operation. Compared with the count result of a single number, these feature maps containing character positions are more valuable and can provide richer spatial position information.

[0029] As Figure 3 shown, the multi-scale covered attention refinement module (MARM) is an improvement to the original covered attention module. Covered attention has always been the key to handwritten mathematical formula recognition. The CoMER model calculates the covered attention separately through the ARM module and then optimizes it by subtraction. In the original ARM, for the feature extraction of covered attention, only a 5×5 convolutional kernel is used. The MARM module we proposed adds convolutional kernels of multiple scales (3×3 and 5×5) to jointly calculate the refined value of covered attention on the basis of the original model.

[0030] Specifically, the MARM module first performs multi-scale feature extraction on the covered attention matrix, then fuses the features of different scales to form a more refined attention representation. Finally, the original attention matrix is combined with the refined attention matrix through a subtraction operation, effectively reducing the attention drift problem. Through the MARM module, the accuracy of the attention module can be further improved, especially showing obvious advantages when dealing with complex nested-structured mathematical formulas.

[0031] During the training process, the model optimizes two main loss functions: the sequence recognition loss L rec and the count loss L count . Specifically, L rec calculates the sequence prediction loss based on cross-entropy: where represents the true label of the t-th symbol, represents the predicted probability of the model for the t-th symbol, and T represents the sequence length.

[0032] Counting loss L count Then, based on weak supervision learning, it is used to optimize the position-aware module: where and respectively represent the true labels of the nesting level and the relative position relationship between symbols, and represent the corresponding predicted probabilities.

[0033] The two losses are weighted by the weight coefficients λ1 and λ2: L total = λ1·L rec + λ2·L count where, by default, λ1:λ2 = 2:1 to balance the learning objectives of the two tasks. In the model inference stage, only the main sequence prediction task is used, and no counting prediction is required, so no additional computational cost is added.

[0034] This invention is consistent with the baseline model CoMER, using the DenseNet structure as the backbone network, three-layer TransformerDecoder as the decoder, and a linear layer for symbol classification. By default, the weight α in PAM feature fusion is 0.5. To improve the PAM learning speed, the ratio of the two loss functions is set as L counting :L rec = 2:1. The training loss of the PAM module decreases rapidly, and it can more effectively provide position information for the recognition backbone network.

[0035] The experimental results are shown in Table 1. The method of this invention has achieved significant performance improvements on multiple standard datasets. CROHME is a publicly available single-line handwritten mathematical expression benchmark, and the training set consists of 8836 expression images. The CROHME 2014 / 2016 / 2019 test sets contain 986, 1147, and 1199 expression images respectively. On these test sets, the method of this invention reaches accuracies of 64.06%, 62.51%, and 66.38% respectively, exceeding the existing best methods.

[0036] Table 1: Test results on CROHME

[0037] For different error tolerances, that is, the recognition rates when allowing ≤1, ≤2, and ≤3 symbol-level errors, the present invention also performs excellently. Specifically, on CROHME2014, the corresponding indicators are 79.19%, 85.69%, and 89.44%; on CROHME2016, the corresponding indicators are 78.81%, 84.92%, and 89.97%; on CROHME2019, the corresponding indicators are 83.66%, 88.66%, and 91.58%. These results all exceed the existing best methods.

[0038] To verify the effectiveness of the position-aware module and the multi-scale coverage attention mechanism, a series of ablation experiments were conducted:

[0039] 1. After removing the position-aware module, the model accuracy decreased by 0.51%, 0.78%, and 3.58% on CROHME2014, CROHME2016, and CROHME2019 respectively, proving the importance of position awareness for complex structure recognition;

[0040] 2. When using single-scale coverage attention, the model accuracy decreased by 0.71% and 0.44% on CROHME2014 and CROHME2016 respectively, indicating that multi-scale attention can capture complex attention patterns more effectively;

[0041] Table 2: Ablation experiment table on CROHME

[0042] In addition, by visualizing the attention weight distribution, it can be observed that the attention in the method of the present invention is more concentrated on the characters to be recognized currently, reducing the attention drift phenomenon, especially being more prominent when dealing with complex nested structures. When dealing with nested structures such as fractions, exponents, etc., the traditional method is prone to attention dispersion or misalignment, while the multi-scale coverage attention mechanism of the present invention can accurately locate the characters to be recognized currently, thereby improving the recognition accuracy.

[0043] The introduction of the position-aware module enables the model to better understand the two-dimensional structural relationship in mathematical formulas. Through weakly supervised counting learning, the generated position-aware map clearly shows the distribution of each symbol in the image, providing important spatial information guidance for the Transformer decoder. Especially for complex nested structures, such as the numerator and denominator in fractions, exponents, and subscripts, the position-aware map can accurately represent their relative position relationships, greatly enhancing the model's ability to grasp structural features.

[0044] The advantages of the multi-scale coverage attention mechanism are mainly reflected in processing long sequences and complex structures. By performing refined processing on the coverage attention through multi-scale convolution, it can better capture local and global attention patterns, reducing the problem of attention drift in complex structures. Especially for formulas with multiple levels of nesting, such as continued fractions and multiple integrals, the multi-scale coverage attention can maintain the coherence and accuracy of attention, ensuring the smooth progress of the decoding process.

[0045] Another advantage of the method of the present invention is the adoption of a weak supervision learning strategy. By counting the visible characters in the LaTeX sequence and constructing a count encoding as a weak supervision signal, without additional annotation work, it can guide the learning of the position perception module. This low-cost supervision method enables the model to be directly trained on existing datasets without complex preprocessing or special annotation work, having high practicality.

[0046] At the same time, the fusion strategy adopted by the present invention is also worthy of note. By performing weighted fusion of the position perception map and the original feature map instead of simply splicing or replacing, it retains the information richness of the original features while enhancing the expression ability of spatial positions. The use of layer normalization further ensures the stability and effectiveness of the fused features, making the model training more stable and converging faster.

[0047] In practical applications, the method of the present invention has good robustness and adaptability. For handwritten mathematical formulas of different styles, including unclear handwriting, complex structures, and symbol deformations, etc., it can maintain a high recognition accuracy. Especially for formulas with complex structures, such as multi-layer nested fractions, multiple integrals, matrices, etc., the advantages of the method of the present invention are more obvious, and the improvement in recognition accuracy is more significant.

[0048] In summary, the present invention proposes a method for recognizing handwritten mathematical formulas based on coverage attention and position perception. By explicitly extracting character position information through the position perception module and using the multi-scale coverage attention mechanism to optimize the attention calculation process, it effectively improves the recognition accuracy of handwritten mathematical formulas, especially showing obvious advantages when dealing with complex structures and nested relationships. While maintaining high accuracy, this method does not require additional annotation work, having high practical value and application prospects.

Claims

1. A handwritten mathematical formula recognition method based on coverage attention and position perception, characterized in that Including: S1. Construct a location perception module, generate a character location map through counting tags and weakly supervised learning, and enhance the feature extraction ability; S2. Construct a multi-scale coverage attention mechanism MARM to correct the interference of structural symbols in the sequence-based decoder architecture and accurately capture the attention of handwritten mathematical expression image recognition; S3. Fuse the original features with the location perception features to enhance the encoder's perception ability of spatial locations; S4. By simultaneously optimizing the two tasks of expression recognition and location recognition, explicitly realize the learning of location-aware symbolic feature representation and train based on this framework; S5. Input the image to be recognized into the trained location perception model PAMER to obtain the LaTeX expression of the mathematical formula.

2. The handwritten mathematical formula recognition method based on coverage attention and position perception according to claim 1, wherein The step S1 of constructing the location perception module specifically includes: S101. Multi-scale feature extraction: Use a multi-core convolution strategy, specifically using two different sizes of convolution kernels, 3×3 and 5×5, aiming to extract multi-scale and multi-resolution feature information from the input feature map, effectively capturing the details and overall structural features of handwritten mathematical symbols; S102. Feature enhancement: Introduce a channel attention mechanism SE block, and significantly improve the feature expression ability and discriminability by adaptively adjusting the channel weights, enhancing the model's sensitivity to handwritten mathematical symbols; S103. Location map generation: Use a fully connected layer to intelligently compress the number of channels to the number of symbol classes C, and accurately generate a location perception map with a value range in (0,1) through the Sigmoid function, realizing the probability modeling of the symbol spatial distribution; S104. Count vector calculation: Apply a sum pooling operation to the generated location perception map to obtain a highly concentrated count vector, providing accurate location statistical information and constraint conditions for the subsequent training stage.

3. The handwritten mathematical formula recognition method based on coverage attention and position perception according to claim 2, wherein The feature enhancement process of the SE block includes: S1021 Global average pooling: Obtain cross-channel global semantic information by performing average pooling on the spatial dimension of the feature map, providing a compressed representation for subsequent channel attention; S1022 Non-linear transformation: Use the first fully connected layer combined with the ReLU activation function to introduce non-linear transformation, enhancing the feature expression ability and non-linear modeling ability; S1023 Channel weight generation: Use the second fully connected layer and the Sigmoid function to adaptively generate normalized channel weight coefficients, realizing the accurate modeling of the importance of each channel; S1024 Feature recalibration: Dynamically apply the learned weights to the original features through the channel product method to achieve adaptive enhancement of the features and interaction between channels.

4. The handwritten mathematical formula recognition method based on coverage attention and position perception according to claim 1, characterized in that The specific implementation scheme of the feature fusion in step S3 is: S31 Multi-feature mapping: Linearly combine the original feature map output by the encoder and the location perception map according to a predetermined ratio, making full use of the spatial location information and the original feature representation; S32 Normalization processing: Perform layer normalization on the fused feature map to stabilize the feature scale, suppress internal covariate shift, and improve the generalization performance.

5. The handwritten mathematical formula recognition method based on coverage attention and position perception according to claim 1, characterized in that The multi-scale coverage attention mechanism in step S2 specifically includes: S21 multi-scale attention calculation: comprehensively using 3×3 and 5×5 convolution kernels to calculate the coverage attention refinement value from multiple angles and scales, enhancing the perception ability of the attention mechanism; S22 attention mechanism optimization: accurately embedding the calculated coverage attention refinement value into the attention mechanism of the decoder to reduce attention drift and defocusing problems; S23 dynamic attention weight: during the decoding process, dynamically calculating the attention weight of the image region according to the current decoding step to achieve adaptive regional attention; S24 attention debiasing: in the attention calculation of each decoding step, explicitly subtracting the attention weight of the recognized entity symbol in the previous step to avoid repeated attention and information redundancy.

6. The handwritten mathematical formula recognition method based on coverage attention and position perception according to claim 1, wherein Step S4 realizes multi-task learning by designing a comprehensive loss function: Loss function = λ1 * L expr + λ2 * L count Where: L expr represents the expression recognition loss, calculated as L count Represents the counting task loss, which measures the difference between the predicted counting vector and the true distribution; λ1 and λ2 are weight coefficients for balancing the importance of different tasks, and the optimal values are determined through cross-validation; Through the multi-task learning strategy, simultaneously optimize the accuracy of expression recognition and the position perception ability to achieve accurate recognition of handwritten mathematical formulas.

Citation Information

Patent Citations

  • Coding and decoding-based mathematical formula identification method and device, and readable storage medium

    CN114255379A

  • Handwritten mathematical formula identification method based on coding and decoding and self-attention model

    CN114926838A

  • Handwritten mathematical formula identification method based on multi-scale encoder network

    CN117496535A

  • Handwritten mathematical formula identification method

    CN117542064A

  • Deep learning-based test paper handwritten mathematical formula identification method and system

    CN117809320A

Cited By

  • Handwritten mathematical expression recognition method based on skeleton shaping and character counting module

    CN120524991A

  • Mathematical formula identification method and electronic equipment

    CN120673424A

  • Mathematical formula recognition method and electronic device

    CN120673424B

  • Handwritten mathematical formula identification method based on rotation position coding

    CN121904783A