Handwritten mathematical formula dynamic identification method based on multi-modal feature fusion
By adopting a multimodal feature fusion method in handwritten mathematical formula recognition and combining image and stroke features, the problem of low accuracy in confusing symbol recognition is solved, and higher recognition accuracy and user writing efficiency are achieved.
Patent Information
- Application Number
- CN202510578983.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Traditional handwritten mathematical formula recognition algorithms have low recognition accuracy when dealing with confusing symbols, especially when symbols such as plus signs and multiplication signs, letters and numbers with similar shapes and visual similarity, they cannot accurately distinguish these symbols.
The dynamic recognition method of handwritten mathematical formula based on multimodal feature fusion is adopted. Through a heterogeneous double-branch space-time fusion network, combined with image features and stroke features, deep fusion is carried out in the two dimensions of space and time, improving the recognition accuracy of confusing symbols.
Through the fusion of multimodal features, the recognition accuracy of confusing symbols is significantly improved, the similar symbols are effectively distinguished, the writing efficiency of users is improved, and the model's distinction ability is further improved through correction methods.
Smart Images

Figure CN120088800A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of handwritten mathematical formula recognition, and particularly to a dynamic recognition method for handwritten mathematical formulas based on multi-modal feature fusion. Background Art
[0002] In the field of handwritten mathematical formula recognition, traditional handwritten mathematical formula recognition algorithms usually face the problem of low recognition accuracy for easily confused symbols. Especially when dealing with symbols that are morphologically similar and visually close, such as plus and multiplication signs, letters and numbers, etc., the existing technologies often cannot accurately distinguish these symbols. The existing solutions mainly rely on single-modal information of images or strokes, but these methods often fail to fully utilize stroke order or context information when dealing with the ambiguity of handwritten symbols, resulting in misrecognition. Summary of the Invention
[0003] In order to solve the above technical problems existing in the prior art, the present invention proposes a dynamic recognition method for handwritten mathematical formulas based on multi-modal feature fusion, namely a heterogeneous dual-branch spatio-temporal fusion network for the recognition of easily confused symbols. This network combines image features and stroke features and performs deep fusion in both spatial and temporal dimensions, thereby improving the recognition accuracy of easily confused symbols. The specific technical solution is as follows: A dynamic recognition method for handwritten mathematical formulas based on multi-modal feature fusion, comprising: Step 1: Obtain handwritten mathematical formula data, encode and convert the stroke trajectory sequence information included in the handwritten mathematical formula data using the HSV color model to generate an HSV pseudo-color image, then fuse the HSV pseudo-color image with the binary image included in the handwritten mathematical formula data to obtain a temporally rich image, and then input the temporally rich image into a convolutional neural network; Step 2: Through the convolutional neural network, extract features from the temporally rich image and the stroke trajectory sequence included in the mathematical formula data respectively to obtain image features and stroke features, use the image features to enhance the stroke features, and then fuse the image features with the enhanced stroke features to obtain multi-modal features; Step 3: Through a decoder, enhance the input multi-modal features using positional encoding, and then combine an attention mechanism and a multi-scale symbol counting module to perform symbol prediction to generate a mathematical formula symbol sequence; Step 4: Design rules to correct the generated mathematical formula symbol sequence.
[0004] Further, in step 1, the stroke trajectory sequence information contained in the handwritten mathematical formula data is encoded and transformed using the HSV color model. Specifically, the writing order, stroke length, and formula complexity of each stroke are respectively mapped to the Hue, Saturation, and Value of the HSV model to generate a pseudo-color image, that is, the writing order of the stroke is mapped to the Hue value, the stroke length is associated with the Saturation, and after normalizing the number of strokes, it is combined with the Value to generate an image with temporal information.
[0005] Further, in step 2, the DenseNet-BC network is used to extract features from the temporally rich image to obtain image features, and the selective state space model mamba is used to process the stroke trajectory sequence to obtain stroke features.
[0006] Further, in step 2, the image features and stroke features are combined. First, a prior alignment relationship is established through the coordinate sequences of the image and the stroke, and then through feature transformation and reset operations, the expressive ability of the stroke features is enhanced.
[0007] Further, in step 2, when fusing the image features with the enhanced stroke features, first, a one-dimensional convolutional layer is used to align the image and stroke features to ensure that the feature sizes of the two modalities are the same, and then the Concat operation is used to splice the image features and stroke features in the channel dimension, thereby increasing the number of features.
[0008] Further, in step 3, the decoder performs a dimensional transformation on the input multi-modal features through a 1×1 convolutional layer, and then adds a positional encoding for feature enhancement.
[0009] Further, the multi-scale symbol counting module performs multi-scale symbol counting processing on the image features to predict various symbol count values, that is, the statistical quantities of various symbols from a global perspective.
[0010] Further, the attention mechanism selects and weights feature maps of different scales, and generates a context vector through the fusion of low-resolution and high-resolution features.
[0011] Further, the decoder calculates the probability distribution of the current symbol based on the current hidden layer state, context vector, various symbol count values, and the prediction result of the previous symbol, selects the symbol with the highest probability through the probability distribution, and outputs it, gradually generating a complete mathematical formula symbol sequence. The previous hidden layer state is obtained by taking the previous symbol and inputting its feature encoding into the decoder GRU layer.
[0012] In step 4, according to the probability distribution of symbol recognition, combined with the statistical results and the confusion matrix, the symbol sequence of the generated mathematical formula is corrected through LaTeX syntax rules.
[0013] Beneficial effects: The method of the present invention improves the accuracy of identifying easily confused characters by combining image features and stroke features, enhances symbol category guidance using context information, enriches the image representation for initial alignment using temporal information, and introduces enhanced stroke features for detailed differentiation, which can effectively distinguish similar symbols written by the re-identified user and improve the recognition accuracy; moreover, a method for correcting similar symbols is proposed for the statistical probability of the recognition test results, and output rules are set for differentiation to enhance the model's ability to distinguish such symbols and improve the user's writing efficiency. Description of the Drawings
[0014] Figure 1 is the network architecture diagram of the dynamic recognition method for handwritten mathematical formulas based on multi-modal feature fusion according to the embodiment of the present invention; Figure 2 is the effect diagram of the temporally rich image according to the embodiment of the present invention; Figure 3 is the schematic diagram of the network structure with enhanced stroke features according to the embodiment of the present invention; Figure 4 is the schematic diagram of the decoder combined with the multi-scale symbol counting module according to the embodiment of the present invention; Figure 5 is the schematic diagram of the multi-scale symbol counting module according to the embodiment of the present invention. Detailed Embodiments
[0015] In order to make the objectives, technical solutions, and technical effects of the present invention clearer and more understandable, the following further elaborates on the present invention in detail in conjunction with the drawings in the specification and embodiments.
[0016] As Figure 1 shown, this embodiment discloses a dynamic recognition method for handwritten mathematical formulas based on multi-modal feature fusion, including: Step 1, as Figure 2 shown, obtain handwritten mathematical formula data, encode and convert the stroke trajectory sequence information in the data using the HSV color model to generate an HSV pseudo-color image with temporal information for the formula strokes, fuse the HSV pseudo-color image with the binary image in the data to obtain a temporally rich image, and then input the temporally rich image into a convolutional neural network for feature processing.
[0017] Specifically, each handwritten mathematical formula data includes an image and corresponding stroke trajectory sequence information. The image of the handwritten formula is generally a binary image, and the stroke sequence is composed of point coordinates and stroke pause flags.
[0018] Encoding and converting the stroke trajectory sequence information in the data using the HSV color model specifically involves mapping the writing order, stroke length, and formula complexity of each stroke to the Hue, Saturation, and Value of the HSV model respectively to generate a pseudo-color image, that is, mapping the writing order of the stroke to the Hue value, associating the stroke length with Saturation, and combining the normalized number of strokes with Value to generate an image with temporal information.
[0019] The generated HSV pseudo-color image is fused with the original binary image as part of the input image to ensure that the convolutional neural network can process both the spatial features of the image and the temporal features of the strokes.
[0020] Step 2, as Figure 3 shown, through the convolutional neural network, feature extraction is performed on the temporally rich image and the stroke trajectory sequence respectively to obtain image features and stroke features, and the image features are used to enhance the stroke features. Then, the image features are fused with the enhanced stroke features to obtain multi-modal features.
[0021] Specifically, the DenseNet-BC network is used to extract features from the temporally rich image to obtain image features. The DenseNet-BC network, through dense connections and Bottleneck layer design, can retain low-level detail information and obtain high-level semantic information when extracting image features.
[0022] The selective state space model mamba is used to process the stroke trajectory sequence to obtain stroke features. The selective state space model captures the long temporal dependence features of the strokes through dynamic modeling of hidden states and is suitable for processing local dynamics and global associations in the stroke writing order.
[0023] In the initial alignment stage of the image and the stroke sequence, a prior alignment relationship is established through the coordinate sequences of the image and the strokes to help the network capture cross-modal dependence relationships.
[0024] The feature map of the temporally rich image extracted by the convolutional neural network is combined with the stroke features extracted by the selective state space model. Through feature transformation and reset operations, the expressive ability of the stroke features is enhanced to ensure that the network can better capture detail information when processing handwritten formulas.
[0025] Regarding the fusion of the image features and the enhanced stroke features, first, a one-dimensional convolutional layer is used to align the image and stroke features to ensure that the feature sizes of the two modalities are the same. Then, the Concat operation is used to splice the image features and the stroke features in the channel dimension, thereby increasing the number of features and enabling the model to learn more feature representations of the same formula.
[0026] By introducing time-series rich images, enriching the images using the writing order of strokes and the contextual relationships between strokes, it helps the convolutional neural network better process the dynamic information in the images. Combining the spatial information of the images and the time-series information of the strokes for preliminary alignment, establishing a priori alignment relationships, helps the network capture cross-modal dependencies, and significantly improves the recognition ability of easily confused symbols.
[0027] Step 3, as Figure 4 shown, input the multi-modal features into the decoder, enhance them through positional encoding, and then combine the attention mechanism and the multi-scale symbol counting module for symbol prediction to generate a mathematical formula symbol sequence.
[0028] Specifically, the present invention applies a decoder integrated with a multi-scale symbol counting module to map the feature information of images and strokes into the final handwritten mathematical formula symbols. The specific process includes: Input feature processing: Perform dimensional transformation on the input multi-modal features through a 1×1 convolutional layer to convert them into intermediate layer features, ensuring that the number of channels of each feature map is the same.
[0029] Positional encoding enhancement: In order to better capture the spatial position relationships of symbols, the decoder adds positional encoding to the input features to enhance the input feature maps, providing the spatial position information of symbols to distinguish symbols at different spatial positions, that is, enhancing the decoder's discrimination ability for each position in the image. This is particularly important in handwritten mathematical formulas. For example, in fractional or radical symbols, the hierarchical relationships of the symbols are very important, and positional encoding can help the decoder recognize these relationships.
[0030] Symbol counting perception: As Figure 5 shown, perform multi-scale symbol counting processing on the image features to predict the count values of various symbols, that is, the statistics of various symbols.
[0031] To solve the problem of symbol recognition in complex formulas, the decoder of the present invention introduces a multi-scale symbol counting module. Through the multi-scale symbol counting module, combining local symbol features with global symbol category statistical information can provide more accurate context information for each symbol prediction, help recognize complex formulas, and effectively reduce the situation of repeated recognition or omission of complex formula symbols. This module determines the meaning and position of symbols according to the context information at different scales, enabling the decoding process to better handle the recognition problems of complex symbols and long formulas.
[0032] Attention weighting: The attention mechanism is used to select and weight feature maps at different scales. By fusing low-resolution and high-resolution features, a fine-grained context vector is generated for accurate prediction of the next symbol.
[0033] During the decoding process, attention weights are calculated based on the context information of the current symbol, and different focus areas are assigned according to the weights to ensure that the decoder focuses on the regions most important for predicting the current symbol. This mechanism enables the decoder to focus on local features crucial for symbol recognition.
[0034] Symbol sequence decoding output: At each time step, the generated context vector is input into a series of GRU layers. These layers calculate the probability distribution of the current symbol based on the current hidden layer state, context vector, various symbol count values, and the prediction result of the previous symbol. The symbol with the highest probability is selected through the probability distribution and output, gradually generating a complete mathematical formula symbol sequence. The previous hidden layer state is obtained by taking the feature encoding of the previous symbol and inputting it into the GRU layer.
[0035] The decoder in this embodiment not only focuses on local features but also further adjusts the prediction result of each symbol through the count vector of the global context, thereby reducing symbol repetition and misrecognition.
[0036] Step 4: Correct the generated mathematical formula symbol sequence in combination with LaTeX syntax rules, as follows.
[0037] Rule-based correction method: After symbol recognition, a rule-based correction method is designed to correct common confusing symbols, further improving the model's performance in complex formulas. For example, for misrecognition between "0" and "O", the rules are set as follows: If the string contains both letters and the ◯ symbol, replace all ◯ with the letter "O". If the string contains only the ◯ symbol, replace it with the number "0".
[0038] Probability distribution correction method: According to the probability distribution of symbol recognition, combined with statistical results and confusion matrices, the recognition results are corrected through LaTeX syntax rules to improve the discrimination of confusing symbols.
[0039] LaTeX syntax rule correction: According to LaTeX syntax rules, post-process the recognition results to ensure the correct output of symbols. For example, to handle the misrecognition of parentheses and the number "1", check if the parentheses are closed. If the left parenthesis is not closed, replace it with "1".
[0040] Symbol output: The corrected symbol results are output, and finally, a mathematical formula text that conforms to LaTeX syntax rules is obtained. At this time, the recognition results contain all correct symbols and meet the syntax requirements, suitable for further formula parsing and calculation.
[0041] Result verification and display: Verify the recognition results, especially evaluate the discrimination accuracy of easily confused symbols, and ensure that the algorithm has a high recognition accuracy in different scenarios.
[0042] The above is only the preferred embodiment of the present invention and does not impose any form of limitation on the present invention. Although the implementation process of the present invention has been described in detail above, those familiar with the art can still modify the technical solutions recorded in the foregoing examples or make equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for dynamic recognition of handwritten mathematical formulas based on multimodal feature fusion, characterized in that: include: Step 1, obtaining handwritten mathematical formula data, using the HSV color model to encode and convert the stroke trajectory sequence information contained in the handwritten mathematical formula data, generating an HSV pseudo-color image, and then fusing the HSV pseudo-color image with the binary image contained in the handwritten mathematical formula data to obtain a time-series enriched image, and then inputting the time-series enriched image into a convolutional neural network; Step 2: Through the convolutional neural network, feature extraction is performed on the time-series rich image and the stroke trajectory sequence to obtain image features and stroke features, and the stroke features are enhanced by using the image features, and then the image features are fused with the enhanced stroke features to obtain multimodal features; Step 3: Through the decoder, position encoding is used to enhance the multimodal features of the input, and then the attention mechanism and multi-scale symbol counting module are combined to perform symbol prediction to generate a mathematical formula symbol sequence; Step 4: Design rules to correct the generated mathematical formula symbol sequence.
2. The method for dynamic recognition of handwritten mathematical formulas according to claim 1, characterized in that: In step 1, the stroke trajectory sequence information contained in the handwritten mathematical formula data is encoded and converted using the HSV color model, specifically: the writing order of each stroke, the stroke length and the formula complexity are respectively mapped to the hue, saturation and value of the HSV model to generate a pseudo-color image, that is, the writing order of the strokes is mapped to the hue value, the stroke length is associated with the saturation, the number of strokes is normalized and combined with the value to generate an image with time series information.
3. The method for dynamic recognition of handwritten mathematical formulas according to claim 1, characterized in that: In step 2, a DenseNet-BC network is used to extract features from the temporally rich image to obtain image features, and a selective state space model mamba is used to process the stroke trajectory sequence to obtain stroke features.
4. The method for dynamic recognition of handwritten mathematical formulas according to claim 1, characterized in that: In step 2, the image features and stroke features are combined, and a priori alignment relationship is established through the coordinate sequence of the image and the strokes. Then, the expressiveness of the stroke features is enhanced through feature conversion and resetting operations.
5. The method for dynamic recognition of handwritten mathematical formulas according to claim 1, characterized in that: In step 2, the image features are fused with the enhanced stroke features. First, a one-dimensional convolutional layer is used to align the image and stroke features to ensure that the feature sizes of the two modalities are consistent. Then, a Concat operation is used to concatenate the image features and stroke features in the channel dimension, thereby increasing the number of features.
6. The method for dynamic recognition of handwritten mathematical formulas according to claim 1, characterized in that: In step 3, the decoder transforms the dimension of the input multimodal features through a 1×1 convolutional layer, and then adds position encoding for feature enhancement.
7. The method for dynamic recognition of handwritten mathematical formulas according to claim 6, characterized in that: The multi-scale symbol counting module performs multi-scale symbol counting processing on the image features to predict various symbol count values, that is, statistics for various symbols from a global perspective.
8. The method for dynamic recognition of handwritten mathematical formulas according to claim 7, characterized in that: The attention mechanism selects and weights feature maps of different scales, and generates a context vector by fusing low-resolution and high-resolution features.
9. The method for dynamic recognition of handwritten mathematical formulas according to claim 8, characterized in that: The decoder calculates the probability distribution of the current symbol based on the current hidden layer state, context vector, various symbol count values and the prediction result of the previous symbol, selects the symbol with the highest probability through the probability distribution, and outputs it, gradually generating a complete mathematical formula symbol sequence. The previous hidden layer state uses the previous symbol, takes its feature code and inputs it into the decoder GRU layer to obtain it.
10. The method for dynamic recognition of handwritten mathematical formulas according to claim 8, characterized in that: In step 4, according to the probability distribution of symbol recognition, combined with the statistical results and the confusion matrix, the generated mathematical formula symbol sequence is corrected through LaTeX grammar rules.
Citation Information
Patent Citations
Timber pile information acquisition method, system and device based on neural network
CN109977960A
Handwritten mathematical expression recognition method and device based on deep learning
CN110766012A
Multi-modal video Chinese subtitle recognition method based on dense connection convolutional network
CN113221900A
Handwriting recognition method and device, electronic equipment and storage medium
CN115984877A
Transform network-based base station optimization site selection method and system
CN118488465A
Cited By
Multimodal OCR (Optical Character Recognition) method and system for handwritten compositions of pupils
CN121617103A
Digital twinborn regulation and control method for facility poultry breeding environment based on deep fusion of spatio-temporal data
CN121934400A
Facility poultry breeding environment digital twin regulation method based on spatiotemporal data deep fusion
CN121934400B