Bayesian prior handwritten mathematical expression recognition method based on context consistency
Through the Bayesian prior multi-source input fusion recognition method based on context consistent Bayesian priors, the features of handwritten mathematical expression pictures and context information are extracted, and the problem of low recognition accuracy in the prior art is solved, achieving higher recognition accuracy.
Patent Information
- Application Number
- CN202510112704.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
Existing handwriting mathematical expression recognition methods are difficult to improve recognition accuracy while ensuring correct grammar, especially in the face of unstable handwriting quality, diversity of Latex sequence expression results and irregular writing.
A Bayesian prior multi-source input fusion recognition method is adopted based on context consistent Bayesian priors. By collecting handwritten mathematical expression pictures and corresponding context information, a multi-source input fusion recognition model is trained, mathematical formula picture features, context information features and Latex fill-in-the-blank target sequence features are extracted, and feature fusion processing is performed to output recognition results.
It effectively reduces the syntax errors in predicting Latex sequences and the uncertainty of the text to be recognized, improves the recognition accuracy of handwritten mathematical expressions, and can provide high application value in test papers or homework correction scenarios.
Smart Images

Figure CN120047962A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a Bayesian prior handwritten mathematical expression recognition method based on context consistency. Background Art
[0002] In the field of artificial intelligence, handwritten mathematical expression recognition (HMER) is a technology that converts handwritten mathematical expression images into text sequences of expressions conforming to the Latex syntax structure. Currently, it is widely used in human-computer interaction scenarios such as online education, manuscript digitization, and automatic grading. However, accurately recognizing handwritten mathematical expressions is a challenging task, which not only requires the ability to recognize individual characters, but also needs to understand the structural relationships between characters to accurately parse the entire mathematical expression.
[0003] In the scenario of test paper or homework grading, the writing quality of a large number of handwritten texts is relatively poor, the expression is not precise enough, and the composite function operation relationships between characters are not clearly written. Therefore, the task of recognizing handwritten mathematical expressions not only needs to deal with the instability of handwritten quality, but also needs to solve the diversity of Latex sequence expression results and the ambiguity that is difficult to eliminate in the representation of mathematical Latex sequences that conform to grammar. Existing handwritten mathematical expression recognition methods cannot reasonably model the grammar of complex mathematical expressions, and it is difficult to ensure the recognition accuracy when comparing with the correct answer while ensuring the grammar is correct. It is easy to cause problems such as syntax errors in predicting Latex sequences and the uncertainty of the text to be recognized (such as the numbers 6 / 9 / 0 and the English characters b / q / o, the number 2 and the Chinese character B, the Greek letter omega and the English letter w), as well as ambiguities caused by non-standard writing (the mutual mathematical operation relationships between characters in trigonometric functions, power functions, and exponential functions are not clear). Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a Bayesian prior handwritten mathematical expression recognition method based on context consistency, which can effectively improve the recognition accuracy of handwritten mathematical expressions.
[0005] The purpose of the present invention can be achieved by the following technical solutions: A Bayesian prior handwritten mathematical expression recognition method based on context consistency, comprising the following steps:
[0006] S1. Collect handwritten mathematical expression pictures, construct a data set, and prepare context information corresponding to each handwritten mathematical expression picture. The data set includes multiple handwritten mathematical expression pictures and corresponding Latex sequence texts;
[0007] S2. Using the dataset and the prepared context information, train a multi-source input fusion recognition model. The multi-source input fusion recognition model is used to extract the feature representations of the mathematical formula picture, the context information, and the Latex filling target sequence, and through feature fusion processing, output the recognition result;
[0008] S3. Input the handwritten mathematical expression picture to be recognized, the corresponding context information, and the Latex filling text sequence into the multi-source input fusion recognition model, and output the corresponding recognition result, including the Latex sequence text and the confidence level.
[0009] Further, the Latex sequence texts corresponding to the multiple handwritten mathematical expression pictures in step S1 are all subjected to normalization processing. The normalization processing process includes:
[0010] Use spaces to separate the keywords, mathematical operation symbols, independent numerical variables, and constants in Latex with spaces;
[0011] Only use a pair of curly braces {} to separate the sub-formulas in the composite function, and remove redundant {};
[0012] The context information corresponding to each handwritten mathematical expression picture in step S1 is used to provide reference information for the content of the handwritten mathematical expression, specifically including: relevant materials corresponding to the mathematical expression, the whole or part of the problem text.
[0013] Further, the multi-source input fusion recognition model in step S2 includes a first module, a second module, a third module, and a fourth module. The first module, the second module, and the third module are respectively connected to the fourth module. The first module is used to extract the feature representation of the mathematical formula picture;
[0014] The second module is used to extract the feature representation of the context information;
[0015] The third module is used to extract the feature representation of the Latex filling target sequence;
[0016] The fourth module is used to fuse the feature representations output by the first module, the second module, and the third module to determine the recognition result.
[0017] Further, step S2 adopts a data perturbation training method for model training. Specifically, during the training process, the context information and the Latex filling text sequence are randomly set as ineffective "dummy" variables or 0 variables, so that the model can adapt to the missing situation.
[0018] Further, the first module specifically adopts a Densenet, VGG, Resnet or Transformer feature extraction network structure in deep learning.
[0019] Further, the second module specifically adopts a VGG, Densenet, Resnet or Transformer image feature extraction network structure in deep learning. The working process of the second module is as follows: Print the text sequence of the context in printed form on an image sheet, and then perform image feature extraction on the image sheet to obtain a feature representation of the context information.
[0020] Further, the second module specifically adopts an RNN / LSTM, Bert, GPT, CLIP or Transformer deep learning model. The working process of the second module is as follows: Perform feature extraction on the text sequence of the context in the manner of Natural Language Processing in natural language processing to obtain a feature representation of the context information.
[0021] Further, the Latex fill-in-the-blank target sequence is specifically constructed by replacing the key characters with blank symbols according to the Latex sequence label corresponding to the picture, so as to construct a Latex fill-in-the-blank text sequence.
[0022] Further, the third module specifically adopts a Densenet, VGG, Resnet or Transformer feature extraction model. The working process of the third module is as follows: Print the Latex fill-in-the-blank text sequence label on an image sheet, and then perform feature extraction on the image sheet by using the deep learning image feature extraction method.
[0023] Further, the third module specifically adopts an RNN / LSTM, Bert, GPT, CLIP or Transformer deep model. The working process of the third module is as follows: Use a deep learning network model of Natural Language Processing to perform feature extraction on the Latex fill-in-the-blank text sequence label.
[0024] Compared with the prior art, the present invention has the following advantages:
[0025] The present invention collects handwritten mathematical expression pictures, constructs a data set (including multiple handwritten mathematical expression pictures and corresponding LaTeX sequence texts), and prepares context information corresponding to each handwritten mathematical expression picture. Then, using the data set and the prepared context information, a multi-source input fusion recognition model is trained. This multi-source input fusion recognition model is used to extract the feature representations of the mathematical formula picture, the context information feature representation, and the feature representation of the LaTeX filling target sequence, and through feature fusion processing, an identification result including the LaTeX sequence text and the confidence level is output. Thus, by using additional multimodal information such as text or images associated with the answer, effective context and LaTeX grammar guidance are formed for the recognition result, thereby effectively reducing problems such as syntax errors in predicting the LaTeX sequence, the uncertainty of the text to be recognized, and the ambiguity caused by non-standard writing, and greatly improving the recognition accuracy.
[0026] The present invention prepares context information matching the mathematical expression picture data, including relevant materials corresponding to the mathematical expression, or the whole or part of the question text (such as the LaTeX sequence of the mathematical expression in the question stem), which can provide reference information for the content of the handwritten mathematical expression for the recognition model, thereby suppressing ambiguous text recognition results.
[0027] The present invention normalizes the text of the LaTeX sequence label to have the same expression style, which can solve the problem of the diversity of LaTeX sequence expression results and is beneficial to improving the efficiency and accuracy of subsequent model training.
[0028] The multi-source input fusion recognition model constructed by the present invention includes a first module to a fourth module. Among them, the first module, the second module, and the third module are respectively used to extract the feature representation of the mathematical formula picture, the context information feature representation, and the feature representation of the LaTeX filling target sequence. The fourth module then performs fusion processing on the extracted features, thereby achieving the effect of multi-modal feature information fusion recognition and ensuring the recognition accuracy.
[0029] When training the multi-source input fusion recognition model, the present invention adopts a data perturbation training method, that is, randomly setting the context information and the LaTeX filling text sequence as ineffective "dummy" variables or 0 variables during the training process, so that the model can adapt to the missing situation. While ensuring that the training model maintains a low false positive rate, it significantly improves the recognition accuracy and has high application value in the artificial intelligence marking scenarios of test papers or homework. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a schematic flowchart of the method of the present invention;
[0031] Figure 2 and Figure 3Schematic diagram of the application process in the embodiment. Detailed implementation manners
[0032] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] Embodiment
[0034] As Figure 1 shown, a Bayesian prior handwritten mathematical expression recognition method based on context consistency includes the following steps:
[0035] S1. Collect handwritten mathematical expression pictures, construct a data set, and prepare context information corresponding to each handwritten mathematical expression picture. The data set includes multiple handwritten mathematical expression pictures and corresponding Latex sequence texts;
[0036] S2. Use the data set and the prepared context information to train a multi-source input fusion recognition model. The multi-source input fusion recognition model is used to extract feature representations of mathematical formula pictures, context information feature representations, and feature representations of Latex filling target sequences, and output recognition results through feature fusion processing;
[0037] S3. Input the handwritten mathematical expression picture to be recognized, the corresponding context information, and the Latex filling text sequence into the multi-source input fusion recognition model, and output the corresponding recognition results, including the Latex sequence text and the confidence level.
[0038] The main contents of this embodiment applying the above solution are as follows:
[0039] The first step: Collect and obtain handwritten mathematical expression pictures to establish a data set for model training. Use the Latex sequence text corresponding to the mathematical expression content in the handwritten mathematical expression pictures as data labels, and prepare context information matching the above-mentioned mathematical expression picture data.
[0040] Among them, the content of the picture data includes handwritten mathematical expression information;
[0041] In addition, the text of the Latex sequence label corresponding to the picture data is also normalized to have the same expression style. The specific normalization of the Latex sequence text includes: 1) Separate the keywords, mathematical operation symbols, independent numerical variables, and constants in Latex with spaces; 2) Only use a pair of curly braces {} to separate the sub-formulas in the composite function, and remove redundant {}.
[0042] For the context information matched by handwritten mathematical expressions, in the context of test paper or homework grading, it mainly refers to the relevant materials corresponding to the mathematical expressions to be recognized, or the whole or part of the question text (including but not limited to the Latex sequence of the mathematical expressions in the question stem). The context information can provide reference information for the recognition model about the content of the handwritten mathematical expressions to be recognized and suppress ambiguous text recognition results.
[0043] The second step: As Figure 2 shown, use Module 1 and Module 2 to extract the feature representation I of the mathematical formula image and the feature representation C of the context information respectively.
[0044] Among them, Module 1 can be any image feature extraction method, including but not limited to feature extraction network structures such as Densenet, VGG, Resnet, and Transformer in deep learning. The model outputs the image feature I of the mathematical expression to be recognized.
[0045] For Module 2, an optional solution is to print the text sequence of the context in printed form on an image sheet, and then perform image feature extraction on the image sheet, including but not limited to image feature extraction network structures such as VGG, Densenet, Resnet, and Transformer in deep learning. The feature representation of the extracted context information is denoted as C.
[0046] For Module 2, another optional solution is to extract features from the text sequence of the context using the method of Natural Language Processing (NLP), including but not limited to NLP deep learning models such as RNN / LSTM, Bert, GPT, CLIP, and Transformer. The feature representation of the extracted context information is denoted as C.
[0047] The third step: According to the Latex sequence label corresponding to the picture, replace the key characters with blank symbols to construct a Latex fill-in-the-blank text sequence, and use Module 3 to extract the feature representation L of the Latex fill-in-the-blank target sequence;
[0048] Among them, the LaTeX fill-in-the-blank text sequence is obtained by replacing the key characters in the LaTeX sequence tags corresponding to the mathematical expressions. For example, mathematical variables / constants are replaced with extra spaces or "#", constructing the LaTeX fill-in-the-blank text sequence. For instance, the LaTeX sequence tag "\frac{\sqrt{3}}{2}\neq\frac{1}{2}" → LaTeX fill-in-the-blank text sequence "\frac{\{}}{}\\frac{}{}" or "\frac{\#{#}}{#}\#\frac{#}{#}", forming the fill-in-the-blank text sequence to be completed. The blank replacement symbols are not limited to spaces or "#", as long as they are symbols not commonly seen in test papers / homework answers.
[0049] Module 3, an alternative solution is to print the LaTeX fill-in-the-blank text sequence tags on image chips, and then use deep learning image feature extraction methods to extract features from the image chips, including but not limited to feature extraction models such as Densenet, VGG, Resnet, Transformer, etc. The feature representation of the LaTeX fill-in-the-blank text sequence extracted is denoted as L.
[0050] Module 3, another alternative solution is to use a deep learning network model for natural language processing to extract features from the LaTeX fill-in-the-blank text sequence tags, including but not limited to NLP deep models such as RNN / LSTM, Bert, GPT, CLIP, Transformer, etc. The feature representation of the LaTeX fill-in-the-blank text sequence extracted is denoted as L.
[0051] Step 4: Fuse the mathematical expression image features I, context information features F, and the feature representation L of the LaTeX fill-in-the-blank text sequence output by Modules 1, 2, and 3 using Module 4 and output the recognition result, where the recognition result includes the predicted LaTeX sequence text and the confidence level.
[0052] Step 5: In practical applications, as Figure 3 shown, Modules 1, 2, 3, and 4 can be an end-to-end recognition model for multi-input fusion, that is, using the description of 3 inputs in total and 1 output as a whole, rather than independently trained modules. Among them, it includes Input 1 - the mathematical expression image, Input 2 - the context information, and Input 3 - the LaTeX fill-in-the-blank text sequence, and outputs the recognition result, including the LaTeX sequence text and the confidence level. It should be noted that the image of the mathematical expression is a necessary input, while the context information and the LaTeX fill-in-the-blank text sequence are optional inputs.
[0053] Step 6: The image of the mathematical expression is a necessary input for this solution, while the context information and the LaTeX fill-in-the-blank text sequence are optional inputs. That is to say, this solution can also give reasonable inference results when only the image of the mathematical expression is input and the context information and the LaTeX fill-in-the-blank text sequence are missing. The role of the context information and the LaTeX fill-in-the-blank text sequence is to significantly improve the recognition effect. To achieve the recognition effect of this optional input, this solution adopts a data perturbation training method, randomly setting the context information and the LaTeX fill-in-the-blank text sequence as ineffective "dummy" variables or 0 variables during the training process, so that the model can adapt to the missing situation.
Claims
1. A context-consistent Bayesian prior handwritten mathematical expression recognition method, characterized in that: The following steps are involved: S1. Collect handwritten mathematical expression images, construct a data set, and prepare context information corresponding to each handwritten mathematical expression image, wherein the data set includes multiple handwritten mathematical expression images and corresponding Latex sequence texts; S2. Using the data set and the prepared context information, a multi-source input fusion recognition model is trained to extract feature representations of mathematical formula images, feature representations of context information, and feature representations of Latex fill-in-the-blank target sequences, and outputs recognition results through feature fusion processing; S3. Input the handwritten mathematical expression image to be recognized and the corresponding context information and the Latex fill-in-the-blank text sequence into the multi-source input fusion recognition model, and output the corresponding recognition result, including the Latex sequence text and confidence.
2. A context-consistent Bayesian prior handwritten mathematical expression recognition method according to claim 1, characterized in that: In step S1, the Latex sequence texts corresponding to the multiple handwritten mathematical expression images are all normalized, and the normalization process includes: Use spaces to separate Latex keywords, mathematical operators, independent numeric variables and constants; Only a pair of curly braces {} can be used to separate sub-formulas in a composite function, and redundant {} should be removed; The context information corresponding to each handwritten mathematical expression image in step S1 is used to provide reference information of the content of the handwritten mathematical expression, specifically including: relevant materials corresponding to the mathematical expression, and the whole or part of the question text.
3. The context-consistent Bayesian prior handwritten mathematical expression recognition method according to claim 1, characterized in that: The multi-source input fusion recognition model in step S2 includes a first module, a second module, a third module and a fourth module, wherein the first module, the second module and the third module are respectively connected to the fourth module, and the first module is used to extract the feature representation of the mathematical formula picture; The second module is used to extract context information feature representation; The third module is used to extract the feature representation of the Latex fill-in-the-blank target sequence; The fourth module is used to fuse the feature representations output by the first module, the second module, and the third module to determine the recognition result.
4. The context-consistent Bayesian prior handwritten mathematical expression recognition method according to claim 1, characterized in that: The step S2 adopts data perturbation training to perform model training, specifically, during the training process, context information and Latex fill-in-the-blank text sequences are randomly set to ineffective "dummy" variables or 0 variables, so that the model can adapt to the missing situation.
5. The context-consistent Bayesian prior handwritten mathematical expression recognition method according to claim 3, characterized in that: The first module specifically adopts the Densenet, VGG, Resnet or Transformer feature extraction network structure in deep learning.
6. The context-consistent Bayesian prior handwritten mathematical expression recognition method according to claim 3, characterized in that: The second module specifically adopts the VGG, Densenet, Resnet or Transformer image feature extraction network structure in deep learning. The working process of the second module is: printing the context text sequence on the image slice in print, and then performing image feature extraction on the image slice to obtain the feature representation of the context information.
7. The context-consistent Bayesian prior handwritten mathematical expression recognition method according to claim 3, characterized in that: The second module specifically adopts RNN / LSTM, Bert, GPT, CLIP or Transformer deep learning model. The working process of the second module is: the context text sequence is subjected to feature extraction by natural language processing to obtain feature representation of context information.
8. The context-consistent Bayesian prior handwritten mathematical expression recognition method according to claim 1, characterized in that: The Latex fill-in-the-blank target sequence specifically replaces key characters with blank symbols based on the Latex sequence tags corresponding to the images, thereby constructing a Latex fill-in-the-blank text sequence.
9. The context-consistent Bayesian prior handwritten mathematical expression recognition method according to claim 3, characterized in that: The third module specifically adopts Densenet, VGG, Resnet or Transformer feature extraction model, and the working process of the third module is: printing the Latex fill-in-the-blank text sequence label on the image slice, and then using the deep learning image feature extraction method to extract features from the image slice.
10. The context-consistent Bayesian prior handwritten mathematical expression recognition method according to claim 3, characterized in that: The third module specifically adopts RNN / LSTM, Bert, GPT, CLIP or Transformer deep model, and the working process of the third module is: using the deep learning network model of Natural Language Processing to extract features of Latex fill-in-the-blank text sequence labels.