Handwritten mathematical expression recognition and comparison method based on large model reasoning
Through the handwritten mathematical expression recognition and comparison method based on large model inference, the large language model is fine-tuned in combination with the output results and input data of the OCR recognition model, and the context information and latex sequence format information are used to solve the complexity and accuracy of handwritten mathematical expression recognition in the prior art, achieving higher recognition accuracy and effectiveness of heterogeneous answer comparison.
Patent Information
- Application Number
- CN202510112702.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing math test paper review/homework correction system faces problems such as complex structure, diversity and nesting when recognizing handwritten mathematical expressions, which leads to high difficulty in recognition and low accuracy.
The handwritten mathematical expression recognition and comparison method based on large model reasoning is adopted. The OCR recognition model is pre-trained, and the large language model is fine-tuned in combination with the output results and input data of the OCR recognition model, and context information and latex sequence format information are used to disambiguate and constrain the recognition results.
It improves the recognition accuracy of handwritten mathematical expressions and effectively solves the problem of heterogeneous answer comparison. By providing additional context information and latex sequence format information, syntax errors and ambiguity problems are reduced.
Smart Images

Figure CN120047961A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and in particular to a method for recognizing and comparing handwritten mathematical expressions based on large model reasoning. Background Technique
[0002] Existing mathematical test paper marking / homework correction systems generally use OCR (Optical Character Recognition) recognition to recognize handwritten mathematical expressions as Latex formula sequences. By recognizing the handwritten mathematical expressions in the image as Latex character sequences, and then comparing the recognition results with the Latex character sequences in the standard answers to determine whether the answers are correct.
[0003] However, in practical applications, due to the diverse and complex forms of mathematical expressions, their structures will pose great challenges in the recognition process: the existence of two-dimensional structures, the diversity of formula formats, the complexity of formula structures, especially the nesting of various structures, difficult-to-distinguish similar-shaped characters (such as "Z" and "2", "O" and "0"), the combination of various characters (non-conventional symbols, letters, numbers, operators), etc. all make it difficult to parse the formula in recognition. Handwritten mathematical expressions are even more prone to irregular writing (such as Summary of the Invention
[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method for recognizing and comparing handwritten mathematical expressions based on large model reasoning, which can improve the accuracy of heterogeneous answer comparison while accurately recognizing handwritten mathematical expressions.
[0005] The purpose of the present invention can be achieved by the following technical solutions: A method for recognizing and comparing handwritten mathematical expressions based on large model reasoning, including the following steps:
[0006] S1. Pre-train an OCR recognition model for handwritten mathematical expressions;
[0007] S2. Based on the output result of the OCR recognition model, fine-tune the large language model (Large Language Model, LLM) in combination with the input data of the OCR recognition model;
[0008] S3. Input the latex text sequence corresponding to the standard answer into the fine-tuned large language model, and output the recognition and comparison result.
[0009] Furthermore, the input data of the OCR recognition model includes the handwritten mathematical expression image to be recognized and optional input information, and the optional input information includes context text information and latex sequence format information.
[0010] Furthermore, the handwritten mathematical expression image to be recognized contains a handwritten recognizable mathematical expression;
[0011] Among the optional input information, the context text information is specifically the title of the mathematical expression to be recognized or relevant knowledge text information;
[0012] The latex sequence format information is specifically to replace the latex variables corresponding to the standard answer with the dummy variable space character "" or "#", and only retain the text part of the syntax.
[0013] Furthermore, the output of the OCR recognition model includes the intermediate features of the formula OCR recognition model, the recognition result probability tensor, and the finally recognized latex sequence.
[0014] Furthermore, the step S2 includes the following steps:
[0015] S21. Construct a fine-tuning dataset based on the output result of the OCR recognition model and the input data of the OCR recognition model;
[0016] S22. Use the fine-tuning dataset to perform fine-tuning training on the large language model to obtain a fine-tuned large language model, which is used to understand the probability distribution tensor text, integrate the pre-trained OCR recognition results, context information, and latex sequence format, and output a global confidence output. Among them, during fine-tuning training, randomly perturb the pre-trained OCR recognition results, context information, and latex sequence format, and disrupt the matching degree by randomly perturbing the content of the three pieces of information.
[0017] Furthermore, the specific process of the step S21 is:
[0018] Convert the intermediate features or the recognition result probability tensor Pm of the OCR recognition model into a text sequence tokens in order, and combine the OCR recognition result latex sequence Pred, context information Context, and latex sequence format to construct a fine-tuning dataset containing the user's question and the corresponding output of the model.
[0019] Furthermore, the user's question and the corresponding output of the model in the fine-tuning dataset adopt the following set format:
[0020] A. User's question: "According to the predicted probability distribution [Pm], given the dictionary [Dict], what is the confidence that the prediction result is [Pred]?";
[0021] Model output: "According to the probability distribution, the confidence that its prediction result is correct should be [pho]";
[0022] B. User's question: "Based on the predicted probability distribution [Pm] and the dictionary [Dict], with the question context information being [Context] and the reference latex format being [refL], what should the prediction result be?";
[0023] Model output: "Based on the above information, the prediction result should be [Lab], and its prediction confidence is [pho]";
[0024] C. User's question: "Based on the predicted probability distribution [Pm] and the dictionary [Dict], what should the prediction result be?";
[0025] Model output: "Based on the predicted probability distribution [Pm] and the dictionary [Dict], the prediction result should be [Lab], and its prediction confidence is [pho]";
[0026] D. User's question: "Is the prediction result [Pred] equivalent to the standard answer [Lab]?";
[0027] The model output includes the intermediate reasoning process, the equivalence judgment result, and the confidence [pho].
[0028] Furthermore, during the fine-tuning training process of step S22, the confidence [pho] is calculated from the true similarity between the OCR prediction result [Pred] and the standard answer [Lab].
[0029] Furthermore, step S22 specifically uses incremental fine-tuning, LORA (Low-Rank Adaptation of LLMs), or full-scale fine-tuning to perform fine-tuning training on the large language model.
[0030] Furthermore, during the fine-tuning training process of step S22, context information is used to eliminate ambiguity, the latex sequence format is used as a grammar constraint, and the overall confidence is determined using text sequence tokens.
[0031] Compared with the prior art, the present invention has the following advantages:
[0032] The present invention first pre-trains an OCR recognition model for handwritten mathematical expressions, and then fine-tunes a large language model based on the output results of the OCR recognition model (including intermediate features of the OCR recognition deep neural network, recognition result probability tensors, and the finally recognized latex sequence), combined with the input data of the OCR recognition model (context text information, latex sequence format information); finally, the latex text sequence corresponding to the standard answer is input into the fine-tuned large language model to output a recognition comparison result. Thus, by providing additional context information related to the image to be recognized and combining the intermediate results of OCR recognition as text to the large language model to predict the latex syntax sequence and confidence of the handwritten mathematical expression, it reduces the syntax errors of predicting the latex character sequence and the ambiguity problem of the text to be recognized, can improve the accuracy of handwritten mathematical expression recognition, and based on the reasoning ability of the large model, can effectively solve the problem of heterogeneous answer comparison when the standard answer has non-unique representations.
[0033] On the one hand, the present invention uses context prompt information to disambiguate the OCR recognition results, and on the other hand, uses latex sequence format information to reduce the difficulty caused by the diversity of latex syntax and the inconsistency between the recognition result and the standard answer, thereby realizing effective prompting / constraint of the latex syntax of the recognition result and ensuring the accuracy of the handwritten mathematical expression recognition result.
[0034] The present invention fine-tunes the large language model so that the large language model can accurately understand the probability distribution tensor text, integrate the pre-trained OCR recognition results, context information, and latex sequence format constraints, and can give a global confidence output. During training, incomplete matching is caused by random perturbations of the above three pieces of information. By randomly perturbing the content of the three pieces of information, the matching degree is disrupted, thereby improving the robustness of the model. In addition, the present invention constructs a fine-tuning data set for training the large language model, which can enhance and improve the reasoning ability of the large language model in mathematical formula recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is a schematic flowchart of the method of the present invention;
[0036] Figure 2 is a schematic diagram of the application process of the embodiment;
[0037] Figure 3 is a schematic diagram of a data sample in the fine-tuning data set in the embodiment;
[0038] Figure 4 is a schematic diagram of the application effect of the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0040] Embodiment
[0041] As Figure 1 shown, a method for recognizing and comparing handwritten mathematical expressions based on large model reasoning includes the following steps:
[0042] S1. Pre-train an OCR recognition model for handwritten mathematical expressions;
[0043] S2. Based on the output result of the OCR recognition model, fine-tune the large language model by combining the input data of the OCR recognition model;
[0044] S3. Input the latex text sequence corresponding to the standard answer into the fine-tuned large language model, and output the recognition and comparison result.
[0045] This embodiment applies the above solution. As Figure 2 shown, the main contents are as follows:
[0046] The first step: Pre-train an OCR recognition model for handwritten mathematical expressions. The OCR recognition model inputs the image of the handwritten mathematical expression to be recognized, as well as fuses optional context text information and latex sequence format information, and outputs the complete recognized latex sequence.
[0047] Among them, the image of the handwritten mathematical expression to be recognized contains recognizable mathematical expressions.
[0048] In the optional input part, the context text information refers to the title of the mathematical expression to be recognized or relevant knowledge text information, and inputting the context prompt information into the model realizes the disambiguation of the OCR recognition result;
[0049] The latex sequence format information is to reduce the difficulty caused by the diversity of latex syntax and the inconsistency between the recognition result and the standard answer. In this solution, the latex variables corresponding to the standard answer are replaced with dummy variable space characters "" or "#", and only the text part of the syntax is retained. The latex sequence format information is input into the model to realize the effective prompt / constraint of the latex syntax of the recognition result.
[0050] The output result includes the intermediate features of the OCR recognition deep neural network of the mathematical expression, the recognition result probability tensor, and the finally recognized latex sequence.
[0051] Step 2: Convert the intermediate features or probability distribution tensors output by the pre-trained handwritten mathematical expression OCR recognition model into a probability distribution sequence text tokens according to the rules and input them into the large language model LLM. The OCR model output can also be combined with text such as context information and latex format sequences and input into the large language model LLM again to re-give the recognized latex sequence that conforms to the context and latex format constraints, as well as the overall confidence of the recognition result.
[0052] Specifically, first convert the recognition probability tensor Pm of OCR into a text sequence tokens in order, so that it can be input into the large language model;
[0053] Optionally, input the OCR recognition result latex sequence Pred, context information Context, latex format sequence, and tokens converted from Pm into the large language model together;
[0054] Require the large language model to output the latex expression and overall confidence pho corresponding to the handwritten mathematical expression that best matches the current text information. Among them, the context information plays a role in eliminating possible ambiguities, the latex format sequence plays a grammatical constraint on the output result, and the overall confidence can be obtained from the tokens information converted from the probability tensor.
[0055] In practical applications, use the intermediate features of the deep neural network of the OCR recognition model as Pm and convert them into a text sequence tokens in a fixed order, so that it can be input into the large language model. At this time, Pm is not a probability tensor, but since it contains all the information necessary to generate the probability tensor, this approach is feasible.
[0056] Step 3: Fine-tune the large language model to make it easier to understand this task: understand the probability distribution tensor text, integrate the pre-trained OCR recognition result, context information, and latex sequence format constraints, and be able to give a global confidence output. In this embodiment, during training, the above three pieces of information are randomly perturbed to cause incomplete matching, and the content of the three pieces of information is randomly perturbed to disrupt the matching degree, thereby improving the robustness of the model.
[0057] In addition, in order to improve and enhance the reasoning ability of the large language model LLM in mathematical formula recognition, this solution constructs a fine-tuning data set for training the large language model LLM, such as Figure 3 shown, the sample formats in the data set include the following types:
[0058] "User: Given the predicted probability distribution [Pm] and the dictionary [Dict], what is the confidence level that the prediction result is [Pred]? LLM: According to the probability distribution, the confidence level that its prediction result is correct should be [pho].";
[0059] "User: Given the predicted probability distribution [Pm], the dictionary [Dict], the context information of the question is [Context], and the reference latex format is [refL], what should the prediction result be? LLM: Based on the above information, the prediction result should be [Lab], and the confidence level of its prediction is [pho]."
[0060] "User: Given the predicted probability distribution [Pm] and the dictionary [Dict], what should the prediction result be? LLM: Based on the predicted probability distribution [Pm] and the dictionary [Dict], the prediction result should be [Lab], and the confidence level of its prediction is [pho]."
[0061] The confidence level [pho] in the training stage is calculated from the true similarity between the OCR prediction result [Pred] and the standard answer [Lab], including but not limited to Jaccard similarity, cosine similarity, etc.
[0062] In addition, when the user asks about the equivalence of the comparison between the prediction result [Pred] and the standard answer [Lab], the intermediate reasoning process is written out and told to the large language model; an expression for the prediction result [Pred] that is not equivalent to the standard answer [Lab] can also be set. In this case, through reasoning, it is informed that the mathematical expression is not equivalent after reasoning, and its confidence level [pho].
[0063] In practical applications, fine-tuning of the large model can be carried out in various ways such as incremental fine-tuning, LORA, full-scale fine-tuning, etc.
[0064] Step 4: Input the latex text sequence corresponding to the standard answer into the large language model LLM to realize the logical reasoning comparison of the standard answer and the latex text sequence recognized from the picture based on the large language model (as Figure 4 shown), realize the judgment of the essential equivalence of different mathematical expressions, and solve the problem of polysemy in mathematical formula recognition.
[0065] In summary, this solution provides additional context information related to the image to be recognized, combines the intermediate results of OCR recognition as text and provides them to the large LLM model to predict the Latex syntax sequence and confidence of handwritten mathematical expressions, reduces the syntax errors in predicting the Latex character sequence and the ambiguity problem of the text to be recognized, and improves the recognition accuracy while ensuring a low false positive rate through appropriate data perturbation training techniques. In addition, based on the reasoning ability of the large model, it effectively solves the problem of heterogeneous answer comparison when the standard answer has non-unique representations.
Claims
1. A handwritten mathematical expression recognition and comparison method based on large model reasoning, characterized in that: The following steps are involved: S1, pre-training to obtain the OCR recognition model of handwritten mathematical expressions; S2. Based on the output result of the OCR recognition model, the large language model is fine-tuned in combination with the input data of the OCR recognition model; S3. Input the latex text sequence corresponding to the standard answer into the fine-tuned large language model, and output the recognition and comparison results.
2. According to claim 1, a handwritten mathematical expression recognition and comparison method based on large model reasoning is characterized in that: The input data of the OCR recognition model in the step includes the handwritten mathematical expression image to be recognized and optional input information, and the optional input information includes context text information and latex sequence format information.
3. According to claim 2, a handwritten mathematical expression recognition and comparison method based on large model reasoning is characterized in that: The handwritten mathematical expression image to be recognized contains a handwritten recognizable mathematical expression; In the optional input information, the context text information is specifically the title of the mathematical expression to be recognized or related knowledge text information; The latex sequence format information specifically replaces the latex variables corresponding to the standard answers with dummy variable space characters "" or "#", retaining only the grammatical text.
4. According to claim 2, a handwritten mathematical expression recognition and comparison method based on large model reasoning is characterized in that: The output of the OCR recognition model includes the intermediate features of the OCR recognition model, the recognition result probability tensor and the final recognized latex sequence.
5. According to claim 4, a handwritten mathematical expression recognition and comparison method based on large model reasoning is characterized in that: The step S2 comprises the following steps: S21, constructing a fine-tuning data set based on the output result of the OCR recognition model and the input data of the OCR recognition model; S22. Fine-tune the large language model using the fine-tuning dataset to obtain a fine-tuned large language model for understanding the probability distribution tensor text, integrating the pre-trained OCR recognition results, context information and latex sequence format, and outputting a global confidence output. During the fine-tuning training, the pre-trained OCR recognition results, context information and latex sequence format are randomly perturbed to disrupt the matching degree by randomly perturbing the contents of the three information.
6. The handwritten mathematical expression recognition and comparison method based on large model reasoning according to claim 5 is characterized in that: The specific process of step S21 is as follows: The intermediate features of the OCR recognition model or the recognition result probability tensor Pm are converted into text sequence tokens in order, and combined with the OCR recognition result latex sequence Pred, context information Context, and latex sequence format, a fine-tuning dataset containing user questions and corresponding model outputs is constructed.
7. The handwritten mathematical expression recognition and comparison method based on large model reasoning according to claim 6 is characterized in that: The user questions and the corresponding outputs of the model in the fine-tuning dataset are set in the following format: A. User question: "Based on the predicted probability distribution [Pm], given the dictionary [Dict], what is the confidence level of the predicted result [Pred]?"; Model output: "According to the probability distribution, the confidence level that the prediction is correct should be [pho]"; B. User question: "Based on the predicted probability distribution [Pm] and dictionary [Dict], the question context information is [Context], and the reference latex format is [refL], what should the prediction result be?"; Model output: "Based on the above information, the prediction result should be [Lab], and the confidence level of the prediction is [pho]"; C. User question: "Based on the predicted probability distribution [Pm] and dictionary [Dict], what should the prediction result be?"; Model output: "Based on the predicted probability distribution [Pm] and dictionary [Dict], the predicted result should be [Lab], and the predicted confidence is [pho]"; D. User question: "Are the predicted result [Pred] and the standard answer [Lab] equivalent?"; The model output includes the intermediate reasoning process, equivalence judgment results and confidence [pho].
8. The handwritten mathematical expression recognition and comparison method based on large model reasoning according to claim 7 is characterized in that: During the fine-tuning training process in step S22, the confidence [pho] is calculated by the true similarity between the OCR prediction result [Pred] and the standard answer [Lab].
9. The handwritten mathematical expression recognition and comparison method based on large model reasoning according to claim 5 is characterized in that: The step S22 specifically uses incremental fine-tuning, LORA or full fine-tuning to perform fine-tuning training on the large language model.
10. The handwritten mathematical expression recognition and comparison method based on large model reasoning according to claim 5 is characterized in that: During the fine-tuning training process of step S22, context information is used to eliminate ambiguity, latex sequence format is used as grammatical constraints, and text sequence tokens are used to determine the overall confidence.
Citation Information
Cited By
Formula identification and model training method and device, related equipment and program product
CN120544225A