Handwriting vertical recognition and calculation system and method based on position forest

By introducing position forest coding and implicit attention correction modules into the handwriting vertical recognition system, the problem of difficulty in identifying complex handwriting vertical mathematical expressions in the prior art is solved, and efficient and accurate handwriting mathematical expression recognition and calculation are achieved, which is suitable for education and scientific research applications in multiple fields.

CN119942565APending Publication Date: 2025-05-06ROBOTICS RESEARCH CENTER OF YUYAO CITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510138715.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and process complex handwritten vertical mathematical expressions, especially when identifying positional relationships and hierarchies between symbols.

Method used

A handwritten vertical recognition and computing system based on location forest is adopted, which includes image acquisition, preprocessing, recognition, calculation and display modules. The identification module uses DenseNet backbone network, location forest coding and implicit attention correction module to analyze the relative position relationship between symbols through location forest coding to achieve efficient and accurate handwritten mathematical expression recognition.

Benefits of technology

It realizes the accurate identification and calculation of complex handwritten vertical mathematical expressions, can handle cross-row formulas and complex structures, provides clear visual calculation results, and is suitable for fields such as education and scientific research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942565A_ABST
    Figure CN119942565A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing and mode recognition, and discloses a handwriting vertical recognition and calculation system and method based on a position forest, and the system comprises an image collection and input module, an image preprocessing module, a recognition module, a calculation module and a display module. The image acquisition and input module is used for writing a mathematical formula, identifying the boundary of the handwritten formula and intercepting an image area containing a complete formula; the image preprocessing module is used for preprocessing the image; the identification module is used for realizing handwritten mathematical expression identification; the calculation module is used for decoding according to the recognition result, extracting key characters from the generated latex sequence, and calculating the result according to a python calculation formula; and the display module is used for displaying a calculation result on a user interface. According to the invention, the calculator input form is innovated, the calculation mode relying on the user hand drawing mobile terminal and digital pen input and the handwriting vertical recognition and calculation are realized, and the blank in the vertical recognition field is filled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing and pattern recognition, and in particular relates to a handwritten vertical recognition and calculation system and method based on position forest. Background Art

[0002] As a special form of calculation, vertical form is a very common arithmetic operation in mathematics education in primary and secondary schools. In order to better support students' learning and teachers' teaching, a vertical form recognition network is urgently needed to solve the vertical form calculation problem. In the automatic correction system, recognizing students' vertical form writing and judging its correctness is an important function.

[0003] Handwritten mathematical expression recognition (HMER) has a wide range of applications in human-computer interaction scenarios, such as digital education and automated office. Existing technologies mainly rely on sequence-based models that solve this task by directly predicting LaTeX sequences. These methods only implicitly learn the grammatical rules provided by LaTeX and preliminarily solve the recognition of single-line handwritten formulas. Since they cannot accurately describe the positional relationship and hierarchical structure between symbols, they cannot effectively recognize vertical formulas, such as matrices, determinants, and other cross-row formulas. They perform poorly when faced with complex structural relationships and diverse writing styles. Summary of the invention

[0004] The purpose of the present invention is to provide a handwritten vertical recognition and calculation system and method based on location forest to solve the above-mentioned technical problems.

[0005] In order to solve the above technical problems, the specific technical solutions of the handwritten vertical recognition and calculation system and method based on location forest of the present invention are as follows: A handwritten vertical recognition and calculation system based on location forest, comprising an image acquisition and input module, an image preprocessing module, a recognition module, a calculation module and a display module; The image acquisition and input module is used to write mathematical formulas, and identify the boundaries of handwritten formulas through image processing technology, and intercept the image area containing the complete formula; the image is uploaded to the cloud or locally deployed for processing; The image preprocessing module is used to preprocess the image; The recognition module: a sequence-based encoder-decoder method comer is used as a baseline model, which consists of a DenseNet backbone network, a position forest encoding, and an implicit attention correction module to achieve efficient and accurate handwritten mathematical expression recognition; The calculation module: decodes the recognition result, extracts key characters from the generated latex sequence, and calculates the result according to the python calculation formula; The display module displays the calculation results on the user interface to provide a clear visualization effect.

[0006] The present invention also discloses a handwritten vertical type recognition and calculation method of a handwritten vertical type recognition and calculation system based on a location forest, comprising the following steps: S1. Image acquisition and input: S2. Image preprocessing: S3. Image recognition: S4. Calculate the recognition results: S5. Display the calculation results.

[0007] Furthermore, the S1 comprises the following steps: The user writes mathematical formulas on the handwriting input pad of the mobile device. The device uses image processing technology to identify the boundaries of the handwritten formula and capture the image area containing the complete formula, ensuring that all symbols and structures are accurately captured.

[0008] Furthermore, the S2 comprises the following steps: Grayscale: Convert a color image to a grayscale image; Binarization: Convert a grayscale image into a binary image; Stroke Erosion: Perform stroke erosion operation on binary images.

[0009] Furthermore, the S3 image recognition sequence-based encoder-decoder method comer is used as a baseline model, which consists of a DenseNet backbone network, a position forest encoding and an implicit attention correction module.

[0010] Furthermore, the S3's DenseNet backbone network first extracts rich 2D visual features from the input image; these features are then input into the attention-based Transformer decoder to obtain discriminative symbol features; a parallel linear head is then deployed to recognize LaTeX expressions, and together with symbol recognition, a position forest is introduced for joint optimization, where each identifier represents a string and indicates its position information; using this encoding, two position forest heads are then used to parse their nesting levels and relative positions in the forest; during the reasoning process: the input image passes through the backbone, decoder, and expression recognition head in sequence to predict the LaTeX sequence.

[0011] Furthermore, the position forest encoding of S3: a position forest structure is proposed to model mathematical expressions and parse the relative position relationship between symbols, which is jointly optimized with the symbol recognition task to promote the learning of position-aware symbol-level feature representation. Each symbol is assigned a position identifier to indicate its relative spatial position in the two-dimensional image. The encoding process includes segmenting the LaTeX sequence into multiple substructures, then constructing a tree structure according to the relative positions of the symbols within the substructures, and finally combining them into a position forest.

[0012] Furthermore, the implicit attention correction module of S3: introduces an implicit attention correction module, which adaptively introduces zero attention as a refinement term, utilizes past alignment information to optimize attention weights, models mathematical expressions as a position forest structure, and parses the nesting level and relative position between symbols, thereby explicitly implementing position-aware symbol-level feature representation learning in a sequence-based encoder-decoder model to obtain accurate LaTeX sequences, especially when recognizing complex expressions.

[0013] Further, the S4 comprises the following steps: Decode the recognition result, extract key characters from the generated latex sequence, and calculate the result according to the Python calculation formula.

[0014] Furthermore, the S5 comprises the following steps: Result transmission: If the inference is performed on the server, the calculation results are transmitted back to the local interface through the HTTP protocol. If it is deployed locally, the calculation results are directly displayed on the local interface. Result display: Display the calculation results on the user interface to provide clear visualization effects.

[0015] The handwritten vertical formula recognition and calculation system and method based on location forest of the present invention have the following advantages: The present invention provides a complete system flow from image acquisition, formula recognition, LaTeX conversion, formula calculation to result display. The system can efficiently process the mathematical formulas handwritten by the user, automatically perform symbolic and numerical calculations, and provide a friendly user interaction experience. It is applicable to multiple fields such as education and scientific research. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a system module block diagram of the present invention; Figure 2 It is a schematic diagram of the model framework of the present invention. DETAILED DESCRIPTION

[0017] In order to better understand the purpose, structure and function of the present invention, the handwritten vertical recognition and calculation system and method based on location forest of the present invention are further described in detail below with reference to the accompanying drawings.

[0018] The handwritten vertical recognition and calculation system based on location forest of the present invention provides two solutions: server cloud deployment and local deployment. Figure 1 As shown, the system includes an image acquisition and input module, an image preprocessing module, a recognition module, a calculation module and a display module.

[0019] The image acquisition and input module is used to write mathematical formulas, and recognize the boundaries of handwritten formulas through image processing technology, and intercept the image area containing the complete formula; Image preprocessing module: used to preprocess images to improve recognition accuracy; Recognition module: A sequence-based encoder-decoder approach comer is used as the baseline model, which consists of a DenseNet backbone network, a position forest encoding, and an implicit attention correction module to achieve efficient and accurate recognition of handwritten mathematical expressions (especially vertical expressions); Calculation module: decodes the recognition results, extracts key characters from the generated latex sequence, and calculates the results according to the Python calculation formula; Display module: displays the calculation results on the user interface, providing clear visualization effects.

[0020] like Figure 2 As shown, the handwritten vertical recognition and calculation system and method based on location forest of the present invention includes the following steps: S1. Image acquisition and input: The user writes mathematical formulas on the handwriting input board of the mobile device. The device uses image processing technology to identify the boundaries of the handwritten formula and capture the image area containing the complete formula to ensure that all symbols and structures are accurately captured. Image upload: In order to optimize transmission efficiency, the captured image is compressed to reduce transmission time, and then uploaded to the cloud server through the network for subsequent processing. If used in local deployment, it does not need to be uploaded to the cloud, but can be sent to the local software for inference, thereby reducing latency and protecting user privacy.

[0021] S2. Image preprocessing: Before entering the recognition stage, the image needs to go through a series of preprocessing steps to improve recognition accuracy: Grayscale: Convert color images to grayscale images to reduce the interference of color information on the recognition process. Binarization: Convert grayscale images to binary images to make handwriting more prominent and the background more concise. Stroke erosion: Perform stroke erosion operations on binary images to remove small noise points, improve the cleanliness of the image, and provide a better basis for subsequent recognition.

[0022] S3. Image recognition: namely, the handwritten vertical form recognition network, the sequence-based encoder-decoder method comer is used as the baseline model, which consists of a DenseNet backbone network, a position forest encoding, and an implicit attention correction module to achieve efficient and accurate recognition of handwritten mathematical expressions (especially vertical forms).

[0023] The DenseNet backbone network first extracts rich 2D visual features from the input image. These features are then fed into the attention-based Transformer decoder to obtain discriminative symbol features. A parallel linear head is then deployed to recognize LaTeX expressions to ensure the professionalism and compatibility of the output format. Together with symbol recognition, position forests are introduced for joint optimization to promote the learning of position-aware symbol-level feature representations. Each identifier represents a string that indicates its position information. Using this encoding, two position forest heads are then used to parse their nesting levels and relative positions in the forest. This encoding method allows the system to parse the nesting levels and relative position relationships between symbols without additional manual annotation work. During reasoning: the input image passes through the backbone, decoder, and expression recognition head in sequence to predict the LaTeX sequence. Note that the position forest encoding and position forest head are removed during reasoning, which does not incur additional latency or computational cost.

[0024] Position Forest Encoding: A position forest structure is proposed to model mathematical expressions and parse the relative position relationship between symbols. It is jointly optimized with the symbol recognition task to promote the learning of position-aware symbol-level feature representation. Each symbol is assigned a position identifier to indicate its relative spatial position in a two-dimensional image. This encoding method allows the system to parse the nesting level and relative position relationship between symbols without additional annotation work. The encoding process involves splitting the LaTeX sequence into multiple substructures, then building a tree structure based on the relative positions of the symbols inside the substructures, and finally combining them into a position forest.

[0025] Implicit Attention Correction Module: An implicit attention correction module is introduced to improve the attention accuracy in the HMER task under the sequence-based decoding architecture. By adaptively introducing zero attention as a refinement term, past alignment information is used to optimize the attention weights to obtain more refined feature representations. Mathematical expressions are modeled as a position forest structure, and the nesting level and relative position between symbols are parsed to explicitly achieve position-aware symbol-level feature representation learning in the sequence-based encoder-decoder model. "_", "_", "{", and "}" can be efficiently generated to obtain accurate LaTeX sequences, especially when recognizing complex expressions.

[0026] S4. Calculate the recognition results: Decode the recognition results, extract key characters such as numbers and calculation symbols from the generated latex sequence, and calculate the results according to the Python calculation formula. This process ensures the accuracy of the calculation and can handle complex mathematical expressions, including but not limited to addition, subtraction, multiplication, division and their combinations.

[0027] S5. Display the calculation results: Result transmission: If the inference is performed on the server, the calculation results are transmitted back to the local interface through the HTTP protocol. If it is deployed locally, the calculation results are directly displayed on the local interface. Result display: The calculation results are displayed on the user interface to provide clear visualization effects.

[0028] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.

Claims

1. A handwritten vertical recognition and calculation system based on location forest, characterized in that: It includes image acquisition and input module, image preprocessing module, recognition module, calculation module and display module; The image acquisition and input module is used to write mathematical formulas, and identify the boundaries of handwritten formulas through image processing technology, and intercept the image area containing the complete formula; the image is uploaded to the cloud or locally deployed for processing; The image preprocessing module is used to preprocess the image; The recognition module: a sequence-based encoder-decoder method comer is used as a baseline model, which consists of a DenseNet backbone network, a position forest encoding, and an implicit attention correction module to achieve efficient and accurate handwritten mathematical expression recognition; The calculation module: decodes the recognition result, extracts key characters from the generated latex sequence, and calculates the result according to the python calculation formula; The display module displays the calculation results on the user interface to provide a clear visualization effect.

2. A method for handwritten vertical type recognition and calculation of the handwritten vertical type recognition and calculation system based on location forest as claimed in claim 1, characterized in that: The steps include: S1. Image acquisition and input: S2. Image preprocessing: S3. Image recognition: S4. Calculate the recognition results: S5. Display the calculation results.

3. The method according to claim 2, characterized in that The S1 comprises the following steps: The user writes mathematical formulas on the handwriting input pad of the mobile device. The device uses image processing technology to identify the boundaries of the handwritten formula and capture the image area containing the complete formula, ensuring that all symbols and structures are accurately captured.

4. The method according to claim 2, characterized in that: The S2 comprises the following steps: Grayscale: Convert a color image to a grayscale image; Binarization: Convert a grayscale image into a binary image; Stroke Erosion: Perform stroke erosion operation on binary images.

5. The method according to claim 2, characterized in that: The S3 image recognition sequence-based encoder-decoder method comer is used as the baseline model, which consists of a DenseNet backbone network, a position forest encoding, and an implicit attention correction module.

6. The method according to claim 5, characterized in that The S3 DenseNet backbone network first extracts rich 2D visual features from the input image; these features are then fed into an attention-based Transformer decoder to obtain discriminative symbol features; Then a parallel linear head is deployed to recognize LaTeX expressions. Together with symbol recognition, a position forest is introduced for joint optimization. Each identifier represents a string, indicating its position information. Using this encoding, two position forest heads are then used to parse their nesting levels and relative positions in the forest. During the inference process: the input image passes through the backbone, decoder, and expression recognition head in sequence to predict the LaTeX sequence.

7. The method according to claim 5, characterized in that The position forest encoding of S3: A position forest structure is proposed to model mathematical expressions and parse the relative position relationship between symbols, which is jointly optimized with the symbol recognition task to promote the learning of position-aware symbol-level feature representation. Each symbol is assigned a position identifier to indicate its relative spatial position in the two-dimensional image. The encoding process includes segmenting the LaTeX sequence into multiple substructures, then constructing a tree structure according to the relative positions of the symbols within the substructures, and finally combining them into a position forest.

8. The method according to claim 5, characterized in that The implicit attention correction module of S3: introduces an implicit attention correction module, which explicitly implements position-aware symbol-level feature representation learning in a sequence-based encoder-decoder model by adaptively introducing zero attention as a refinement term, leveraging past alignment information to optimize attention weights, modeling mathematical expressions as a position forest structure, and parsing the nesting levels and relative positions between symbols to obtain accurate LaTeX sequences, especially when recognizing complex expressions.

9. The method according to claim 2, characterized in that: The S4 comprises the following steps: Decode the recognition result, extract key characters from the generated latex sequence, and calculate the result according to the Python calculation formula.

10. The method according to claim 2, characterized in that The S5 comprises the following steps: Result transmission: If the inference is performed on the server, the calculation results are transmitted back to the local interface through the HTTP protocol. If it is deployed locally, the calculation results are directly displayed on the local interface. Result display: Display the calculation results on the user interface to provide clear visualization effects.