Handwritten mathematical formula data set construction method and system based on multi-modal acquisition and semantic annotation
Through the construction method of handwritten mathematical formula data set of multimodal acquisition and semantic annotation, the digital pen and multi-stage generation adversarial network MSF-Gan are used to solve the problem of high data acquisition cost in handwritten mathematical expression recognition, and the support for efficient generation and recognition model training is achieved.
Patent Information
- Application Number
- CN202510529599.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-12
AI Technical Summary
The existing handwritten mathematical expression recognition methods rely on a large amount of real data to train, but the acquisition cost is high and the quality is difficult to grasp, resulting in limited performance improvement.
The construction method of handwritten mathematical formula data set using multimodal acquisition and semantic annotation includes automatically synthesizing high-quality handwritten mathematical formula images based on digital pen acquisition and multi-stage generation of adversarial network MSF-Gan, and combining manual annotation to reduce costs and enhance data diversity.
It realizes rapid and efficient generation of reliable data, significantly improves the generalization ability of handwritten formula recognition models, and supports large-scale training and recognition tasks.
Smart Images

Figure CN120472477A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and educational technology, and in particular relates to a method and system for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation. Background Art
[0002] Handwritten mathematical expressions (HMEs) have widespread applications in education and scientific research. However, due to the structural complexity and diverse handwriting of handwritten mathematical expressions, automatic recognition of handwritten mathematical expressions (HMERs) remains a challenging task. Existing methods rely primarily on large amounts of real-world data for training, but acquiring real-world data is expensive and difficult to ensure quality. Therefore, developing a method and system for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation is crucial for improving the performance of HMER systems. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation to solve the above-mentioned technical problems.
[0004] To solve the above technical problems, the specific technical solutions of the method and system for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation are as follows: A method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation includes the following steps: Step 1: Capture handwritten mathematical formulas using a digitizer pen: Automatically generate well-structured mathematical formulas to provide standardized input for subsequent handwritten collection; Step 2: Build a model HMEG for end-to-end handwritten mathematical formula generation based on symbolic graph guidance: Through the multi-stage generative adversarial network MSF-Gan, high-quality handwritten mathematical formula images are automatically synthesized to supplement manually collected data, reducing annotation costs and enhancing data diversity.
[0005] Furthermore, the step 1 includes the following steps: Step 1.1: Generate a regularized LaTeX sequence; Step 1.2. Collect handwriting data using a digitizer pen.
[0006] Furthermore, the step 1.1 includes the following steps: Symbol library construction: Based on K12 education needs, a target dictionary containing numbers, operators, and complex symbols is established; Random formula generation: Generate random operands and operators; Avoid mathematical errors; Convert to LaTeX format; Output: Generates a standard LaTeX string for front-end rendering.
[0007] Furthermore, the step 1.2 includes the following steps: Step 1.2.1: Data collection process: Participant recruitment: Invite 100 volunteers; Record metadata; Writing task: The system displays the formula image rendered by LaTeX; Volunteers used digitizing pens to write on the electronic canvas; The collected data include: handwriting coordinate sequence and pressure data; Function support: Provides undo, redo, and eraser tools to simulate real writing experience; Step 1.2.2: Data annotation and quality control: Including three-level annotation system: Difficulty level criteria: Invalid writing errors, incomplete, or illegible handwriting; The number of simple characters is less than 20, and there is no complex structure; Difficulty containing complex symbols or illegible handwriting; Complex symbol definition: Complex characters: Chinese symbols, special symbols; Complex structures: fractions, matrices, calculus symbols, geometric symbols; Invalid data filtering: Line break detection: check whether the written line break is consistent with the formula structure; Manual review: Labelers verify the validity of the data to ensure that the labeling is accurate.
[0008] Furthermore, the step 2 includes the following steps: Step 2.1: Layout prediction stage; Step 2.2: Mask optimization stage; Step 2.3: Image generation stage.
[0009] Furthermore, the step 2.1 includes the following steps: Input: LaTeX sequence; Processing: Use the GCN graph convolutional network to parse the formula structure, establish the topological relationship between symbols, and predict the bounding box position of each symbol; Output: A two-dimensional layout of the formula.
[0010] Furthermore, the step 2.2 includes the following steps: Input: layout prediction results; Processing: Use the U-Net structure generator to convert the bounding box into a pixel-level soft mask to solve the layout ambiguity problem; Output: high-resolution mask map.
[0011] Furthermore, the step 2.3 includes the following steps: Input: optimized mask map; Processing: Generate realistic handwriting based on a diffusion model, simulating natural writing characteristics: connected strokes, jitter, pressure changes, and style diversity; Output: Final handwritten formula image.
[0012] The present invention also discloses a multi-source handwritten mathematical formula data set management and automated testing system, including an online collection and annotation module, a data management module and an automated testing module. The online collection and annotation module includes a data collection process and quality control rules. The data collection process includes user login, handwriting input, and intelligent annotation. User login supports multi-identity login and hierarchical authority management. Handwriting input uses a passive pen to write input, records the stroke trajectory in real time, and supports writing area screenshots, handwriting undo / redo, and canvas clearing functions. Intelligent annotation includes automatic generation of LaTeX sequences, user verification mechanism, and automatic classification and storage of data. The user verification mechanism: if the recognition is correct, the user confirms the submission; if the recognition is incorrect, the user corrects the annotation. Data is automatically classified and stored in a correct sample library and an incorrect sample library. The quality control rules include validity judgment criteria and a hierarchical annotation system. Valid data of the validity judgment criteria are data with complete formula structure, clear handwriting, and conformity to LaTeX semantics; invalid data are data with illegal line breaks, illegible handwriting, and unfinished writing. The hierarchical annotation system includes judgment criteria and processing methods for valid data, boundary data, and invalid data. The data management module includes platform function design and data return mechanism. The platform function design includes visual dashboard, data management and authority management. The visual dashboard includes real-time data statistics, collection progress monitoring and user contribution ranking; data management includes multi-dimensional retrieval, batch export and version control; authority management includes role classification and operation log audit; data return mechanism includes error sample recovery and dynamic amplification strategy. The data return mechanism includes error sample recovery and dynamic amplification strategy. Error sample recovery includes automatic classification and identification of error cases and adding them to the training set after review by annotators. The dynamic amplification strategy includes automatic detection of hotspot symbols and active learning sampling. The automated testing module includes test process design and continuous integration support. Test process design includes test task creation and performance evaluation. Test task creation supports batch uploading and customized test parameters. Performance evaluation includes core indicators and anomaly detection. Core indicators include recognition accuracy, processing latency, and peak memory usage. Anomaly detection includes classification of typical error patterns and automatic archiving of failure cases. Continuous integration support includes API interface and benchmark testing. API interface includes RESTful interface to support automated pipeline and automatic generation of test reports. Benchmark testing includes cross-model comparative testing and historical performance trend analysis.
[0013] The method and system for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation of the present invention have the following advantages: the method of the present invention can realize batch generation and management of handwritten formula data. The training of the handwritten formula recognition model requires the support of a large-scale dataset, and synthetic data enhancement significantly improves the generalization ability of the model in various recognition tasks, combined with manual annotation and drawing. The dataset construction method of the present invention can generate reliable data quickly and efficiently. At present, there are few categories of public datasets available in the field of handwritten mathematical formula recognition, which makes it difficult to meet all the symbols covered by the application, and the corpus information is seriously insufficient, which brings difficulties to the model's understanding of the image context. The present invention solves these problems through advanced technologies and algorithms, and realizes a method and system for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation, which provides strong support for the training of the handwritten formula recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flow chart of the method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation of the present invention; Figure 2 A schematic diagram showing the handwritten formula recognition method of the present invention; Figure 3 This is a schematic diagram of the handwritten formula recognition interactive interface of the present invention; Figure 4 This is a functional diagram of each module of the handwritten mathematical formula collection system of the present invention; Figure 5 Schematic diagram of the overall framework of the multi-stage formula generation algorithm net (MSF-Gan) based on the end-to-end framework of the present invention; Figure 6 This is a schematic diagram of the MSF-Gan model generation effect of the present invention; Figure 7 Schematic diagram of the online data statistics platform of the present invention; Figure 8 This is a schematic diagram of the automated test interface of the present invention; DETAILED DESCRIPTION
[0015] In order to better understand the purpose, structure and function of the present invention, the following is a further detailed description of the method and system for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation in conjunction with the accompanying drawings.
[0016] The method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation of the present invention comprises the following steps: Step 1: Capture handwritten mathematical formulas using a digitizer pen Step 1.1: Generate a regular LaTeX sequence Purpose: To automatically generate well-structured mathematical formulas to provide standardized input for subsequent handwriting collection.
[0017] Method: Symbol library construction: Based on K12 education needs, a target dictionary containing numbers, operators, and complex symbols (such as integrals, matrices, etc.) is established.
[0018] Random formula generation: Generate random operands (integers, variables) and operators (+, -, ×, ÷, etc.).
[0019] Avoid mathematical errors (such as dividing by 0 and repeating operators consecutively).
[0020] Convert to LaTeX format (such as ×→\times, ÷→\div).
[0021] Output: Generates a standard LaTeX string (such as $1 + 1 = 2$) for front-end rendering.
[0022] Advantages: Ensures the legality of formula structures and reduces invalid samples. Supports large-scale generation and covers a wide range of mathematical expressions.
[0023] Step 1.2. Collect handwriting data using a digitizer pen Purpose: To collect real and diverse handwritten mathematical formula data for K12 education scenarios.
[0024] Step 1.2.1: Data collection process Participant Recruitment: Invite 100 volunteers (of different ages and writing habits).
[0025] Record metadata (age, gender, writing hand, etc.).
[0026] Writing tasks: The system displays the formula image rendered by LaTeX.
[0027] Volunteers write on an electronic canvas using a digital pen (such as a Wacom tablet / Hanwang education tablet).
[0028] The collected data includes: Handwriting coordinate sequence (sampling rate of 500Hz) Pressure data (2048-level pressure sensitivity) Function support: Provide tools such as undo, redo, and eraser to simulate a real writing experience.
[0029] Step 1.2.2: Data annotation and quality control Three-level annotation system: Difficulty level Judgment criteria Invalid Writing errors, incomplete, unrecognizable handwriting Simple Character quantity < 20, no complex structure (such as x^2 + 1) Difficult Containing complex symbols (such as integral ∫, matrix) or scribbled handwriting Definition of complex symbols: Complex characters: Chinese symbols ("and", "or"), special symbols (°, ①~⑨).
[0030] Complex structures: fractions, matrices, calculus symbols (∫dx), geometric symbols (∥, ⌒).
[0031] Filtering of invalid data: Line break detection: Compare whether the writing line break is consistent with the formula structure (mark as invalid if inconsistent).
[0032] Manual review: The annotator verifies the data validity to ensure accurate annotation.
[0033] Step 2: Build a model (HMEG) for generating end-to-end handwritten math formulas guided by a symbol graph Through a multi-stage generative adversarial network (MSF-Gan), automatically synthesize high-quality handwritten math formula images as a supplement to manually collected data, reducing the annotation cost and enhancing data diversity. Adopt a three-stage generation process: Step 2.1: Layout prediction stage Input: LaTeX sequence (such as $x^2 + y = 1$).
[0034] Processing: Analyze the formula structure through a GCN (Graph Convolutional Network) to establish the topological relationship between symbols. Predict the bounding box position of each symbol (for example, x² should be above the fraction line).
[0035] Output: A two-dimensional layout diagram (symbol position heat map) of the formula.
[0036] Step 2.2: Mask optimization stage Input: Layout prediction results.
[0037] Processing: Use the U-Net structure generator to convert bounding boxes into pixel-level soft masks. This solves layout ambiguity issues (such as accurate segmentation of overlapping subscripts and subscripts).
[0038] Output: High-resolution mask image (distinguishing strokes from background).
[0039] Step 2.3: Image generation phase Input: Optimized mask map.
[0040] Processing: Generate realistic handwriting based on a diffusion model, simulating natural handwriting characteristics: ligatures, jitter, pressure variations (achieved through noise injection), and stylistic diversity (variation in handwriting between different writers).
[0041] Output: Final handwritten formula image (PNG / INKML format).
[0042] The present invention's multi-source handwritten mathematical formula dataset management and automated testing system utilizes an integrated "collection-annotation-management-testing" design to achieve full lifecycle management of handwritten mathematical formula data. The system includes an online collection and annotation module, a data management module, and an automated testing module. Through human-in-the-loop interaction, it ensures closed-loop optimization of data quality and model iteration.
[0043] The online data collection and annotation module includes a data collection process and quality control rules. The data collection process includes user login, handwriting input, and intelligent annotation. User login supports multiple identities (student / teacher / annotator) and hierarchical permission management. Handwriting input uses a passive pen (supporting 2048 levels of pressure sensitivity), records stroke trajectories in real time (sampling rate ≥ 500Hz), and supports screenshots of writing areas, handwriting undo / redo, and canvas clearing. Intelligent annotation includes automatic generation of LaTeX sequences, user verification mechanisms, and automatic data classification and storage. The user verification mechanism: correct recognition: confirmation submission; incorrect recognition: annotation correction; data is automatically classified and stored in a correct sample library and an incorrect sample library (for model iteration). Quality control rules include validity judgment criteria and a hierarchical annotation system. Valid data in the validity judgment criteria refers to data with complete formula structure, clear and legible handwriting, and conformity to LaTeX semantics. Invalid data refers to data with illegal line breaks, illegible handwriting, and unfinished writing. The hierarchical annotation system is shown in the following table:
[0044] The data management module encompasses platform functionality design and a data reflow mechanism. Platform functionality includes a visual dashboard, data management, and permissions management. The dashboard includes real-time data statistics (total volume / category distribution), collection progress monitoring, and user contribution rankings. Data management includes multi-dimensional search (by formula type, collection time, author, etc.), batch export (supporting INKML, PNG, and LaTeX formats), and version control (supporting dataset differential comparison). Permission management includes role grading (administrator / quality inspector / ordinary user) and operation log auditing. The data reflow mechanism includes error sample recovery and dynamic amplification strategies. Error sample recovery involves automatic classification and identification of erroneous cases, and their addition to the training set after human review. The dynamic amplification strategy includes automatic detection of hotspot symbols (prioritizing low-accuracy symbols) and active learning sampling (prioritizing samples with high uncertainty).
[0045] The automated testing module includes test process design and continuous integration support. Test process design includes test task creation and performance evaluation. Test task creation supports batch uploads (ZIP / folder) and custom test parameters (recognition timeout threshold and accuracy calculation method). Performance evaluation includes core metrics and anomaly detection. Core metrics include recognition accuracy (by symbol / formula grading), processing latency (P50 / P90 / P99), and peak memory usage. Anomaly detection includes typical error pattern classification and automatic archiving of failure cases. Continuous integration support includes API interfaces and benchmarking. The API interface includes a RESTful interface that supports automated pipelines and automatic test report generation (PDF / HTML). Benchmarking includes cross-model comparative testing and historical performance trend analysis.
[0046] This system has the following advantages: 1. Full process automation: - Complete closed loop from data acquisition to model testing - Reduce manual intervention by more than 70% 2. Educational scenario adaptation: - Support K12 special symbols (∵∴∥, etc.) - Handwriting-formula dual-mode storage 3.Quality assurance system: - Three-level data quality inspection process - Dynamic difficult sample mining 4. Performance indicators: - Daily processing capacity: 200,000+ formulas - Average annotation time: 1.2 minutes / sample - Test set coverage: 100% coverage of 385 types of symbols By tightly integrating data collection, quality management, and model testing, this system has established a standardized infrastructure for handwritten mathematical formula research, providing reliable data support and evaluation tools for educational AI applications. The system has been widely deployed in scenarios such as K12 math homework grading, achieving a 97% accuracy rating on a dataset exceeding 30 million results. Example
[0047] Design a technical route for real-scene data collection and annotation based on an online platform. Figure 1 shown.
[0048] Step 1: Collect handwritten mathematical formulas based on the digitizer pen; The login interface of the handwritten mathematical formula data collection module is as follows Figure 2 As shown in the figure, the user selects an identity to log in. This handwritten mathematical formula recognition software implements functions such as passive pen writing, writing area screenshot, handwritten mathematical formula recognition, and withdrawal. The main interface consists of a submit button, a writing area, and functional modules. The recognition method proposed in Chapter 2 is used for reasoning, and a human-in-the-loop design method is introduced. The user confirms and judges to collect data. The handwritten formula recognition interactive interface is shown in the figure. Figure 3 As shown in Figure 3, the functions of each module of the handwritten mathematical formula acquisition system are shown in Figure 3. The user first opens the software to activate the system and uses a passive pen to write a mathematical formula in a specified area on the screen. After the user completes writing, the front-end captures the drawn area on the screen, sets the image size to the writing area, and records the stroke trajectory. The model recognition results are transmitted back to the software page for display, and the user determines whether the recognition is correct. If the recognition is correct, click the "Recognize Correct" button to submit the result. A "Submission Received" notification will be received, indicating that the submission was successful. If the recognition is incorrect, click the "Recognize Error" button. The back-end will further organize the data based on the recognition results and classify the captured user data into two groups: correct and incorrect, preparing for subsequent training data.
[0049] Step 2: Build a model for end-to-end handwritten mathematical formula generation based on symbolic graph guidance (HMEG) Design a multi-stage text generation network (Multi-stage Formula Generation Algorithm Net, MSF-Gan) based on an end-to-end framework. Figure 5As shown, it is intended to serve as a supplementary method for manual collection, reducing the high cost and difficulty of automation of manual collection. Data can be automatically generated based on the diffusion model generation method. The overall model framework is designed to be end-to-end (graph → layout → mask → image), including layout prediction, mask optimization and image generation, which are jointly optimized in an end-to-end manner. The label domain and the symbol domain are aligned through the GCN network, and the bounding box is mapped to a pixel-level soft mask to solve the ambiguity problem of the layout area. Text-generated images are used as augmented data as a useful supplement to enhance the diversity of the dataset, improve the robustness of the model for fuzzy data recognition, and deal with the problems of connected strokes, repeated strokes, jitter and blurring in practical applications, ensuring balanced recognition performance of the model for various characters and scenes. Figure 6 This is the effect diagram generated by the MSF-Gan model, and 3,000 pictures are generated as supplements.
[0050] Step 3: Design a multi-source handwritten mathematical formula dataset management and automated testing system Design a data management module based on the online platform, Figure 7 It is an online data collection platform interface, which is designed with menu bar, toolbar, address bar, task bar, title bar, status bar and user login window. It displays the total amount of data currently collected and its category. You can select the corresponding tool in the menu bar and the data collection status by selecting the formula data bar. In order to quickly verify the quality of the data set and the performance of the algorithm, you can upload the data set in batches to view the recognition effect, which is convenient for users to test in the interactive interface. An automated test module is designed. Figure 8 This section introduces the interface design and functionality. In the automated testing interface, users can add automated testing, enter a task name, and specify test data. Once the task is completed, the status bar displays the total number of test data, success rate, abnormal data alerts, and average recognition time, allowing users to easily understand the automated testing process. The index bar also allows users to search for recognition results and time intervals.
[0051] It will be understood that the present invention is described by way of some embodiments, and it will be appreciated by those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are intended to be protected by the present invention.
Claims
1. A method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation, characterized in that: The steps include: Step 1: Capture handwritten mathematical formulas using a digitizer pen: Automatically generate well-structured mathematical formulas to provide standardized input for subsequent handwritten collection; Step 2: Build a model HMEG for end-to-end handwritten mathematical formula generation based on symbolic graph guidance: Through the multi-stage generative adversarial network MSF-Gan, high-quality handwritten mathematical formula images are automatically synthesized to supplement manually collected data, reducing annotation costs and enhancing data diversity.
2. The method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation according to claim 1 is characterized in that: The step 1 comprises the following steps: Step 1.1: Generate a regularized LaTeX sequence; Step 1.
2. Collect handwriting data using a digitizer pen.
3. The method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation according to claim 2 is characterized in that: The step 1.1 includes the following steps: Symbol library construction: Based on K12 education needs, a target dictionary containing numbers, operators, and complex symbols is established; Random formula generation: Generate random operands and operators; Avoid mathematical errors; Convert to LaTeX format; Output: Generates a standard LaTeX string for front-end rendering.
4. The method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation according to claim 2, characterized in that: The step 1.2 includes the following steps: Step 1.2.1: Data collection process: Participant recruitment: Invite 100 volunteers; Record metadata; Writing task: The system displays the formula image rendered by LaTeX; Volunteers used digitizing pens to write on the electronic canvas; The collected data include: handwriting coordinate sequence and pressure data; Function support: Provides undo, redo, and eraser tools to simulate real writing experience; Step 1.2.2: Data annotation and quality control: Including three-level annotation system: Difficulty level criteria: Invalid writing errors, incomplete, or illegible handwriting; The number of simple characters is less than 20, and there is no complex structure; Difficulty containing complex symbols or illegible handwriting; Complex symbol definition: Complex characters: Chinese symbols, special symbols; Complex structures: fractions, matrices, calculus symbols, geometric symbols; Invalid data filtering: Line break detection: check whether the written line break is consistent with the formula structure; Manual review: Labelers verify the validity of the data to ensure that the labeling is accurate.
5. The method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation according to claim 1, characterized in that: The step 2 comprises the following steps: Step 2.1: Layout prediction stage; Step 2.2: Mask optimization stage; Step 2.3: Image generation stage.
6. The method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation according to claim 5, characterized in that: The step 2.1 includes the following steps: Input: LaTeX sequence; Processing: Use the GCN graph convolutional network to parse the formula structure, establish the topological relationship between symbols, and predict the bounding box position of each symbol; Output: A two-dimensional layout of the formula.
7. The method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation according to claim 5, characterized in that: The step 2.2 includes the following steps: Input: layout prediction results; Processing: Use the U-Net structure generator to convert the bounding box into a pixel-level soft mask to solve the layout ambiguity problem; Output: high-resolution mask map.
8. The method for constructing a handwritten mathematical formula dataset based on multimodal acquisition and semantic annotation according to claim 5, characterized in that: The step 2.3 includes the following steps: Input: optimized mask map; Processing: Generate realistic handwriting based on a diffusion model, simulating natural writing characteristics: connected strokes, jitter, pressure changes, and style diversity; Output: Final handwritten formula image.
9. A multi-source handwritten mathematical formula data set management and automated testing system, characterized in that: Including online collection and annotation module, data management module and automated testing module, The online collection and annotation module includes data collection processes and quality control rules. The data collection process includes user login, handwriting input, and intelligent annotation. User login supports multiple identities and hierarchical authority management. Handwriting input uses a passive pen to write input, records stroke trajectories in real time, and supports screenshots of writing areas, handwriting undo / redo, and canvas clearing functions. Intelligent annotation includes automatic generation of LaTeX sequences, user verification mechanism, and automatic classification and storage of data. The user verification mechanism: if the recognition is correct, the user confirms and submits it; if the recognition is incorrect, the user corrects the annotation. Data is automatically classified and stored into a correct sample library and an incorrect sample library. Quality control rules include validity judgment criteria and a graded annotation system. Valid data according to the validity judgment criteria are data with complete formula structure, legible handwriting, and conformity to LaTeX semantics; invalid data are data with illegal line breaks, illegible handwriting, and incomplete writing. The hierarchical annotation system includes the criteria and treatment methods for valid data, boundary data, and invalid data; The data management module includes platform function design and data return mechanism. The platform function design includes visual dashboard, data management and permission management. The visual dashboard includes real-time data statistics, collection progress monitoring and user contribution ranking. Data management includes multi-dimensional retrieval, batch export, and version control; permission management includes role classification and operation log auditing; data reflow mechanisms include error sample recovery and dynamic amplification strategies. Error sample recovery includes automatic classification and identification of error cases and adding them to the training set after human review. Dynamic amplification strategies include automatic detection of hotspot symbols and active learning sampling. The automated testing module includes test process design and continuous integration support. Test process design includes test task creation and performance evaluation. Test task creation supports batch uploading and customized test parameters. Performance evaluation includes core indicators and anomaly detection. Core indicators include recognition accuracy, processing latency, and peak memory usage. Anomaly detection includes classification of typical error patterns and automatic archiving of failure cases. Continuous integration support includes API interface and benchmark testing. API interface includes RESTful interface to support automated pipeline and automatic generation of test reports. Benchmark testing includes cross-model comparative testing and historical performance trend analysis.