Artificial Intelligence-Based Real-Time Recognition Method and System for Handwritten Formulas on a Whiteboard
By constructing high-density handwriting areas, combining handwriting pressure and inclination characteristics, using MobileNetV3-Small and Transformer decoder to generate LaTeX expressions, the missed and mis-checking problems of whiteboard system in dense writing areas are solved, and the accuracy of symbol detection and system efficiency are improved.
Patent Information
- Application Number
- CN202510509321.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-22
AI Technical Summary
The existing whiteboard system is prone to missed or missed in dense writing areas, and lacks the ability to distinguish similar symbols, resulting in delayed and inaccurate formula recognition.
By constructing high-density handwriting areas, combining handwriting pressure and inclination characteristics, the stroke features are extracted using the MobileNetV3-Small model, the symbol relationship diagram is constructed, and the LaTeX expression is generated using the Transformer decoder, the threshold value and adjacency matrix are dynamically adjusted, and the symbol detection is optimized.
It significantly improves the accuracy of symbol detection and the accuracy of LaTeX expressions, solves the problems of missed and missed detection in dense writing areas, reduces CPU/GPU load, and is suitable for low-power scenarios.
Smart Images

Figure CN120071364B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent whiteboards, and specifically, to a real-time recognition method and system for handwritten formulas on a whiteboard based on artificial intelligence. Background Art
[0002] An artificial intelligence whiteboard is an interactive writing platform integrating artificial intelligence technology, which can perceive, analyze, and process the user's handwritten input in real time. When writing mathematical formulas, physical equations, etc. on the whiteboard in class or remote teaching, the system can recognize and generate standard LaTeX code or visual mathematical expressions in real time. However, when writing complex formulas, there are often problems such as easy missed detection or misdetection in the dense writing area and insufficient ability to distinguish similar symbols.
[0003] At the same time, when writing quickly, the system has a recognition delay due to computing power limitations and cannot reconstruct a coherent path in time. In view of this, a real-time recognition method and system for handwritten formulas on a whiteboard based on artificial intelligence are provided. Summary of the Invention
[0004] The purpose of the present invention is to provide a real-time recognition method and system for handwritten formulas on a whiteboard based on artificial intelligence to solve the problems of easy missed detection or misdetection in the dense writing area and insufficient ability to distinguish similar symbols when writing complex formulas as mentioned in the above background art.
[0005] To achieve the above purpose, the present invention provides a real-time recognition method for handwritten formulas on a whiteboard based on artificial intelligence, including the following steps:
[0006] S1. Capture the sequence of touch points input by the user through the artificial intelligence whiteboard, and construct a handwritten path trajectory based on the sequence of touch points , and delimit the high-density handwriting area;
[0007] S2. Detect symbols in the high-density handwriting area to locate potential symbol areas, and extract the stroke feature vectors of the potential symbol areas through the MobileNetV3-Small model ;
[0008] S3. Use the extracted stroke feature vectors , and fuse the handwriting pressure and inclination features to construct a symbol relationship graph, and use the Transformer combined with the MLP classification module to dynamically correct the node connection weights of the symbol relationship graph for constructing the two-dimensional spatial relationship of superscripts, subscripts, fractions, and integral structures;
[0009] S4. Generate an initial LaTeX sequence through the Transformer decoder based on the node features and dynamic adjacency matrix in the symbol relationship graph, and perform dynamic programming on the initial LaTeX sequence through the CYK algorithm to convert the symbol relationship graph into a LaTeX expression.
[0010] As a further improvement of this technical solution, in S1, the specific steps involved in delimiting the high-density handwriting area are as follows:
[0011] S1.1. Combine the handwriting pressure with the inclination , and discretize the handwriting path trajectory into a dense point sequence ;
[0012] S1.2. Merge the original touch point sequence with the dense point sequence to form an enhanced point set , and calculate the grid density ;
[0013] S1.3. Delimit the high-density handwriting area through dynamic threshold segmentation using the grid density .
[0014] As a further improvement of this technical solution, in S1.3, the specific steps involved in delimiting the high-density handwriting area are as follows:
[0015] For all non-empty grid cells , calculate their global density mean ;
[0016] Based on the global density mean , calculate the standard deviation of the density distribution ;
[0017] Take the linear combination of the global density mean and the standard deviation as the dynamic density threshold ;
[0018] Traverse all grid cells . If , then mark the grid cell as a high-density area to form a candidate set ;
[0019] Model the high-density candidate set as a graph ;
[0020] Traverse the graph based on the breadth-first search algorithm, and merge adjacent high-density grids into connected regions ;
[0021] The process of traversing the graph using the repeated breadth - first search algorithm, starting from an unprocessed starting point each time to generate new connected regions until all high - density grids are visited, and finally generating a set of independent high - density handwriting regions . .
[0022] As a further improvement of this technical solution, in S2, the specific steps involved in symbol detection and positioning of potential symbol regions for high - density handwriting regions are as follows:
[0023] For each connected region in the set of high - density handwriting regions , calculate its center coordinates and the width and height of the circumscribed rectangle , and generate a set of multi - scale candidate boxes ;
[0024] Use the set of multi - scale candidate boxes as the input of YOLOv8 - Nano, and output the initial detection boxes by YOLOv8 - Nano;
[0025] Sort the initial detection boxes in descending order of confidence to obtain an ordered list ;
[0026] Calculate the average grid density covered by each detection box :
[0027] Among them,
[0028] In the formula, represents the grid cells into which the image is divided; represents the th initial detection box; represents the density value of the grid cell ;
[0029] Take the ratio of the local density mean to the global density mean as the NMS threshold adjustment factor, and based on the NMS threshold adjustment factor, adjust the initial NMS threshold to generate a dynamic NMS threshold ; The dynamic threshold is , where is the global density mean; In the formula, represents the dynamically adjusted NMS threshold;
[0030] If the detection box Intersection over Union with the detection box If the Intersection over Union is less than a certain value, then keep the detection box with a higher density mean, delete the other box, and repeat this step until all candidate boxes are processed; preferentially keep the detection results in the high-density area;
[0031] Build a Feature Pyramid Network in MobileNetV3-Small, and perform multi-scale symbolic feature extraction on the detection boxes in the high-density area through the Feature Pyramid Network, and output the final symbol position and category.
[0032] As a further improvement of this technical solution, in step S2, the specific steps involved in extracting the stroke feature vector of the potential symbol area through the MobileNetV3-Small model are as follows:
[0033] According to the detection box coordinates , crop the symbol area from the original image , and perform preprocessing to obtain ;
[0034] Introduce deformable convolution in the Bottleneck layer of MobileNetV3-Small to dynamically learn the convolution kernel offset and the direction angle , and obtain a feature map with an adaptive stroke direction ;
[0035] Use the feature map as the input of the local stroke attention module of MobileNetV3-Small, and add stroke direction weights in the local stroke attention module to obtain a feature map of the key stroke area , which is used to strengthen the feature response of the key stroke areas (such as intersection points and endpoints);
[0036] Perform cross-level feature fusion on the feature maps extracted from different Bottleneck layers to obtain a fused multi-scale feature map ;
[0037] Reduce the dimension of the multi-scale feature map to generate a stroke feature vector .
[0038] As a further improvement of this technical solution, in step S3, the specific steps involved in constructing a symbol relationship graph using the stroke feature vector are as follows:
[0039] Fuse the handwriting pressure and the tilt angle in the stroke feature vector Obtain the fused stroke feature vector based on the features ;
[0040] For the handwriting pressure Initial importance weight ;
[0041] Discretize the inclination angle into an 8 - direction encoding ;
[0042] For any two nodes, calculate the distance between the centers of their bounding rectangles , and define the dynamic adjacency threshold :
[0043]
[0044] In the formula, represents the dynamic adjacency threshold of the node pair and ; represents the weight coefficient for balancing the influence of the spatial scale represents the weight coefficient for balancing the influence of the feature difference represents the cosine similarity represents the width of the bounding rectangle of the th node; represents the height of the bounding rectangle of the th node; represents the width of the bounding rectangle of the th node; represents the height of the bounding rectangle of the
[0045] And calculate the inclination difference of the inclination angle ;
[0046] Based on the fused stroke feature vector , the direction encoding and the inclination difference , initialize the edge weight for each pair of connected nodes ;
[0047] Sum up the edge weights of all nodes to construct the adjacency matrix .
[0048] As a further improvement of this technical solution, use Transformer combined with the MLP classification module to dynamically correct the node connection weights of the symbol relationship graph. The specific steps involved are:
[0049] Before constructing the adjacency matrix perform self-attention encoding on the stroke feature vectors using a Transformer encoder to obtain context-sensitive representations ;
[0050] For any pair of nodes construct a joint feature vector ;
[0051] Introduce an MLP classification module to classify the joint feature vector and predict the correction factor ;
[0052] Based on the correction factor update the dynamic adjacency threshold , and finally determine whether to establish a connection in the adjacency matrix according to the updated dynamic adjacency threshold .
[0053] As a further improvement of this technical solution, in S4, for each node in the symbol relationship graph , predict the syntactic role using a multi-layer perceptron;
[0054] Based on the final adjacency matrix and the syntactic role , construct a syntactic structure tree;
[0055] Use a tree-shaped long short-term memory network to encode the syntactic structure tree from bottom to top;
[0056] The Transformer decoder generates a LaTeX sequence in an autoregressive manner, predicting the next token at each step based on the syntactic tree encoding and the attention mechanism;
[0057] During the decoding process of the Transformer decoder, introduce the relative position encoding between symbols to adjust the attention scores;
[0058] Convert the symbol relationship graph into a LaTeX expression through dynamic programming using the CYK algorithm on the initial LaTeX sequence.
[0059] On the other hand, the present invention provides an artificial intelligence-based real-time recognition system for whiteboard handwritten formulas, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the steps of the artificial intelligence-based real-time recognition method for whiteboard handwritten formulas described in any one of the above.
[0060] Compared with the prior art, the beneficial effects of the present invention are:
[0061] 1. In the real-time recognition method and system for whiteboard handwritten formulas based on artificial intelligence, dynamic threshold segmentation is adopted and combined with handwritten pressure and inclination information. By dynamically adjusting the sampling density and NMS threshold, the detection results of high-density regions are preferentially retained, and the response of key regions is enhanced by combining stroke direction features, significantly improving the accuracy of symbol detection and solving the problems of easy omission or misdetection in dense writing regions and insufficient ability to distinguish similar symbols in traditional methods.
[0062] 2. In the real-time recognition method and system for whiteboard handwritten formulas based on artificial intelligence, when constructing a symbol relationship graph, stroke features, pressure, and inclination are synchronously fused. Through dynamic adjacency threshold and context-sensitive weight correction, the relative position and direction between symbols are accurately modeled, improving the accuracy of the syntax structure tree and solving the problem that traditional methods are difficult to effectively capture the spatial relationships of structures such as superscripts, subscripts, fractions, and integrals.
[0063] 3. In the real-time recognition method and system for whiteboard handwritten formulas based on artificial intelligence, based on the CYK algorithm combined with a symbol space scoring function, and by enhancing Transformer decoding through relative position encoding, geometric constraints such as spatial distance and direction angle are incorporated into grammar rules, and the optimal parse tree is generated by dynamic programming to ensure the strict correspondence between LaTeX expressions and handwritten formulas;
[0064] At the same time, through lightweight network design and parallel architecture, the CPU / GPU load is significantly reduced while ensuring accuracy, which is suitable for low-power scenarios of intelligent whiteboards. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is the overall method flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0067] Embodiment 1: Please refer to Figure 1 As shown, this embodiment provides a real-time recognition method for whiteboard handwritten formulas based on artificial intelligence, including the following steps:
[0068] S1. Construct a handwritten trajectory by capturing the touch point sequence input by the user through an artificial intelligence whiteboard, and delimit a high-density handwriting area;
[0069] Specifically, in this embodiment, the AI whiteboard records the user's input in the form of timestamps and coordinate points, so the touch point sequence is a data stream containing time and coordinates, and its touch point sequence is specifically:
[0070] ;
[0071] Among them, represents the coordinates in the coordinate system of the AI whiteboard; represents the timestamp of the th touch point, and satisfies ; represents the total number of touch points within a single stroke;
[0072] The touch point sequence captures the user's writing action through high-frequency sampling (100Hz) to ensure that the point density is consistent with the real handwriting;
[0073] Then, based on the touch point sequence a handwritten path trajectory is constructed. The specific steps involved are:
[0074] Perform Gaussian filtering and noise reduction on the original touch point sequence to eliminate the touch screen jitter noise and generate a smooth touch point sequence ;
[0075] Among them,
[0076]
[0077] and ;
[0078] In the formula, represents the th touch point after filtering; represents the th point in the original touch point sequence ; is the summation variable, representing the traversal range of the touch point index; represents the standard deviation of the Gaussian function; represents the normalization coefficient of the Gaussian function; represents the exponential decay term of the Gaussian kernel;
[0079] Since large-scale interpolation may increase the CPU load, therefore, the Catmull-Rom interpolation algorithm with dynamic density control is used to prevent the CPU load from being too high:
[0080] Define the interpolation step size adaptive adjustment strategy:
[0081]
[0082] In the formula, represents the reference interpolation interval; , represents the average writing speed; , represents the speed sensitivity coefficient; represents the th interpolation step of the th touch point;
[0083] The interpolation step adaptive adjustment strategy is used to reduce the interpolation density in the high-speed writing area and increase the interpolation density in the low-speed area, so as to achieve dynamic balance of the calculation load;
[0084] For every four consecutive touch points an independent spline segment is constructed:
[0085]
[0086] In the formula, represents the Catmull-Rom basis function (cubic polynomial);
[0087] is the normalized time parameter; represents the th independent spline segment constructed from four consecutive touch points;
[0088] The interpolation calculations of each segment are independent of each other and meet the parallelization conditions;
[0089] Furthermore, a parallel computing architecture is adopted to implement real-time interpolation calculation, which is used to reduce the calculation time consumption;
[0090] Specifically, the touch point sequence is divided into processing units, each unit contains 4 consecutive touch points, and is assigned to an independent computing thread;
[0091] The interpolation operation is executed in parallel by using a compute shader, and real-time interpolation calculation is realized through the GPU parallel pipeline;
[0092] At the same time, at the connection of segments a continuity constraint is imposed, and by constraining the continuity of the first-order derivatives of adjacent segments, the connection mutation caused by parallel segmentation is eliminated:
[0093] Among them, the continuity constraint is specifically:
[0094]
[0095] In the formula: represents the spline segment at The left derivative at a moment; Denote the spline segment At The right derivative at a moment;
[0096] Construct a circular buffer to store the interpolation results of adjacent segments, and use the double buffering technique to achieve lock-free data synchronization to ensure the spatio-temporal continuity of the trajectory output.
[0097] In this embodiment, the specific steps involved in delimiting the high-density handwriting area are as follows:
[0098] S1.1. Combine the handwriting pressure And the inclination angle , and discretize the handwriting path trajectory Into a dense point sequence :
[0099] Among them, for the trajectory point Generated by interpolation, its corresponding handwriting pressure is 、The inclination angle is ; Among them, Represents the trajectory point coordinates, calculated by the interpolation algorithm, Represents the th The coordinate value of the sampling point on the spline segment;
[0100] When discretely sampling each spline segment, the trajectory point Generated by interpolation is the function At time The value at;
[0101] On the normalized time parameter Sampled at a step size of To obtain a dense point sequence :
[0102] , ;
[0103] In the formula, Denote the spline segment The start time of; Denote the end time of the spline segment; Denote the sampling time interval; Denote the number of sampling points of the spline segment;
[0104] Dynamically adjust the sampling density through the handwriting pressure And the inclination angle To retain the handwriting details;
[0105] Merge the sampling points of all spline segments into a discretized point sequence:
[0106]
[0107] wherein, represents the total number of spline segments; represents the set of all discretized point sequences; represents the last discretized sampling point of the -th spline segment; represents the number of sampling points on the -th spline segment minus 1. If a spline segment is discretized into 5 points, then , and the indices are 0, 1, 2, 3, 4;
[0108] Furthermore, the writing pressure can reflect the weight of the stroke, and the inclination angle can indicate the writing direction. By fusing these physical features, the system can more accurately distinguish symbols with similar visual forms but different writing dynamics. A stroke with a greater pressure may correspond to the starting stroke of a bold symbol or an integral symbol, and the change in the inclination angle can distinguish the fractional horizontal line (horizontal inclination angle) from italic letters (tilt angle);
[0109] S1.2. Merge the original touch point sequence and the dense point sequence to form an enhanced point set , and calculate the grid density :
[0110] The original touch point sequence Specifically:
[0111] ;
[0112] The trajectory point sequence generated by the interpolation algorithm (cubic spline interpolation):
[0113]
[0114] wherein, represents the interpolated abscissa; represents the interpolated ordinate; represents the interpolated pressure value; represents the -th original touch point; represents the abscissa of the -th original touch point; represents the ordinate of the -th original touch point; represents the timestamp of the -th original touch point; represents the inclination angle of the interpolation point (the tilt angle of the pen, also used for dynamic sampling); represents an index variable; represents the total number of original touch points;
[0115] Then the enhanced point set is:
[0116]
[0117] In the formula, represents the set of all dense point sequences; , represents the sampling time step after dynamic adjustment; represents the original touch point sequence in the th point timestamp;
[0118] In this embodiment, through secondary sampling, the point density is increased to the 400 dpi level, and the spatial resolution of each grid cell is aligned with the target point density of 400 dpi;
[0119] Furthermore, define the quantization grid
[0120] where inches, representing the physical size of the grid cell; represents the grid index, covering the effective writing area of the whiteboard; represents the width of the grid cell; represents the height of the grid cell; represents the row index of the grid ( axis direction), covering the effective writing area; represents the row index of the grid ( axis direction), covering the effective writing area;
[0121] Then for each grid cell , calculate the touch point coverage:
[0122] For each point in the enhanced point set , then the division basis of the grid cell is:
[0123] ,
[0124] In the formula, represents the timestamp;
[0125] Grid density is defined as the number of points falling into the grid cell :
[0126]
[0127] In the formula, represents the indicator function, which counts the frequency of touch points falling into the grid; when belongs to it takes the value of 1, otherwise it takes 0; represents the coordinate component of point ; represents the total number of points;
[0128] By statistically analyzing the distribution of the grid density , it is detected whether the touch points evenly cover or there are blind areas.
[0129] S1.3. Divide the grid density by dynamic threshold segmentation to delimit the high-density handwriting area;
[0130] Define the dynamic density threshold of the high-density handwriting area:
[0131]
[0132] Among them, for the grid cell , its global density mean value is statistically analyzed:
[0133]
[0134]
[0135] Among them, is the total number of rows and columns for quantifying the grid; represents the grid cell 's density value;
[0136] In the formula, is the adjustment coefficient, and the empirical value range is , which is used to dynamically adapt to the writing density fluctuation and adjust the value through fast writing (low density) and fine writing (high density); is the global density mean, representing the average value of the densities of all grid cells; is the global density standard deviation, which is used to measure the dispersion degree of the grid density; represents the total number of rows of the quantified grid (the number of grid cells divided in the horizontal direction); represents the total number of columns of the quantified grid (the number of grid cells divided in the vertical direction); represents the point density within the grid cell (that is, the total number of touch points and interpolation points falling into this grid); represents the The grid cell in the th row and
[0137] Traverse all grid cells and determine whether it is a high-density area:
[0138] If , mark the grid cell as a high-density cell to form a candidate set ;
[0139] Construct a graph model:
[0140] Model the high-density candidate set as a graph where the vertex set , and the edge set represents the adjacent grid relationship (using the 8-neighborhood connection rule). If two grid cells are adjacent in the 8-neighborhood (sharing an edge or a corner), there is an edge between their corresponding vertices;
[0141] Traverse the graph through the breadth-first search (BFS) algorithm and merge adjacent high-density grids into connected regions :
[0142]
[0143] where is the th connected component, containing the index set of all grid cells belonging to the same connected region ; represents the th independent high-density handwriting region, generated by merging all grid cells within the connected component ; represents the serial number index of the connected component;
[0144] Repeat the process of traversing the graph through the breadth-first search (BFS) algorithm starting from an unprocessed starting point each time to generate new connected regions until all high-density grids have been visited, and finally obtain independent high-density handwriting region sets , where , provides input for subsequent formula recognition, and each region represents a connected block of high-density handwriting;
[0145] represents the final output set of independent high-density regions; represents the total number of independent high-density regions (i.e., the number of connected components); represents the th independent high-density handwriting region 。
[0146] S2. Use YOLOv8 - Nano to perform symbol detection on the high - density handwriting area to locate potential symbol areas, and extract the stroke feature vectors of the potential symbol areas through the MobileNetV3 - Small model;
[0147] Specifically, the specific steps involved in performing symbol detection on the high - density handwriting area to locate potential symbol areas are as follows:
[0148] For each connected region in the high - density handwriting area set , calculate its center coordinates and the width and height of the circumscribed rectangle ;
[0149] Generate a multi - scale candidate box set , where the scale factor , and the candidate box coordinates are:
[0150]
[0151] In the formula, represents the abscissa of the center of the connected region ; represents the ordinate of the center of the connected region ; represents; represents the width of the circumscribed rectangle of the connected region ; represents the height of the circumscribed rectangle of the connected region ; represents the candidate box of the th scale generated based on the connected region ;
[0152] Input the candidate boxes into the YOLOv8 - Nano network to output the initial detection boxes and their confidence levels;
[0153] Sort the initial detection boxes in descending order of confidence level and denote them as ;
[0154] Calculate the average grid density covered by each detection box , where in the formula, represents the grid cells into which the image is divided; represents the th initial detection box; represents the density value of the grid cell ;
[0155] Take the local average density The ratio with the global density mean is used as the NMS threshold adjustment factor, and based on the NMS threshold adjustment factor, the initial NMS threshold is used to generate a dynamic NMS threshold ; the dynamic threshold is , where is the global density mean; in the formula, represents the dynamically adjusted NMS threshold;
[0156] If the intersection over union (IoU) of the detection box and is , then the detection box with a higher density mean is retained, and the other box is deleted. Repeat this step until all candidate boxes are processed; preferentially retain the detection results in the high-density area, and retain more candidate boxes in the high-density area to improve the symbol detection recall rate;
[0157] Specifically, the improved density-weighted (NMS) algorithm solves the problems of over-suppression or missed detection caused by uneven regional density in traditional fixed-threshold NMS in complex handwritten formulas by dynamically adjusting the threshold ;
[0158] In the densely written area , increase the threshold to allow boxes with higher overlap to coexist and avoid misdeleting adjacent symbols (such as superscripts and subscripts);
[0159] In the sparse area , decrease the threshold to reduce missed detections;
[0160] Construct a Feature Pyramid Network (FPN) in MobileNetV3-Small to fuse the low-level feature map and the high-level feature map , :
[0161] Detect small-scale symbols at the layer and detect large-scale symbols at the layer. Extract multi-scale symbol features through the Feature Pyramid Network, that is, perform multi-scale symbol feature extraction on the area containing symbols (i.e., the detection boxes in the high-density area) through the Feature Pyramid Network, and output the final symbol position and category.
[0162] Specifically, take the output feature maps (shallow layer, high resolution), (middle layer), (deep layer, low resolution) of the backbone network of MobileNetV3-Small as the input;
[0163] Through the Feature Pyramid Network for , , Perform cross-layer fusion and finally output a multi-scale feature pyramid; , and As outputs, corresponding to different detection scales respectively. On each level (P3, P4, P5), deploy the YOLOv8-Nano detection head to predict the symbols corresponding to the scales respectively:
[0164] Perform 3×3 convolution on the deep features to reduce the dimension and generate the initial :
[0165]
[0166] Upsample P5 to the size of C4 and add it element-wise to the 1×1 convolution result of C4 to generate :
[0167]
[0168] Upsample P4 to the size of C3 and add it element-wise to the 1×1 convolution result of C3 to generate :
[0169]
[0170] Then the final feature pyramid:
[0171] : The highest resolution (such as 1 / 8), used to detect small symbols (such as dots, apostrophes, commas);
[0172] : Medium resolution (such as 1 / 16), used to detect regular symbols (such as letters, numbers);
[0173] : The lowest resolution (such as 1 / 32), used to detect large symbols (such as integral signs, fraction bars).
[0174] The YOLOv8-Nano detection head predicts the symbol positions and categories on P3, P4, and P5. Each detection head generates detection results through convolution operations. The YOLOv8-Nano detection head aggregates the prediction results of all levels. Finally, the detection results are optimized by an improved density-weighted (NMS) algorithm to output the final symbol list.
[0175] Specifically, the dynamic threshold adjustment based on the local / global density ratio, compared with the traditional NMS using a fixed threshold, preferentially retains the results in the high-confidence and high-density regions, overcoming the problem of easy missed detection in dense regions of the traditional method;
[0176] And introduce the handwriting grid density , by calculating the average density of the area covered by the detection box, quantify the complexity of the symbol distribution; for example, there may be dense superscripts and subscripts around the integral symbol, and at this time, the dynamic threshold can avoid misdeleting the key detection box.
[0177] In this embodiment, the specific steps involved in extracting the stroke feature vector of the potential symbol area through the MobileNetV3-Small model are as follows:
[0178] According to the coordinates of the detection box , and crop the symbol area from the original image , and then perform preprocessing to obtain :
[0179] Scale to a fixed size (such as 64×64 pixels) to generate a standardized image ,
[0180] Convert to a single-channel grayscale image and perform pixel value normalization:
[0181]
[0182] In the formula, represents the mean; represents the standard deviation; represents the standardized representation of the symbol area and serves as the only input to the feature extraction network;
[0183] Aiming at the stroke characteristics (direction, curvature, continuity) of handwritten symbols, introduce direction-sensitive convolution and local attention mechanism in MobileNetV3-Small to enhance the ability to capture stroke details:
[0184] Through MobileNetV3-Small, perform feature extraction on , and introduce deformable convolution in the Bottleneck layer of MobileNetV3-Small to dynamically learn the offset of the convolution kernel and the direction angle , to obtain a feature map with an adaptive stroke direction :
[0185] The network levels of MobileNetV3-Small include an initial convolution layer, a Bottleneck layer, a local stroke attention module, multi-level feature fusion, and global feature descriptor generation;
[0186] Among them, the initial convolution layer is used to perform operations on Perform preliminary feature extraction to generate low-level feature maps :
[0187]
[0188] Wherein, represents the low-level feature map output by the initial convolutional layer;
[0189] Bottleneck layer:
[0190] Introduce deformable convolution in the Bottleneck layer:
[0191]
[0192] The input is the low-level feature map generated by the previous layer ;
[0193] Wherein, represents the initial offset position of the fixed convolution kernel; represents the spatial position coordinates on the feature map; represents the learnable offset; represents the direction angle parameter, which is optimized by gradient descent to make the convolution kernel adapt to the stroke direction; represents the feature map with an adaptive stroke direction output by the deformable convolution, that is, the output of the Bottleneck layer; represents the total number of convolution kernels; represents the index variable; represents the weight of the convolution kernel, a parameter automatically learned through model training, used for weighted summation of different positions on the input feature map; represents the input of the Bottleneck layer, that is, the low-level feature map ;
[0194] Local stroke attention module:
[0195] Increase the stroke direction weight
[0196] Then the output of the local stroke attention module is:
[0197]
[0198] Wherein, represents global average pooling; represents convolution operation; represents the Sigmoid activation function, used to map the weight value to the interval [0,1]; represents the stroke direction attention weight matrix; represents element-wise multiplication; Indicates the feature map enhanced by attention;
[0199] Multi-level feature fusion:
[0200] Extract feature maps from the Bottleneck3, 6, and 12 layers of MobileNetV3-Small , and perform fusion:
[0201]
[0202] In the formula, Indicates the output of the 3rd Bottleneck layer; Indicates the output of the 6th Bottleneck layer; Indicates the output of the 12th Bottleneck layer; Indicates the upsampling operation (such as bilinear interpolation); Indicates the fused multi-scale feature map; Indicates element-wise addition for feature fusion;
[0203] Global feature descriptor generation:
[0204]
[0205] In the formula, Indicates global average pooling; Indicates a 512-dimensional fully connected layer; Indicates the unnormalized stroke feature vector;
[0206] Perform L2 normalization on the feature vector to generate the final stroke feature vector :
[0207]
[0208] In the formula, Indicates the L2 norm for calculating the magnitude of the vector; Indicates the normalized stroke feature vector;
[0209] In this embodiment, a bidirectional feature pyramid is constructed in MobileNetV3-Small to fuse semantic features at different levels; for example, shallow features Retain stroke details (such as dots, flicks), and deep features Capture structural information (such as fractions, square roots), solving the problem of large scale differences in handwritten formulas;
[0210] Embed a local stroke attention module in the Feature Pyramid Network (FPN), and add stroke direction weights in the local stroke attention module , preferentially focus on key areas such as stroke intersections and symbol junctions; for example, it can effectively distinguish between and morphological differences;
[0211] S3. Use the extracted stroke feature vectors to construct a symbol relationship graph, and dynamically correct the node connection weights of the symbol relationship graph based on a graph neural network or Transformer, for constructing the two-dimensional spatial relationship of superscripts, subscripts, and fractional and integral structures;
[0212] In this embodiment, the specific steps involved in using the stroke feature vectors to construct a symbol relationship graph are as follows:
[0213] Fuse the handwriting pressure and inclination features in the stroke feature vectors to obtain the fused stroke feature vectors ; the handwriting pressure and inclination serve as supplementary dimensions of the stroke feature vectors to enhance the modeling ability of the physical characteristics of the writing tool (such as the inclination angle of the chalk and the pressure change of the pen), so as to distinguish symbols with similar shapes but different writing dynamics (such as "6" and "9"). Among them, the handwriting pressure and inclination need to be normalized to ensure the scale consistency of the feature vectors and avoid weight deviation caused by differences in physical dimensions, , where represents the concatenation operation along the feature dimension;
[0214] Based on the fused stroke feature vectors construct an adjacency matrix :
[0215] For each node , define the node feature vector :
[0216]
[0217] And for the feature vector of the node , it is concatenated by the stroke feature vectors fusing the handwriting pressure and inclination ;
[0218] Initialize the importance weight for each node , and use the Sigmoid function to normalize the handwriting pressure :
[0219]
[0220] and discretize the dip angle into an 8 - direction encoding:
[0221]
[0222] where, represents the handwritten pressure of the -th node; represents the Sigmoid function, which is used to map the pressure value to a normalized weight range; represents the initial importance weight of the -th node, reflecting the force information of this node during writing; represents the complete feature vector of the -th node that integrates semantic and physical attributes, and its total dimension is ; represents the nib dip angle information of the -th node; represents discretizing the continuous dip angle , into an 8 - direction encoding (such as up, down, left, right and their diagonal directions), so it is represented by an 8 - dimensional vector;
[0223] For any two nodes, the distance between the centers of their circumscribed rectangles is defined as:
[0224]
[0225] where, represents the Euclidean distance between the centers of the circumscribed matrices corresponding to node and node ; represents the coordinates of the center of the rectangle of the -th node;
[0226] Define the dynamic adjacency threshold according to the node size and feature difference:
[0227]
[0228] where, represents the dynamic adjacency threshold of node pair and ; represents the weight coefficient used to balance the influence of spatial scale; represents the weight coefficient used to balance the influence of feature difference; represents the larger one in size between node and node , which is used to characterize the spatial scale; represents the cosine similarity, which is used to measure the fusion feature vectors of two nodes between scales; represents the width of the circumscribed rectangle of the represents the height of the circumscribed rectangle of the represents the width of the circumscribed rectangle of the represents the height of the circumscribed rectangle of the
[0229] If , then a basic connection is established between node and node ;
[0230] And calculate the inclination angle difference of the inclination angle :
[0231]
[0232] If , it is considered that the writing directions are the same, and the weight of the edge is enhanced; represents the nib inclination angle of the th node; represents the nib inclination angle of the th node; represents the absolute difference in the nib inclination angles between node and node . If this difference is less than the set threshold , it is considered that the writing directions of the two are the same, and the edge weight is enhanced accordingly;
[0233] Based on the fused stroke feature vector , direction encoding and inclination angle difference , for each pair of connected nodes initialize the edge weight :
[0234]
[0235] In the formula, represents the initial weight between node and node ; represents concatenating the fusion feature vectors of two nodes to obtain a joint feature representation of a higher dimension; represents from node to node The combined direction encoding vector generated from the direction encoding (or the direction information of the relationship between the two) reflects the relative direction between them (such as "upper right", "lower left", etc.); Denotes element-wise addition, which is used to fuse feature information from different sources (combined features, direction encoding, and inclination difference); Denotes a learnable weight matrix, which is used to linearly transform the fused features to generate appropriate edge weights; Denotes the Sigmoid activation function;
[0236] Sum up the edge weights of all nodes to construct the adjacency matrix (where is the number of nodes), which is defined as:
[0237] If the spatial connection condition or the feature similarity condition is satisfied, then , otherwise ;
[0238] Among them, the spatial connection condition is , and the feature similarity condition is ;
[0239] As mentioned above, denotes whether a connection is established between node and node and the weight of the connection; denotes the feature similarity threshold. When the similarity is greater than this threshold, it is considered that the two are similar enough in features to establish a connection; denotes the cosine similarity between the fused feature vectors of node and node , which is used to measure whether they are similar semantically.
[0240] The Transformer dynamically corrects the node connection weights of the symbol relationship graph. The specific steps involved are:
[0241] Before constructing the adjacency matrix , perform self-attention encoding on the node fused features based on the Transformer encoder to obtain a context-sensitive representation ;
[0242]
[0243] In the formula, denotes the context-sensitive representation obtained by the th node after the self-attention mechanism, which is the result of weighted summation of all node information; denotes the query vector of the th node; Denotes the dimension of the key vector (or query vector), which is used as a scaling factor to prevent the dot product result from being too large and thus affecting the gradient; Denotes the number of elements in the input sequence, i.e., the total number of nodes; Denotes the value vector of the th node; Denotes the key vector of the th node, Denotes the transpose of the key vector of the th node;
[0244] For any pair of nodes Construct the joint feature vector :
[0245] For each pair of nodes and , construct the joint feature vector
[0246]
[0247] where Denotes the context-sensitive feature vector of node after being encoded by the Transformer self-attention mechanism; Denotes the context-sensitive feature vector of node after being encoded by the Transformer self-attention mechanism; Denotes the dimension of the context-sensitive feature vector of each node ; Denotes the element-wise product of the feature vectors of node and node ; Denotes the element-wise difference of the feature vectors of node and node ;
[0248] Introduce an MLP classification module to classify the joint feature vector and predict the correction factor ;
[0249] Use a multi-layer perceptron (MLP) as the classification module to process the joint feature vector and output the correction factor corresponding to the pair of nodes :
[0250]
[0251] where Represents the weight matrix of the first layer of the MLP; Is the bias vector of the first layer of the MLP; Is the bias scalar of the second layer of the MLP; Is the weight matrix of the second layer of the MLP, used to map the hidden layer output to the correction factor ; Represents the non-linear activation function; Represents the Sigmoid function;
[0252] Based on the correction factor Update the dynamic adjacency threshold , and finally, based on the updated dynamic adjacency threshold Determine whether to establish a connection in the adjacency matrix ;
[0253] Combined with the strategy of context information correction, the updated threshold is:
[0254]
[0255] Among them, Represents the adjustable weight factor, used to balance the contributions of the original threshold and the correction factor; Represents the updated dynamic adjacency threshold;
[0256] Finally, according to the distance between the centers of the circumscribed rectangles of the node pairs , construct the final adjacency matrix :
[0257] Then, after optimization, if the spatial connection condition or the feature similarity condition is satisfied, then , otherwise ;
[0258] Among them, the spatial connection condition is , and the feature similarity condition is ; Represents the feature similarity threshold, used to judge the semantic association strength between nodes.
[0259] In this embodiment, the final adjacency matrix Encodes all the structural information of the symbolic relationship graph, and its dynamic nature enables the symbolic relationship graph to adapt to the writing changes of complex formulas. The adjacency matrix Is the mathematical expression of the symbolic relationship graph;
[0260] The symbolic relationship graph provides a directly parsable topological structure for subsequent syntax tree construction and LaTeX generation.
[0261] S4. Based on the node features and dynamic adjacency matrix in the symbol relationship graph, generate an initial LaTeX sequence through a Transformer decoder, and perform dynamic programming on the initial LaTeX sequence through the CYK algorithm to convert the symbol relationship graph into a LaTeX expression;
[0262] In this embodiment, each node in the symbol relationship graph , predict the syntactic role through a multi-layer perceptron (MLP) :
[0263]
[0264] where represents the predicted syntactic role of node ; is a multi-layer perceptron;
[0265] Based on the final adjacency matrix and syntactic role , construct a syntactic structure tree:
[0266] Based on the updated dynamic adjacency threshold , determine the parent-child relationship through breadth-first search;
[0267] Use a tree-shaped long short-term memory network to encode the syntactic structure tree from bottom to top:
[0268] For each node of the syntax tree , aggregate the features of the child nodes to generate a hidden state :
[0269]
[0270] where represents the LSTM unit of the tree-shaped long short-term memory network, aggregating the information of the subtree from bottom to top; represents the hidden state of node , encoding its syntactic role and the features of the child nodes; represents the hidden state of the child node ;
[0271] The Transformer decoder generates a Token sequence based on the syntax tree encoding ;
[0272] At the same time, introduce the relative coordinate difference between symbols , map it to an embedding vector and add it to the attention score:
[0273]
[0274] Among them, represents the horizontal coordinate difference between node and ; represents the vertical coordinate difference between node and ;
[0275] In the formula, represents the "attention demand" of the target position (such as the LaTeX token being decoded) that needs to be generated currently for the input symbol; represents the "feature index" of the input symbol (the node in the symbol relationship graph); represents the actual feature information of the input symbol, which is used to generate the target output after weighted aggregation; represents the core input matrix for generating the attention score; represents calculating the similarity between the query and the key (the original attention weight); represents a learnable function that maps the coordinate difference to an attention bias term; represents the dimension of the key vector;
[0276] The attention score is the weight value obtained through the similarity calculation of and , indicating the importance of the input symbol to the current target position;
[0277] Furthermore, during the decoding process of the Transformer decoder, relative position encoding between symbols is introduced to adjust the attention score:
[0278] Encode spatial constraints (distance, direction, adjacency weight) as grammar rule scores;
[0279] For each candidate rule , define a spatial score function:
[0280]
[0281] In the formula, represents the Euclidean distance between symbol and ; represents the direction angle between symbol and ; represents the ideal direction expected by rule ; represents the edge weight between and in the adjacency matrix; represents the distance decay coefficient, which controls the spatial sensitivity; represents the parent node symbol; all represent candidate child symbols;
[0282] In the CYK algorithm, for the dynamic programming update rule, each cell has a weight of:
[0283]
[0284] In the formula, represents the probability of the grammar rule; represents the spatial score, which is used to amplify the probability of the rule that conforms to the spatial layout; represents covering the input position to of the non-terminal with the maximum score, represents the starting position of the input sequence (the starting index of the symbol in the symbol relationship graph), represents the ending position of the input sequence (the ending index of the symbol in the symbol relationship graph); represents the splitting point; represents from the position to the maximum score that the subsequence can be generated by the non-terminal ; and both represent non-terminals in the context-free grammar (CFG) (such as expressions, terms, factors, etc.);
[0285] Through the CYK algorithm, dynamic programming is performed on the initial LaTeX sequence to convert the symbol relationship graph into a LaTeX expression. The specific expressions involved are:
[0286]
[0287] In the formula, represents the parse tree that conforms to the grammar rules and spatial constraints; represents the logarithmic probability of the grammar rule; represents the logarithmic term of the spatial score, which is used to amplify the reasonable layout rules;
[0288] Embodiment 2: This embodiment provides an artificial intelligence-based real-time recognition system for whiteboard handwritten formulas, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the steps of the artificial intelligence-based real-time recognition method for whiteboard handwritten formulas described in any one of the above.
[0289] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, various changes and improvements will occur to the present invention, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A real-time recognition method for handwritten formulas on a whiteboard based on artificial intelligence, characterized in that, Including the following steps: S1. Capture the touch point sequence input by the user through the artificial intelligence whiteboard, construct the handwritten path trajectory based on the touch point sequence , and delimit the high-density handwriting area; S2. Detect symbols in the high-density handwriting area to locate potential symbol areas, and extract the stroke feature vectors of the potential symbol areas through the MobileNetV3-Small model ; S3. Use the extracted stroke feature vectors , and fuse the handwriting pressure and inclination features to construct a symbol relationship graph, and use Transformer combined with an MLP classification module to dynamically correct the node connection weights of the symbol relationship graph; Among them, the stroke feature vector is used The specific steps involved in constructing the symbol relationship graph are as follows: Fuse the handwriting pressure and tilt angle features in the stroke feature vector to obtain the fused stroke feature vector ; For handwriting pressure Initial importance weight ; Discretize the dip angle into an 8-direction encoding ; For any two nodes, calculate the distance between the centers of their circumscribed rectangles , and define the dynamic adjacency threshold : In the formula, represents the dynamic adjacency threshold of the node pair and ; represents the weight coefficient for balancing the influence of the spatial scale; represents the weight coefficient for balancing the influence of the feature difference; represents the cosine similarity; represents the -th width of the circumscribed rectangle of the -th node; represents the -th height of the circumscribed rectangle of the represents the -th width of the circumscribed rectangle of the represents the -th height of the circumscribed rectangle of the And calculate the inclination angle The inclination angle difference ; Based on the fused stroke feature vectors , direction encoding and dip angle difference , for each pair of connected nodes initialize the edge weights ; Sum up the edge weights of all nodes to construct the adjacency matrix ; S4. Based on the node features and dynamic adjacency matrix in the symbol relationship graph, generate an initial LaTeX sequence through a Transformer decoder, and perform dynamic programming on the initial LaTeX sequence through the CYK algorithm to convert the symbol relationship graph into a LaTeX expression.
2. The real-time recognition method for handwritten formulas on a whiteboard based on artificial intelligence according to claim 1, characterized in that: In the said S1, the specific steps involved in delimiting the high-density handwriting area are: S1.
1. Combine the handwriting pressure with the inclination angle , and discretize the handwriting path trajectory into a dense point sequence ; S1.
2. Merge the original touch point sequence with the dense point sequence to form an enhanced point set , and calculate the grid density ; S1.
3. Divide the grid density Define the high-density handwriting area through dynamic threshold segmentation.
3. The real-time recognition method for handwritten formulas on a whiteboard based on artificial intelligence according to claim 2, characterized in that: In the said S1.3, the specific steps involved in delimiting the high-density handwriting area are: For all non-empty grid cells , calculate their global density means ; Based on the global density mean Calculate the standard deviation of the density distribution ; Take the linear combination of the global density mean and the standard deviation as the dynamic density threshold ; Traverse all grid cells , if , then mark the grid cell as a high-density area to form a candidate set ; Model the high-density candidate set as a graph ; Traverse the graph based on the breadth-first search algorithm , and merge adjacent high-density grids into connected regions ; The process of traversing a graph using the repeated breadth-first search algorithm, starting from an unprocessed starting point each time to generate new connected regions until all high-density grids are visited, and finally generating a set of independent high-density handwriting regions by generating new connected regions until all high-density grids are visited, and finally generating .
4. The real-time recognition method for handwritten formulas on a whiteboard based on artificial intelligence according to claim 3, characterized in that: In the said S2, the specific steps involved in symbol detection in the high-density handwriting area to locate potential symbol areas are: For each connected region in the set of high-density handwriting regions , calculate its center coordinates and the width and height of its bounding rectangle to generate a set of multi-scale candidate boxes ; The multi-scale candidate box set is used as the input of YOLOv8-Nano, and the initial detection boxes are output by YOLOv8-Nano; Sort the initial detection boxes in descending order of confidence to obtain an ordered list ; Calculate each detection box Average grid density covered ; Take the ratio of the local density mean to the global density mean as the NMS threshold adjustment factor, and generate a dynamic NMS threshold based on the NMS threshold adjustment factor and the initial NMS threshold ; ; If the detection box and the detection box have an intersection over union , then retain the detection box with a higher density mean, delete the other box, and repeat this step until all candidate boxes are processed; Construct a feature pyramid network in MobileNetV3-Small, and perform multi-scale symbol feature extraction on the detection boxes in the high-density area through the feature pyramid network to output the final symbol positions and categories.
5. The real-time recognition method for handwritten formulas on a whiteboard based on artificial intelligence according to claim 4, characterized in that: In the said S2, the specific steps involved in extracting the stroke feature vectors of potential symbol areas through the MobileNetV3-Small model are: According to the detection box coordinates , and crop the symbol area from the original image , and then perform preprocessing to obtain ; Introduce deformable convolution in the Bottleneck layer of MobileNetV3-Small to dynamically learn the offset of the convolution kernel and the direction angle to obtain a feature map with an adaptive stroke direction ; Take the feature map as the input of the local stroke attention module of MobileNetV3-Small, and add stroke direction weights in the local stroke attention module to obtain the feature map of the key stroke area ; Cross-level feature fusion is performed on the feature maps extracted from different Bottleneck layers to obtain the fused multi-scale feature maps ; Reduce the dimensionality of the multi-scale feature map to generate a stroke feature vector .
6. The real-time recognition method for handwritten formulas on a whiteboard based on artificial intelligence according to claim 1, characterized in that: Adopt Transformer combined with an MLP classification module to dynamically correct the node connection weights of the symbol relationship graph. The specific steps involved are: Before constructing the adjacency matrix self-attention encoding is performed on the stroke feature vectors based on the Transformer encoder to obtain a context-sensitive representation ; For any node pair Construct a combined feature vector ; Introduce the MLP classification module to classify the joint feature vector and obtain the prediction correction factor ; Based on the correction factor Update the dynamic adjacency threshold , and finally, based on the updated dynamic adjacency threshold Determine whether to establish a connection in the adjacency matrix or not.
7. The real-time recognition method for handwritten formulas on a whiteboard based on artificial intelligence according to claim 6, characterized in that: In S4, for each node in the symbol relationship diagram , predict the syntactic role based on a multi-layer perceptron ; Based on the final adjacency matrix and syntactic roles , construct a syntactic structure tree; Use a tree-shaped long short-term memory network to encode the syntax structure tree from bottom to top; The Transformer decoder generates a LaTeX sequence in an autoregressive manner, predicting the next Token at each step based on the syntax tree encoding and the attention mechanism; During the decoding process of the Transformer decoder, introduce the relative position encoding between symbols to adjust the attention scores; Perform dynamic programming on the initial LaTeX sequence through the CYK algorithm to convert the symbol relationship graph into a LaTeX expression.
8. A real-time recognition system for whiteboard handwritten formulas based on artificial intelligence, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes a computer program to implement the steps of the artificial intelligence-based real-time recognition method for whiteboard handwritten formulas as described in any one of claims 1-7.
Citation Information
Patent Citations
Control method for touch writing acceleration under Android system
CN110737364A
Formula identification method and device and device for formula identification
CN113408417A