White board handwritten formula real-time identification method and system based on artificial intelligence

The touch point sequence input by the user is captured through the artificial intelligence whiteboard, the handwriting path trajectory is constructed and the high-density handwriting area is demarcated. The stroke feature vector is extracted using the MobileNetV3-Small model, and the symbol relationship diagram is constructed based on the handwriting pressure and inclination characteristics. Transformer is used to dynamically correct the node connection weight of the symbol relationship diagram of the symbol relationship diagram, and the CYK algorithm is used to generate LaTeX expressions, which solves the problem of missing or mis-checking of dense writing areas when writing complex formulas, and realizes efficient symbol detection and formula recognition.

CN120071364AActive Publication Date: 2025-05-30GUANGZHOU DAZZLE VIEW INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510509321.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-30
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

When writing complex formulas, dense writing areas are prone to missed or missed, and the ability to distinguish similar symbols is insufficient. When writing quickly, the system has a recognition delay due to computing power limitations, so it is impossible to reconstruct the coherent path in time.

Method used

The artificial intelligence whiteboard captures the touch point sequence input by the user, constructs the handwriting path trajectory, and delineates the high-density handwriting area. The stroke feature vector is extracted using the MobileNetV3-Small model, and the symbol relationship diagram is constructed based on handwriting pressure and inclination characteristics. Transformer is used to dynamically correct the node connection weight of the symbol relationship diagram with the MLP classification module, and LaTeX expression is generated through the CYK algorithm.

Benefits of technology

It significantly improves the accuracy of symbol detection, can effectively distinguish the upper and lower marks, fractions, and integral structures in complex formulas, solves the problem of traditional methods being prone to missed or missed in dense writing areas, and reduces CPU/GPU load, and is suitable for low-power scenarios of smart whiteboards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071364A_ABST
    Figure CN120071364A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent whiteboards, in particular to a whiteboard handwritten formula real-time identification method and system based on artificial intelligence. The method comprises the following steps: constructing a handwriting path track based on a touch point sequence, and delimiting a high-density handwriting region; performing symbol detection on the high-density handwriting region to position a potential symbol region, and extracting a stroke feature vector of the potential symbol region through a MobileNetV3-Small model; using the stroke feature vector to fuse the handwriting pressure and the inclination angle feature to construct a symbol relation graph; and performing dynamic planning on the initial LaTeX sequence through a CYK algorithm on the basis of node characteristics and a dynamic adjacency matrix in the symbol relation graph, and converting the symbol relation graph into a LaTeX expression. According to the method, dynamic threshold segmentation is adopted, handwriting pressure and dip angle information are combined, the sampling density and the NMS threshold value are dynamically adjusted, a high-density area detection result is reserved preferentially, the symbol detection accuracy is remarkably improved, and the problem that missing detection or false detection is prone to occurring in a dense writing area in a traditional method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent whiteboards, and specifically, to a method and system for real-time recognition of handwritten formulas on a whiteboard based on artificial intelligence. Background Art

[0002] An artificial intelligence whiteboard is an interactive writing platform integrated with artificial intelligence technology, which can perceive, analyze, and process the user's handwritten input in real time. When writing mathematical formulas, physical equations, etc. on the whiteboard in class or remote teaching, the system can recognize and generate standard LaTeX code or visual mathematical expressions in real time. However, when writing complex formulas, there are often problems such as easy missed detection or misdetection in dense writing areas and insufficient ability to distinguish similar symbols.

[0003] At the same time, when writing quickly, due to computing power limitations, the system has a recognition delay and cannot reconstruct a coherent path in time. In view of this, a method and system for real-time recognition of handwritten formulas on a whiteboard based on artificial intelligence are provided. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for real-time recognition of handwritten formulas on a whiteboard based on artificial intelligence to solve the problems of easy missed detection or misdetection in dense writing areas and insufficient ability to distinguish similar symbols when writing complex formulas as mentioned in the above background art.

[0005] To achieve the above purpose, the present invention aims to provide a method for real-time recognition of handwritten formulas on a whiteboard based on artificial intelligence, including the following steps: S1. Capture the sequence of touch points input by the user through the artificial intelligence whiteboard, construct a handwritten path trajectory based on the sequence of touch points , and delimit a high-density handwriting area; S2. Detect symbols in the high-density handwriting area to locate potential symbol areas, and extract stroke feature vectors of the potential symbol areas through the MobileNetV3-Small model ; S3. Use the extracted stroke feature vectors , and fuse the handwriting pressure and inclination features to construct a symbol relationship graph, and use Transformer combined with the MLP classification module to dynamically correct the node connection weights of the symbol relationship graph for constructing the two-dimensional spatial relationship of superscripts, subscripts, fractions, and integral structures; S4. Based on the node features and dynamic adjacency matrix in the symbol relationship graph, generate an initial LaTeX sequence through the Transformer decoder, and perform dynamic programming on the initial LaTeX sequence through the CYK algorithm to convert the symbol relationship graph into a LaTeX expression.

[0006] As a further improvement of this technical solution, in S1, the specific steps involved in delimiting the high-density handwriting area are as follows: S1.1. Combine the handwriting pressure and the inclination angle , and discretize the handwriting path trajectory into a dense point sequence ; S1.2. Merge the original touch point sequence and the dense point sequence to form an enhanced point set , and calculate the grid density ; S1.3. Delimit the high-density handwriting area through dynamic threshold segmentation using the grid density .

[0007] As a further improvement of this technical solution, in S1.3, the specific steps involved in delimiting the high-density handwriting area are as follows: For all non-empty grid cells , calculate their density means ; Based on the density means , calculate the standard deviation of the density distribution ; Use the linear combination of the density means and the standard deviation as the dynamic density threshold ; Traverse all grid cells . If , then mark the grid cell as a high-density area to form a candidate set ; Model the high-density candidate set as a graph ; Based on the breadth-first search algorithm, traverse the graph , and merge adjacent high-density grids into connected areas ; Repeat the process of traversing the graph using the breadth-first search algorithm. Each time, start from an unprocessed starting point to generate new connected areas until all high-density grids have been visited, and finally generate independent high-density handwriting area sets .

[0008] As a further improvement of this technical solution, in S2, the specific steps involved in symbol detection in the high-density handwriting area to locate potential symbol areas are as follows: For the high-density handwriting area set For each connected region in , calculate its center coordinates and the width and height of its circumscribed rectangle , and generate a multi-scale candidate box set ; Use the multi-scale candidate box set as the input of YOLOv8-Nano, and output the initial detection boxes by YOLOv8-Nano; Sort the initial detection boxes in descending order according to the confidence level to obtain an ordered list ; Calculate the average grid density covered by each detection box : Among them, where In the formula, represents the grid cells into which the image is divided; represents the th initial detection box; represents the density value of the grid cell ; Use the ratio of the local density average to the global density average as the NMS threshold adjustment factor, and based on the NMS threshold adjustment factor, the initial NMS threshold , generate a dynamic NMS threshold ; The dynamic threshold is , where is the global density average; In the formula, represents the dynamically adjusted NMS threshold; If the intersection over union of the detection box and the detection box , then retain the detection box with a higher density average and delete the other box. Repeat this step until all candidate boxes are processed; Give priority to retaining the detection results in the high-density area; Construct a feature pyramid network in MobileNetV3-Small, and perform multi-scale symbolic feature extraction on the detection boxes in the high-density area through the feature pyramid network, and output the final symbolic position and category.

[0009] As a further improvement of this technical solution, in S2, the specific steps involved in extracting the stroke feature vector of the potential symbol area by the MobileNetV3-Small model are: According to the detection box coordinates , crop the symbol area from the original image, and perform preprocessing to obtain ; Introduce deformable convolution in the Bottleneck layer of MobileNetV3-Small to dynamically learn the offset of the convolution kernel and the direction angle , to obtain a feature map with an adaptive stroke direction ; Use the feature map as the input of the local stroke attention module of MobileNetV3-Small, and add stroke direction weights in the local stroke attention module , to obtain a feature map of the key stroke area , which is used to strengthen the feature response of the key stroke areas (such as intersections and endpoints); Perform cross-level feature fusion on the feature maps extracted from different Bottleneck layers to obtain a fused multi-scale feature map ; Reduce the dimension of the multi-scale feature map to generate a stroke feature vector .

[0010] As a further improvement of this technical solution, in S3, the specific steps for constructing a symbol relationship graph using the stroke feature vector are as follows: Fuse the handwritten pressure and the tilt angle features in the stroke feature vector to obtain a fused stroke feature vector ; For the initial importance weight of the node pressure ; ; Discretize the tilt angle into an 8-direction encoding ; For any two nodes, calculate the distance between the centers of their circumscribed rectangles , and define a dynamic adjacency threshold : In the formula, represents the dynamic adjacency threshold of the node pair and ; represents the weight coefficient for balancing the influence of the spatial scale; represents the weight coefficient for balancing the influence of the feature difference; represents the cosine similarity; represents the th width of the circumscribed rectangle of the th node; The height of the circumscribed rectangle of the -th node; The width of the circumscribed rectangle of the -th node; The height of the circumscribed rectangle of the -th node; And calculate the inclination angle ; Based on the fused stroke feature vector , direction encoding , and inclination angle difference , for each pair of connected nodes Initialize the edge weight ; Sum up the edge weights of all nodes to construct The adjacency matrix .

[0011] As a further improvement of this technical solution, a Transformer combined with an MLP classification module is used to dynamically correct the node connection weights of the symbol relationship graph. The specific steps involved are as follows: Before constructing the adjacency matrix , perform self-attention encoding on the stroke feature vector using a Transformer encoder to obtain a context-sensitive representation ; For any node pair Construct a joint feature vector ; Introduce an MLP classification module to classify the joint feature vector to predict the correction factor ; Based on the correction factor Update the dynamic adjacency threshold , and finally, based on the updated dynamic adjacency threshold Determine whether to establish a connection in the adjacency matrix .

[0012] As a further improvement of this technical solution, in S4, for each node in the symbol relationship graph, predict the syntactic role using a multi-layer perceptron; Based on the final adjacency matrix and the syntactic role , construct a syntactic structure tree; Use a tree-shaped long short-term memory network to encode the syntactic structure tree from bottom to top; The Transformer decoder generates the LaTeX sequence in an autoregressive manner, predicting the next token at each step based on the syntax tree encoding and the attention mechanism; During the decoding process of the Transformer decoder, relative position encoding between symbols is introduced to adjust the attention scores; The initial LaTeX sequence is dynamically programmed through the CYK algorithm to convert the symbol relationship graph into a LaTeX expression.

[0013] On the other hand, the present invention provides an artificial intelligence-based real-time whiteboard handwritten formula recognition system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the steps of the artificial intelligence-based real-time whiteboard handwritten formula recognition method described in any one of the above.

[0014] Compared with the prior art, the beneficial effects of the present invention are: 1. In the artificial intelligence-based real-time whiteboard handwritten formula recognition method and system, dynamic threshold segmentation is adopted and combined with handwritten pressure and inclination information. By dynamically adjusting the sampling density and NMS threshold, the detection results in high-density regions are preferentially retained, and the response of key regions is enhanced by combining stroke direction features, significantly improving the accuracy of symbol detection, and solving the problems of easy omission or misdetection in densely written regions and insufficient ability to distinguish similar symbols in traditional methods.

[0015] 2. In the artificial intelligence-based real-time whiteboard handwritten formula recognition method and system, stroke features, pressure, and inclination are synchronously fused when constructing the symbol relationship graph. Through dynamic adjacency threshold and context-sensitive weight correction, the relative position and direction between symbols are accurately modeled, improving the accuracy of the syntax structure tree, and solving the problem that traditional methods are difficult to effectively capture the spatial relationships of structures such as superscripts, subscripts, fractions, and integrals.

[0016] 3. In the artificial intelligence-based real-time whiteboard handwritten formula recognition method and system, based on the CYK algorithm combined with the symbol space scoring function, and enhanced Transformer decoding through relative position encoding, geometric constraints such as spatial distance and direction angle are incorporated into the syntax rules, and the optimal parse tree is generated through dynamic programming to ensure the strict correspondence between the LaTeX expression and the handwritten formula; At the same time, through lightweight network design and parallel architecture, the CPU / GPU load is significantly reduced while ensuring accuracy, which is suitable for the low-power scenario of intelligent whiteboards. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is the overall method flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment 1: Please refer to Figure 1 As shown, this embodiment provides a real-time recognition method for handwritten formulas on a whiteboard based on artificial intelligence, including the following steps: S1. Construct a handwritten trajectory by capturing the touch point sequence input by the user through an artificial intelligence whiteboard, and delimit the high-density handwriting area; Specifically, in this embodiment, the artificial intelligence whiteboard records the user's input in the form of time stamps and coordinate points, so the touch point sequence is a data stream containing time and coordinates, and its touch point sequence is specifically: ; Among them, represents the coordinate in the coordinate system of the artificial intelligence whiteboard; represents the time stamp of the th touch point, and satisfies ; represents the total number of touch points within a single stroke; The touch point sequence captures the user's writing action through high-frequency sampling (100Hz) to ensure that the point density is consistent with the real handwriting; Then, based on the touch point sequence construct the handwritten path trajectory The specific steps involved are: Perform Gaussian filtering and noise reduction processing on the original touch point sequence to eliminate the touch screen jitter noise and generate a smooth touch point sequence ; Among them, and ; In the formula, represents the th touch point after filtering; represents the th point in the original touch point sequence ; is the summation variable, representing the traversal range of the touch point index; represents the standard deviation of the Gaussian function; represents the normalization coefficient of the Gaussian function; represents the exponential decay term of the Gaussian kernel; Since large-scale interpolation may increase the CPU load, therefore, the Catmull-Rom interpolation algorithm with dynamic density control is used to prevent the CPU load from being too high: Define the interpolation step adaptive adjustment strategy: where represents the reference interpolation interval; , represents the average writing speed; , represents the speed sensitivity coefficient; represents the th interpolation step of the represents the index variable; The interpolation step adaptive adjustment strategy is used to reduce the interpolation density in the high-speed writing area and increase the interpolation density in the low-speed area to achieve dynamic balance of the calculation load; For every four consecutive touch points Construct independent spline segments: where represents the Catmull-Rom basis function (cubic polynomial); is the normalized time parameter; represents the th independent spline segment constructed by four consecutive touch points ; The interpolation calculations of each segment are independent of each other and meet the parallelization conditions; Furthermore, a parallel computing architecture is adopted to implement real-time interpolation calculation for reducing the calculation time-consuming; Specifically, the touch point sequence is divided into processing units, each unit contains 4 consecutive touch points, and is assigned to an independent computing thread; The Compute Shader is used to perform the interpolation operation in parallel, and real-time interpolation calculation is achieved through the GPU parallel pipeline; At the same time, at the connection of segments Apply continuity constraints to eliminate the connection mutation caused by parallel segmentation by constraining the continuity of the first derivative of adjacent segments: Among them, the continuity constraint is specifically: where: represents the left derivative of the spline segment at the moment; represents the spline segment The right derivative at the moment; Construct a circular buffer to store the interpolation results of adjacent segments, and use the double buffering technique to achieve lock-free data synchronization to ensure the spatio-temporal continuity of the trajectory output.

[0020] In this embodiment, the specific steps for demarcating the high-density handwriting area are as follows: S1.1. Combine the handwriting pressure and the inclination angle , and discretize the handwriting path trajectory into a dense point sequence : Among them, for the trajectory point generated by interpolation, its corresponding handwriting pressure is , and the inclination angle is ; among them, represents the trajectory point coordinates, which are calculated by the interpolation algorithm, represents the th coordinate value of the sampling point on the th spline segment; When discretely sampling each spline segment, the trajectory point generated by interpolation is the value of the function at time ; Sampling at a step size on the normalized time parameter to obtain a dense point sequence , ; In the formula, represents the start time of the spline segment ; represents the end time of the spline segment; represents the sampling time interval; represents the number of sampling points of the spline segment; Dynamically adjust the sampling density through the handwriting pressure and the inclination angle to retain the handwriting details; Combine all the sampling points of the spline segments into a discretized point sequence: In the formula, represents the total number of spline segments; represents the set of all discretized point sequences; represents the last discretized sampling point of the th spline segment; Indicates the number of sampling points on the th spline segment minus 1. If a spline segment is discretized into 5 points, then , the indices are 0, 1, 2, 3, 4; Furthermore, the writing pressure can reflect the stroke weight, and the inclination angle can indicate the writing direction. By fusing these physical features, the system can more accurately distinguish symbols with similar visual forms but different writing dynamics. A stroke with a greater pressure may correspond to the starting stroke of a bold symbol or an integral symbol, and the change in the inclination angle can distinguish the fractional horizontal line (horizontal inclination angle) from italic letters (tilt angle); S1.2. Merge the original touch point sequence with the dense point sequence to form an enhanced point set , and calculate the grid density : The original touch point sequence Specifically: ; The trajectory point sequence generated by the interpolation algorithm (cubic spline interpolation): In the formula, represents the abscissa after interpolation; represents the ordinate after interpolation; represents the pressure value after interpolation; represents the th original touch point; represents the abscissa of the th original touch point; represents the ordinate of the th original touch point; represents the timestamp of the th original touch point; represents the inclination angle of the interpolation point (the tilt angle of the pen, also used for dynamic sampling); represents the index variable; represents the total number of original touch points; Then the enhanced point set is: In the formula, represents the set of all dense point sequences; , represents the dynamically adjusted sampling time step; represents the original touch point sequence at the th point's timestamp; In this embodiment, the point density is increased to the 400 dpi level through secondary sampling, and the spatial resolution of each grid cell is aligned with the target point density of 400 dpi; Furthermore, a quantization grid is defined where inches, representing the physical size of the grid cell; represents the grid index, covering the effective writing area of the whiteboard; represents the width of the grid cell; represents the height of the grid cell; represents the row index of the grid ( axis direction), covering the effective writing area; represents the row index of the grid ( axis direction), covering the effective writing area; Then for each grid cell , calculate the number of times the touch point is covered: For the enhanced point set each point in , then the division basis of the grid cell is: , In the formula, represents the timestamp; grid density is defined as the number of points falling into the grid cell : In the formula, represents the indicator function, which statistically counts the frequency of touch points falling into the grid; when belongs to , the value is 1, otherwise it is 0; represents the coordinate component of the point ; represents the total number of points; By statistically analyzing the distribution of the grid density , detect whether the touch points are evenly covered or there are blind spots.

[0021] S1.3. Delimit the high-density handwriting area by dynamically thresholding the grid density ; Define the dynamic density threshold of the high-density handwriting area: where, for the grid cell , statistically calculate its density mean : Among them, is the total number of rows and columns of the quantization grid; represents the density value of the grid cell ; In the formula, is the adjustment coefficient, and the empirical value range is , which is used to dynamically adapt to the writing density fluctuation and adjust the value through fast writing (low density) and fine writing (high density); is the global density mean value, representing the average value of the densities of all grid cells; is the global density standard deviation, which is used to measure the dispersion degree of the grid density; represents the total number of rows of the quantization grid (the number of grid cells divided in the horizontal direction); represents the total number of columns of the quantization grid (the number of grid cells divided in the vertical direction); represents the point density within the grid cell (i.e., the total number of touch points and interpolation points falling into this grid); represents the th row and the th column of the two-dimensional quantization grid; Traverse all grid cells and determine whether it is a high-density area: If , then mark the grid cell as a high-density cell to form a candidate set ; Construct a graph model: Model the high-density candidate set as a graph , where the vertex set , and the edge set represents the adjacent grid relationship (adopting the 8-neighborhood connection rule). If two grid cells are adjacent in the 8-neighborhood (sharing an edge or a corner), then there is an edge between their corresponding vertices; Traverse the graph through the breadth-first search (BFS) algorithm and merge adjacent high-density grids into connected regions : Among them is the th connected component, containing the index set of all grid cells belonging to the same connected region ; represents the A single independent high-density handwriting region is generated by merging all grid cells within the connected component ; Represents the serial number index of the connected component; Repeat the process of the breadth-first search (BFS) algorithm to traverse the graph , starting from an unprocessed starting point each time to generate a new connected region until all high-density grids are visited, and finally obtain A set of independent high-density handwriting regions , where , provides input for subsequent formula recognition, and each region represents a connected block of high-density handwriting; Represents the set of final output independent high-density regions; Represents the total number of independent high-density regions (i.e., the number of connected components); Represents the th independent high-density handwriting region .

[0022] S2. Use YOLOv8-Nano to perform symbol detection on the high-density handwriting region to locate potential symbol regions, and extract the stroke feature vectors of the potential symbol regions through the MobileNetV3-Small model; Specifically, the specific steps involved in performing symbol detection on the high-density handwriting region to locate potential symbol regions are as follows: For each connected region in the high-density handwriting region set , calculate its center coordinates and the width and height of the circumscribed rectangle ; Generate a set of multi-scale candidate boxes , where the scale factor , and the candidate box coordinates are: In the formula, represents the abscissa of the center of the connected region ; represents the ordinate of the center of the connected region ; represents; represents the width of the circumscribed rectangle of the connected region ; represents the height of the circumscribed rectangle of the connected region ; represents the th candidate box generated based on the connected region scale; Input the candidate bounding boxes into the YOLOv8-Nano network to output the initial detection bounding boxes and their confidence levels; Sort the initial detection bounding boxes in descending order of confidence level, denoted as ; Calculate the average grid density covered by each detection bounding box : , where represents the grid cells into which the image is divided; represents the th initial detection bounding box; represents the density value of the grid cell ; Take the ratio of the local density average to the global density average as the NMS threshold adjustment factor, and based on the NMS threshold adjustment factor, the initial NMS threshold , generate the dynamic NMS threshold ; The dynamic threshold is , where is the global density average; where represents the dynamically adjusted NMS threshold; If the intersection over union of the detection bounding box is greater than , then retain the detection bounding box with a higher density average and delete the other box. Repeat this step until all candidate bounding boxes are processed; Give priority to retaining the detection results in the high-density area, and retain more candidate bounding boxes in the high-density area to improve the recall rate of symbol detection; Specifically, the improved density-weighted (NMS) algorithm solves the problems of over-suppression or missed detection caused by uneven regional density in traditional fixed-threshold NMS in complex handwritten formulas by dynamically adjusting the threshold ; In the densely written area , increase the threshold , allow boxes with higher overlap to coexist, and avoid misdeleting adjacent symbols (such as superscripts and subscripts); In the sparse area , lower the threshold to reduce false detections; Construct a feature pyramid network (FPN) in MobileNetV3-Small to fuse the low-level feature map with the high-level feature maps and : Detect small-scale symbols at the layer, and at the The layer detects large-scale symbols and extracts multi-scale symbol features through the Feature Pyramid Network, that is, the Feature Pyramid Network performs multi-scale symbol feature extraction on the area containing symbols (i.e., the detection box of the high-density area), and outputs the final symbol position and category.

[0023] Specifically, the output feature maps of the backbone network of MobileNetV3-Small (shallow layer with high resolution), (middle layer), (deep layer with low resolution) are used as inputs; Through the Feature Pyramid Network, 、 、 are fused across layers, and finally a multi-scale feature pyramid is output; 、 and are used as outputs, corresponding to different detection scales respectively. On each level (P3, P4, P5), a YOLOv8-Nano detection head is deployed to predict symbols corresponding to the respective scales: For the deep feature perform 3×3 convolution for dimensionality reduction to generate the initial : Upsample P5 to the size of C4 and add it element-wise to the 1×1 convolution result of C4 to generate : Upsample P4 to the size of C3 and add it element-wise to the 1×1 convolution result of C3 to generate : Then the final feature pyramid: : The highest resolution (such as 1 / 8), used to detect small symbols (such as dots, apostrophes, commas); : Medium resolution (such as 1 / 16), used to detect regular symbols (such as letters, numbers); : The lowest resolution (such as 1 / 32), used to detect large symbols (such as integral signs, fraction bars).

[0024] The YOLOv8-Nano detection head predicts the symbol position and category on P3, P4, and P5. Each detection head generates detection results through convolution operations. The YOLOv8-Nano detection head aggregates the prediction results of all levels. Finally, the detection results are optimized through an improved density-weighted (NMS) algorithm, and the final symbol list is output.

[0025] Specifically, compared with the traditional NMS that uses a fixed threshold, the dynamic threshold adjustment based on the local / global density ratio gives priority to retaining the results of high-confidence and high-density regions, overcoming the problem of easy missed detection in dense regions by the traditional method; and the handwriting grid density is introduced , and the complexity of the symbol distribution is quantified by calculating the density mean of the area covered by the detection box; for example, there may be dense superscripts and subscripts around the integral symbol, and at this time, the dynamic threshold can avoid misdeleting the key detection box.

[0026] In this embodiment, the specific steps involved in extracting the stroke feature vector of the potential symbol region by the MobileNetV3-Small model are as follows: According to the detection box coordinates , and crop out the symbol region from the original image , and perform preprocessing to obtain : Scale to a fixed size (such as 64×64 pixels) to generate a standardized image , Convert to a single-channel grayscale image and perform pixel value normalization: In the formula, represents the mean value; represents the standard deviation; represents the standardized representation of the symbol region and serves as the only input to the feature extraction network; Aiming at the stroke characteristics (direction, curvature, continuity) of handwritten symbols, a direction-sensitive convolution and a local attention mechanism are introduced in MobileNetV3-Small to enhance the ability to capture stroke details: Through MobileNetV3-Small, is subjected to feature extraction, and deformable convolution is introduced in the Bottleneck layer of MobileNetV3-Small to dynamically learn the convolution kernel offset and the direction angle , to obtain a feature map with an adaptive stroke direction : The network hierarchy of MobileNetV3-Small includes an initial convolution layer, a Bottleneck layer, a local stroke attention module, multi-level feature fusion, and global feature descriptor generation; Among them, the initial convolution layer is used to perform preliminary feature extraction on to generate a low-level feature map In the formula, represents the low-level feature map output by the initial convolutional layer; Bottleneck layer: Deformable convolution is introduced in the Bottleneck layer: The input is the low-level feature map generated by the previous layer ; In the formula, represents the initial offset position of the fixed convolution kernel; represents the spatial position coordinates on the feature map; represents the learnable offset; represents the direction angle parameter, which is optimized by gradient descent to make the convolution kernel adapt to the stroke direction; represents the feature map with an adaptive stroke direction output by the deformable convolution, that is, the output of the Bottleneck layer; represents the total number of convolution kernels; represents the index variable; represents the weight of the convolution kernel, a parameter automatically learned through model training, used to perform weighted summation on different positions on the input feature map; represents the input of the Bottleneck layer, that is, the low-level feature map ; Local stroke attention module: Increase the stroke direction weight Then the output of the local stroke attention module is: In the formula, represents global average pooling; represents convolution operation; represents the Sigmoid activation function, used to map the weight value to the interval [0,1]; represents the stroke direction attention weight matrix; represents element-wise multiplication; represents the feature map after attention enhancement; Multi-level feature fusion: Extract feature maps from the Bottleneck3, 6, and 12 layers of MobileNetV3-Small , and perform fusion: In the formula, represents the output of the 3rd Bottleneck layer; Represents the output of the 6th Bottleneck layer; Represents the output of the 12th Bottleneck layer; Represents an upsampling operation (such as bilinear interpolation); Represents the fused multi-scale feature map; Represents element-wise addition for feature fusion; Global feature descriptor generation: In the formula, Represents global average pooling; Represents a 512-dimensional fully connected layer; Represents the unnormalized stroke feature vector; Performs L2 normalization on the feature vector to generate the final stroke feature vector : In the formula, Represents the L2 norm for calculating the magnitude of a vector; Represents the normalized stroke feature vector; In this embodiment, a bidirectional feature pyramid is constructed in MobileNetV3-Small to fuse semantic features at different levels; for example, shallow features Retain stroke details (such as dots, flicks), deep features Capture structural information (such as fractions, square roots), solving the problem of large scale differences in handwritten formulas; Embed a local stroke attention module in the Feature Pyramid Network (FPN) and add stroke direction weights in the local stroke attention module , giving priority to key areas such as stroke intersections and symbol connections; for example, it can effectively distinguish the from morphological differences; S3. Use the extracted stroke feature vectors Construct a symbol relationship graph and dynamically correct the node connection weights of the symbol relationship graph based on graph neural networks or Transformers for constructing the two-dimensional spatial relationships of superscripts, subscripts, fractions, and integral structures; In this embodiment, using the stroke feature vectors The specific steps involved in constructing the symbol relationship graph are as follows: In the stroke feature vector Fuse the handwritten pressure with the inclination angle features to obtain the fused stroke feature vector ; Handwritten pressure and inclination angle As a supplementary dimension of the stroke feature vector , it is used to enhance the modeling ability of the physical characteristics of the writing tool (such as the tilt angle of the chalk and the pressure change of the pen), so as to distinguish symbols with similar shapes but different writing dynamics (such as "6" and "9"). Among them, the handwriting pressure and the tilt angle need to be normalized to ensure the scale consistency of the feature vector and avoid weight deviation caused by physical dimension differences, , where represents the concatenation operation along the feature dimension; Based on the fused stroke feature vector construct an adjacency matrix : For each node , define the node feature vector : And for the feature vector of node , it is concatenated by the stroke feature vector fusing the handwriting pressure and the tilt angle ; Initialize the importance weight for each node , and normalize the handwriting pressure using the Sigmoid function: And discretize the tilt angle into an 8-direction encoding: where represents the handwriting pressure of the th node; represents the Sigmoid function, which is used to map the pressure value to a normalized weight range; represents the initial importance weight of the th node, reflecting the strength information of this node during writing; represents the complete feature vector of the th node that fuses semantic and physical attributes, and its total dimension is ; represents the nib tilt angle information of the th node; represents discretizing the continuous tilt angle into an 8-direction encoding (such as up, down, left, right and their diagonal directions), so it is represented by an 8-dimensional vector; For any two nodes, the distance between the centers of their bounding rectangles is defined as: where, represents the Euclidean distance between the centers of the bounding rectangles corresponding to node and node ; represents the coordinates of the center of the rectangle of the -th node; According to the differences in node size and features, a dynamic adjacency threshold is defined: where, represents the dynamic adjacency threshold between node pairs and ; represents the weight coefficient used to balance the influence of spatial scale; represents the weight coefficient used to balance the influence of feature differences; represents the larger one of the sizes of node and node , used to characterize the spatial scale; represents the cosine similarity, used to measure the scale between the fusion feature vectors of two nodes ; represents the width of the bounding rectangle of the -th node; represents the height of the bounding rectangle of the -th node; Similarly, represents the width of the bounding rectangle of the -th node; represents the height of the bounding rectangle of the -th node; If , a basic connection is established between node and node ; And the inclination difference of the inclination angle is calculated: If , it is considered that the writing directions are the same, and the weight of the edge is enhanced; represents the nib inclination angle of the -th node; represents the nib inclination angle of the -th node; represents the absolute difference in the nib inclination angles between node and node . If this difference is less than the set threshold , it is considered that the writing directions of the two are the same, and thus the edge weights are enhanced; Based on the fused stroke feature vectors , direction encoding and inclination difference , for each pair of connected nodes the initial edge weights are: In the formula, represents the initial weight between node and node ; represents concatenating the fused feature vectors of the two nodes to obtain a joint feature representation of a higher dimension; represents the joint direction encoding vector generated from the direction encoding of node and node (or the direction information of the relationship between the two), reflecting their relative directions (such as "upper right", "lower left", etc.); represents element-wise addition, used to fuse feature information from different sources (joint features, direction encoding, and inclination difference); represents a learnable weight matrix, used to linearly transform the fused features to generate appropriate edge weights; represents the Sigmoid activation function; Summarize the edge weights of all nodes to construct the adjacency matrix (where is the number of nodes), which is defined as: If the spatial connection condition or the feature similarity condition is satisfied, then , otherwise ; Among them, the spatial connection condition is , and the feature similarity condition is ; Above, represents whether a connection is established between node and node and the weight of the connection; represents the feature similarity threshold. When the similarity is greater than this threshold, it is considered that the two are similar enough in features to establish a connection; represents the cosine similarity between the fused feature vectors of node and node , used to measure whether they are similar semantically.

[0027] The Transformer dynamically corrects the node connection weights of the symbol relationship graph. The specific steps involved are: Before constructing the adjacency matrix the node fusion features are self-attention encoded based on the Transformer encoder to obtain context-sensitive representations ; In the formula, represents the context-sensitive representation obtained by the -th node after the self-attention mechanism, which is the result of weighted summation of all node information; represents the query vector of the -th node; represents the dimension of the key vector (or query vector), which is used as a scaling factor to prevent the dot product result from being too large and thus affecting the gradient; represents the number of elements in the input sequence, that is, the total number of nodes; represents the value vector of the -th node; represents the key vector of the -th node, represents the transpose of the key vector of the -th node; represents the index variable; For any node pair a joint feature vector is constructed: For each pair of nodes and , a joint feature vector In the formula, represents the context-sensitive feature vector of node after being encoded by the Transformer self-attention mechanism; represents the context-sensitive feature vector of node after being encoded by the Transformer self-attention mechanism; represents the dimension of the context-sensitive feature vector of each node ; represents the element-wise product of the feature vectors of node and node ; represents the element-wise difference of the feature vectors of node and node ; An MLP classification module is introduced to classify the joint feature vector and predict the correction factor ; Take the multi - layer perceptron (MLP) as the classification module to process the joint feature vector and output the correction factor corresponding to the node pair : : In the formula, represents the weight matrix of the first layer of the MLP; is the bias vector of the first layer of the MLP; is the bias scalar of the second layer of the MLP; is the weight matrix of the second layer of the MLP, which is used to map the output of the hidden layer to the correction factor ; represents the non - linear activation function; represents the Sigmoid function; Based on the correction factor update the dynamic adjacency threshold , and finally, according to the updated dynamic adjacency threshold judge whether to establish a connection in the adjacency matrix ; Combined with the strategy of correcting the context information, the updated threshold is: Wherein, represents the adjustable weight factor, which is used to balance the contributions of the original threshold and the correction factor; represents the updated dynamic adjacency threshold; Finally, according to the center distance of the circumscribed rectangles between node pairs , construct the final adjacency matrix : Then, after optimization, if the spatial connection condition or the feature similarity condition is satisfied, then , otherwise ; Among them, the spatial connection condition is , and the feature similarity condition is ; represents the feature similarity threshold, which is used to judge the semantic association strength between nodes.

[0028] In this embodiment, the final adjacency matrix encodes all the structural information of the symbolic relationship graph, and its dynamic nature enables the symbolic relationship graph to adapt to the writing changes of complex formulas. The adjacency matrix is the mathematical expression of the symbolic relationship graph; The symbolic relationship graph provides a directly parsable topological structure for subsequent syntax tree construction and LaTeX generation.

[0029] S4. Based on the node features and dynamic adjacency matrix in the symbol relationship graph, generate an initial LaTeX sequence through a Transformer decoder, and perform dynamic programming on the initial LaTeX sequence through the CYK algorithm to convert the symbol relationship graph into a LaTeX expression; In this embodiment, each node in the symbol relationship graph is used to predict the syntactic role through a multi-layer perceptron (MLP) : where represents the predicted syntactic role of node ; is the multi-layer perceptron; Based on the final adjacency matrix and the syntactic role , construct a syntactic structure tree: Based on the updated dynamic adjacency threshold , determine the parent-child relationship through breadth-first search; Use a tree-shaped long short-term memory network to encode the syntactic structure tree from bottom to top: For each node of the syntax tree, aggregate the features of the child nodes to generate a hidden state : where represents the LSTM cell of the tree-shaped long short-term memory network, aggregating the information of the subtree from bottom to top; represents the hidden state of node , encoding its syntactic role and the features of the child nodes; represents the hidden state of the child node ; The Transformer decoder generates a Token sequence based on the syntax tree encoding ; Meanwhile, introduce the relative coordinate difference between symbols, , map it to an embedding vector and add it to the attention score: where , represents the horizontal coordinate difference between node and ; , represents the vertical coordinate difference between node and ; where Indicates the "attention requirement" of the target position to be generated currently (such as the LaTeX token being decoded) for the input symbol; Indicates the "feature index" of the input symbol (a node in the symbol relationship graph); Indicates the actual feature information of the input symbol, which is used to generate the target output after weighted aggregation; Indicates the core input matrix for generating attention scores; Indicates calculating the similarity between the query and the key (the original attention weight); Indicates a learnable function that maps the coordinate difference to an attention bias term; Indicates the dimension of the key vector; The attention score is obtained through and The weight value calculated by the similarity, indicating the importance of the input symbol to the current target position; Furthermore, during the decoding process of the Transformer decoder, relative position encoding between symbols is introduced to adjust the attention score: Encodes spatial constraints (distance, direction, adjacency weight) as grammar rule scores; For each candidate rule , define the spatial score function: In the formula, Indicates the Euclidean distance between symbol and ; Indicates the direction angle between symbol and ; Indicates the ideal direction expected by rule ; Indicates the edge weight between and in the adjacency matrix; Indicates the distance decay coefficient, which controls the spatial sensitivity; Indicates the parent node symbol; Both indicate candidate child symbols; In the CYK algorithm, for the dynamic programming update rule, the weight of each cell is: In the formula, Indicates the grammar rule probability; Indicates the spatial score, which is used to amplify the probability of the rule that conforms to the spatial layout; Indicates covering the non-terminal symbol from to of the input position, represents the starting position of the input sequence (the starting index of the symbol in the symbol relationship graph), represents the ending position of the input sequence (the ending index of the symbol in the symbol relationship graph); represents the splitting point; represents from position to the maximum score that the subsequence can be generated by the non-terminal ; and both represent non-terminals in the context-free grammar (CFG) (such as expressions, terms, factors, etc.); By performing dynamic programming on the initial LaTeX sequence through the CYK algorithm, the symbol relationship graph is converted into a LaTeX expression. The specific expressions involved are: In the formula, represents the parse tree that conforms to the grammar rules and space constraints; represents the logarithmic probability of the grammar rule; represents the logarithmic term of the space score, which is used to amplify the rules of reasonable layout; Embodiment 2: This embodiment provides an artificial intelligence-based real-time recognition system for whiteboard handwritten formulas, including a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the computer program to implement the steps of the artificial intelligence-based real-time recognition method for whiteboard handwritten formulas described in any one of the above.

[0030] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A real-time recognition method for whiteboard handwritten formulas based on artificial intelligence, characterized in that: The following steps are involved: S1. Capture the touch point sequence input by the user through the artificial intelligence whiteboard, and build the handwriting path trajectory based on the touch point sequence , and delineate high-density handwriting areas; S2: Perform symbol detection on high-density handwriting areas to locate potential symbol areas, and extract stroke feature vectors of potential symbol areas through the MobileNetV3-Small model ; S3. Use the extracted stroke feature vector , and incorporate handwriting pressure With inclination The symbolic relationship graph is constructed based on the features, and the node connection weights of the symbolic relationship graph are dynamically corrected using Transformer combined with the MLP classification module; S4. Based on the node features and dynamic adjacency matrix in the symbolic relationship graph, the initial LaTeX sequence is generated through the Transformer decoder, and the initial LaTeX sequence is dynamically programmed through the CYK algorithm to convert the symbolic relationship graph into a LaTeX expression.

2. The method for real-time recognition of whiteboard handwritten formulas based on artificial intelligence according to claim 1, characterized in that: In S1, the specific steps involved in delineating the high-density handwriting area are: S1.

1. Combined with handwriting pressure With inclination , the handwritten path trajectory Discretize into a dense point sequence ; S1.

2. Merge the original touch point sequence With dense point sequence Forming an enhanced point set , and calculate the mesh density ; S1.

3. Set the grid density High-density handwriting areas are delineated through dynamic threshold segmentation.

3. The method for real-time recognition of whiteboard handwritten formulas based on artificial intelligence according to claim 2, characterized in that: In S1.3, the specific steps involved in delineating the high-density handwriting area are: For all non-empty grid cells , and calculate the mean density ; Based on density mean Calculate the standard deviation of a density distribution ; The density mean With standard deviation The linear combination of ; Iterate through all grid cells ,like , then mark the grid cells is a high-density area, forming a candidate set ; The high-density candidate set Modeling as a graph ; Traversing the graph based on the breadth-first search algorithm , merge adjacent high-density grids into connected areas ; Repeated breadth-first search algorithm traverses the graph The process starts from the unprocessed starting point each time and generates a new connected area until all high-density grids are visited, and finally generates A collection of independent high-density handwriting areas .

4. The method for real-time recognition of whiteboard handwritten formulas based on artificial intelligence according to claim 3 is characterized in that: In S2, the specific steps involved in performing symbol detection on the high-density handwriting area to locate the potential symbol area are: For high-density handwriting area collection Each connected region in , calculate its center coordinates and the width and height of the bounding rectangle , generate a set of multi-scale candidate boxes ; The multi-scale candidate box set As YOLOv8-Nano input, YOLOv8-Nano outputs the initial detection box; Sort the initial detection boxes in descending order by confidence to get an ordered list ; Calculate each detection box Mean grid density of coverage ; The local density mean With the global density mean The ratio of is used as the NMS threshold adjustment factor, and the initial NMS threshold is based on the NMS threshold adjustment factor. , generate dynamic NMS threshold ; If the detection frame With detection box The intersection ratio , then keep the detection frame with higher density mean, delete the other frame, and repeat this step until all candidate frames are processed; A feature pyramid network is constructed in MobileNetV3-Small, and multi-scale symbol features are extracted from the detection boxes in high-density areas through the feature pyramid network to output the final symbol position and category.

5. The method for real-time recognition of whiteboard handwritten formulas based on artificial intelligence according to claim 4 is characterized in that: In S2, the specific steps involved in extracting the stroke feature vector of the potential symbol area through the MobileNetV3-Small model are: According to the detection box coordinates , and crop the symbol area from the original image , and Preprocessing to obtain ; Introducing deformable convolution in the Bottleneck layer of MobileNetV3-Small to dynamically learn the convolution kernel offset With direction angle , get the feature map of adaptive stroke direction ; The feature map As the input of the local stroke attention module of MobileNetV3-Small, and increase the stroke direction weight in the local stroke attention module , get the feature map of the key stroke area ; The feature maps extracted from different Bottleneck layers are cross-level fused to obtain the fused multi-scale feature maps. ; Multi-scale feature map Dimensionality reduction to generate stroke feature vector .

6. The method for real-time recognition of whiteboard handwritten formulas based on artificial intelligence according to claim 1, characterized in that: In S3, the stroke feature vector is used The specific steps involved in building a symbolic relationship diagram are: In the stroke feature vector Medium Fusion Handwriting Pressure With inclination Features, get the fused stroke feature vector ; For node pressure Initial importance weights ; The inclination angle Discretized into 8-directional encoding ; For any two nodes, calculate the distance between the centers of their circumscribed rectangles , and define a dynamic adjacency threshold : In the formula, Represents a node pair and Dynamic adjacency threshold of The weight coefficient representing the effect of balancing spatial scale; The weight coefficient representing the impact of the difference in balancing characteristics; represents cosine similarity; Indicates The width of the bounding rectangle of each node; Indicates The height of the bounding rectangle of each node; Indicates The width of the bounding rectangle of each node; Indicates The height of the bounding rectangle of each node; And calculate the inclination The difference in inclination ; Based on the fused stroke feature vector , direction coding and inclination difference , for each pair of connected nodes Initial edge weights ; All nodes The edge weights of The adjacency matrix of .

7. The method for real-time recognition of whiteboard handwritten formulas based on artificial intelligence according to claim 6 is characterized in that: The Transformer combined with the MLP classification module is used to dynamically correct the node connection weights of the symbolic relationship graph. The specific steps involved are: In constructing the adjacency matrix Based on the Transformer encoder, the stroke feature vector Perform self-attention encoding to obtain context-sensitive representation ; For any node pair Constructing joint eigenvectors ; Introduce the MLP classification module to the joint feature vector Perform classification processing and predict correction factors ; Based on the correction factor Update dynamic adjacency threshold , and finally according to the updated dynamic adjacency threshold Determine whether it is in the adjacency matrix Establish a connection.

8. The method for real-time recognition of whiteboard handwritten formulas based on artificial intelligence according to claim 7 is characterized in that: In S4, for each node in the symbolic relationship graph , predicting grammatical roles based on multi-layer perceptron ; Based on the final adjacency matrix and grammatical roles , construct a syntax structure tree; A tree-shaped long short-term memory network is used to encode the syntax structure tree from bottom to top; The Transformer decoder generates LaTeX sequences in an autoregressive manner, predicting the next token at each step based on the syntax tree encoding and attention mechanism; During the decoding process of the Transformer decoder, the relative position encoding between symbols is introduced to adjust the attention score; The initial LaTeX sequence is dynamically programmed using the CYK algorithm to convert the symbolic relationship graph into a LaTeX expression.

9. A real-time recognition system for handwritten formulas on a whiteboard based on artificial intelligence, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor executes a computer program to implement the steps of the real-time recognition method of whiteboard handwritten formulas based on artificial intelligence as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Control method for touch writing acceleration under Android system

    CN110737364A

  • Online handwritten mathematical formula recognition method based on bidirectional Tree-GRU

    CN110929634A

  • Formula identification method and device and device for formula identification

    CN113408417A

  • Handwritten mathematical formula identification method

    CN117542064A

  • Mathematical formula recognition method and apparatus, electronic device, and readable storage medium

    WO2024244760A1

Cited By

  • Handwritten formula image recognition method based on rule injection and related device

    CN121768005A

  • A handwritten formula image recognition method based on rule injection and a related device

    CN121768005B