Dynamic online handwritten flow chart identification method and device
By extracting geometric representations from online stroke data and employing an attention masking mechanism for streaming inference, and utilizing a Transformer encoder and Maskformer decoder for real-time symbol instance recognition, the problem of insufficient real-time performance and accuracy in online handwritten flowchart recognition is solved, achieving efficient human-computer interaction and recognition results.
Patent Information
- Application Number
- CN202510803082.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-11-18
AI Technical Summary
In existing technologies, the recognition methods for online handwritten flowcharts cannot meet the dual requirements of real-time performance and accuracy. They cannot provide accurate recognition results during the user's real-time writing process, resulting in poor human-computer interaction.
The algorithm extracts geometric representations of strokes from online stroke data, performs streaming inference through an attention masking mechanism, and uses a Transformer encoder and a Maskformer decoder for real-time identification of symbol instances, including stroke feature sequence segmentation, historical information masking, and future information masking. It also combines a multilayer perceptron and a target classifier to output symbol category and location information in real time.
It realizes real-time dynamic streaming symbol segmentation and recognition of online handwritten flowcharts, improves the real-time performance and accuracy of recognition, and supports the real-time needs of human-computer interaction.
Smart Images

Figure CN120976606A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a dynamic online handwritten flowchart recognition method and device. BACKGROUND
[0002] Online handwritten documents are generated by users with the help of electronic writing tools, and the path information of the movement of the electronic handwriting pen is completely saved and stored in the form of a structured digital electronic file. Related hardware devices (such as touch screens with pressure sensing or electromagnetic induction plates) can accurately collect multi-dimensional parameters during writing, including the coordinates, time and pen tip pressure of each stroke point. As an important category of online handwritten documents, online handwritten flowcharts are widely used in intelligent education, collaborative office, remote meetings and other scenarios, and have become one of the research focuses of the academic and industrial communities.
[0003] The difficulty of recognizing online handwritten flowcharts lies in their complex two-dimensional spatial structure, which contains various geometric symbols (such as rectangles, circles and diamonds) and connectors (such as arrows). These symbols are related to each other through non-linear space-time relationships, and the differences in user writing styles and the diversity of stroke order further increase the difficulty of recognition.
[0004] In related technologies, complete flowchart data is usually input into recurrent neural networks and graph neural networks to implement online handwritten flowchart recognition. Since users need real-time feedback of the recognition results to support human-computer interaction when writing flowcharts, the above recognition method requires the recognition of complete flowcharts, which makes it difficult for the static recognition method to meet the dual demands of real-time and accuracy. SUMMARY
[0005] The present application provides a dynamic online handwritten flowchart recognition method and device to solve the defect that the complete flowchart needs to be provided to the recognition model in the prior art, which cannot obtain the recognition results in real time to support human-computer interaction, resulting in the difficulty in meeting the dual demands of real-time and accuracy. The method of the present application realizes the real-time and accuracy of online handwritten flowchart recognition.
[0006] The present application provides a dynamic online handwritten flowchart recognition method, comprising: extracting the geometric representation of strokes from online stroke data to obtain a stroke feature sequence; using an attention mask mechanism to perform streaming inference on the stroke feature sequence through target encoding operations to obtain encoded features; the target encoding operations include dividing the stroke feature sequence into non-overlapping blocks, setting each stroke feature in the block to be visible to each other, and setting the current block to only access the previous block in the historical window, and shielding future block information; Based on the encoded features, symbol instances in the online handwriting flowchart are identified in real time, and the category and location information of each symbol instance are output.
[0007] According to the present invention, a method for recognizing dynamic online handwritten flowcharts includes extracting the geometric representation of strokes from online stroke data to obtain a stroke feature sequence, which includes: The online stroke data is preprocessed by resampling, shifting, and height normalization to obtain processed online stroke data; wherein, the online stroke data includes stroke height information and the resampled stroke sequence; A fixed-length feature extraction network extracts fixed-length feature vectors from the processed online stroke data. This network uses sample online stroke data as training samples for a Transformer encoder, the sample fixed-length feature vectors as input to a Maskformer decoder network, and the sample stroke coordinate sequence as the output of the Maskformer decoder. The encoder-decoder architecture is trained by minimizing the difference between the sample stroke coordinate sequence and the resampled stroke sequence. The sample stroke coordinate sequence is reconstructed from the sample fixed-length feature vectors. The geometric representation is obtained based on the fixed-length feature vector and the stroke height information, and the stroke feature sequence is obtained according to the geometric representation corresponding to different numbers of strokes in the online stroke data.
[0008] According to the present invention, a dynamic online handwritten flowchart recognition method is provided, wherein the target encoding operation is implemented by a Transformer encoder, and the Transformer encoder includes a multi-head attention layer; The attention masking mechanism is used to perform streaming inference on the stroke feature sequence through target encoding operations to obtain encoded features including: The stroke feature sequence is subjected to relative temporal encoding to obtain a temporal encoding result; the stroke feature sequence is subjected to relative spatial encoding to obtain a spatial encoding result. Multi-head attention calculation is performed based on the temporal and spatial encoding results. In the multi-head attention layer, the attention mask is used to control the visibility range of each stroke feature, and a filling mask is used to block invalid stroke features to obtain the multi-head attention calculation result. The multi-head attention calculation result is processed by layer normalization and then input into the feedforward neural network. The output of the feedforward neural network is then processed by layer normalization again to obtain the encoded features.
[0009] According to the present invention, a method for recognizing dynamic online handwritten flowcharts includes the following steps: real-time recognition of symbol instances in the online handwritten flowchart based on the encoded features, and outputting category and location information for each symbol instance. Based on the Maskformer decoder, a query embedding vector is generated according to the encoded features and multiple positional encoded vectors; The query embedding vector is processed using a multilayer perceptron to obtain a mask embedding vector, and the encoded features are processed using a multilayer perceptron to obtain a stroke embedding vector. The stroke embedding vector and the mask embedding vector are merged to obtain the predicted binary mask; The target classifier predicts the symbol category based on the mask embedding vector to obtain the category information; the predicted binary mask is then used for symbol segmentation mask prediction to obtain the position information; wherein, the target classifier is determined based on a linear classifier and a Softmax function.
[0010] According to a dynamic online handwritten flowchart recognition method provided by the present invention, the plurality of position encoding vectors include a query vector and an additional query vector; The target classifier is obtained through the following steps: The target classifier is trained using sample radix masks and sample mask embedding vectors as training samples, joint loss as the loss function, and the DETR architecture is trained using a bipartite graph matching mechanism based on one-to-one matching and one-to-many matching methods. The one-to-one matching method is achieved by using the Hungarian algorithm to optimally match K query vectors with m real symbols; the one-to-many matching method is achieved by copying the m real symbols n times and using the Hungarian algorithm to optimally match the K query vectors with m×n symbols; the joint loss is determined based on the one-to-one matching loss function, the one-to-many matching loss function, and the auxiliary loss.
[0011] According to the dynamic online handwritten flowchart recognition method provided by the present invention, the one-to-one matching loss and the one-to-many matching loss are determined by the category prediction cost, the binary mask loss cost and the focus loss cost.
[0012] The present invention also provides a dynamic online handwritten flowchart recognition device, comprising: The feature extraction module is used to extract the geometric representation of strokes from online stroke data to obtain stroke feature sequences; The streaming inference module is used to perform streaming inference on the stroke feature sequence through target encoding operation using an attention masking mechanism to obtain encoded features; the target encoding operation includes dividing the stroke feature sequence into non-overlapping blocks, setting each stroke feature within a block to be mutually visible, setting the current block to only access the preceding blocks in the history window, and masking information of future blocks; The recognition module is used to identify symbol instances in the online handwriting flowchart in real time based on the encoded features, and output the category information and position information of each symbol instance.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the dynamic online handwritten flowchart recognition method as described above.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the dynamic online handwritten flowchart recognition method as described above.
[0015] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the dynamic online handwritten flowchart recognition method as described above.
[0016] The present invention provides a dynamic online handwritten flowchart recognition method and apparatus. This method extracts the geometric representation of strokes from online stroke data to obtain a stroke feature sequence. An attention masking mechanism is then used to perform streaming inference on the stroke feature sequence through target encoding operations to obtain encoded features. Specifically, the stroke feature sequence is segmented into non-overlapping blocks, each stroke feature within a block is made mutually visible, and the current block only accesses preceding blocks within the history window, while future block information is masked. Finally, symbol instances in the online handwritten flowchart are identified in real time based on the encoded features, and the category and position information of each symbol instance are output. This achieves real-time dynamic streaming symbol segmentation and recognition of online handwritten flowcharts, improving the real-time performance and accuracy of online handwritten flowchart recognition. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is one of the flowchart diagrams of the dynamic online handwritten flowchart recognition method provided by the present invention.
[0019] Figure 2 This is one of the flowcharts illustrating the stroke feature sequence extraction method provided by the present invention.
[0020] Figure 3 This is the second flowchart of the stroke feature sequence extraction method provided by the present invention.
[0021] Figure 4 This is a schematic diagram of the working mechanism of the attention mask provided by the present invention.
[0022] Figure 5 This is a schematic diagram of the streaming inference process provided by the present invention.
[0023] Figure 6 This is a flowchart illustrating the symbol instance recognition method provided by the present invention.
[0024] Figure 7 This is the second flowchart illustration of the dynamic online handwritten flowchart recognition method provided by the present invention.
[0025] Figure 8 This is a schematic diagram of the structure of the dynamic online handwritten flowchart recognition device provided by the present invention.
[0026] Figure 9 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0028] The following is combined with Figures 1-9 This invention describes a dynamic online handwritten flowchart recognition method and apparatus.
[0029] Figure 1 This is one of the flowchart illustrations of the dynamic online handwritten flowchart recognition method provided by the present invention, such as... Figure 1 As shown, the method includes the following: Step 110: Extract the geometric representation of the strokes from the online stroke data to obtain the stroke feature sequence.
[0030] In this step, a stroke feature extraction method based on unsupervised pre-training automatically extracts effective geometric feature representations from online stroke data; whereby unsupervised pre-training can capture the inherent geometric information of strokes by utilizing a large amount of unlabeled data.
[0031] For example, pre-trained network models such as Transformer network architecture, LSTM (Long Short-Term Memory), and RNN (Recurrent Neural Network) can be used for stroke feature extraction.
[0032] In this embodiment, the inherent geometric information of the stroke includes the direction of the stroke point (such as east, west, south, north, southeast, etc.), the curvature of the stroke, and the writing speed.
[0033] It should be noted that, compared with the highly subjective manual features, the stroke feature extraction method provided in this embodiment can not only maintain the advantages of unsupervised training, but also provide better stroke geometric feature representation for stroke classification and symbol segmentation.
[0034] Step 120: Use attention masking mechanism to perform streaming inference on the stroke feature sequence through target encoding operation to obtain encoded features; the target encoding operation includes dividing the stroke feature sequence into non-overlapping blocks, setting each stroke feature in the block to be visible to each other, setting the current block to only access the previous block in the history window, and masking the information of future blocks.
[0035] In this step, an innovative stroke masking mechanism is used to simulate the streaming stroke input process. An attention masking strategy is used to shield the following information of incomplete strokes, enabling the model to perform inference based solely on the current and historical stroke points.
[0036] Specifically, a Transformer encoder can be used to process temporal stroke features through a masked self-attention layer, which preserves the long-range dependencies of completed strokes while avoiding the leakage of irrelevant information.
[0037] In some embodiments, a graph neural network can also be used to perform streaming inference on the stroke feature sequence according to the above target encoding operation to obtain the corresponding encoding module.
[0038] In this embodiment, the implementation method of streaming inference includes: in the encoder-decoder Transformer network framework of the network, local modeling is indirectly achieved in a block-based manner, and attention masks are used to shield irrelevant contexts, thereby constructing a streaming dynamic online handwritten flowchart recognition algorithm.
[0039] In this embodiment, the attention mask is used to truncate the history and limit the length of usable future information; for example, when modeling sequential data using an attention mechanism with a Transformer, the attention mask can be applied to the Transformer's attention weight matrix { }, to specify the range of the input sequence involved in the computation; when When the attention mask for a location is set to 0, it represents time. Input It will not be used to calculate time. Output .
[0040] In this embodiment, a Transformer network can be used to perform temporal and spatial encoding on the stroke feature sequence to obtain the corresponding spatiotemporal information and achieve accurate modeling of the contextual relationship of the strokes; then, the attention masking mechanism mentioned above is used to perform streaming inference on the spatiotemporal information through block encoding operations to obtain the corresponding encoded features.
[0041] Step 130: Identify symbol instances in the online handwriting flowchart in real time based on encoding features, and output the category and location information of each symbol instance.
[0042] In this step, in the online handwritten flowchart, a symbol instance refers to an independent graphic element drawn by the user, including geometric symbols (such as node identifiers of types like rectangles, rhombuses, and circles), connectors (usually represented by arrows or straight lines), or compound symbols (such as decision boxes with text); each symbol instance consists of a set of strokes.
[0043] In this embodiment, the identification of symbol instances may include the identification of symbol type and the identification of symbol location.
[0044] Specifically, a corresponding query vector is generated by encoding features and learnable positional encodings. Then, the query embedding vector generated by the query vector is used to predict the symbol category and obtain the category information of each symbol instance.
[0045] In this embodiment, the stroke embedding vector generated using the encoded feature is fused with the query embedding vector, and then the fused vector is used to predict the symbol position to obtain the corresponding position information.
[0046] The dynamic online handwritten flowchart recognition method provided in this invention extracts the geometric representation of strokes from online stroke data to obtain a stroke feature sequence. An attention masking mechanism is then used to perform streaming inference on the stroke feature sequence through target encoding operations to obtain encoded features. Specifically, the stroke feature sequence is segmented into non-overlapping blocks, each stroke feature within a block is made mutually visible, and the current block only accesses preceding blocks within the history window, while future block information is masked. Finally, symbol instances in the online handwritten flowchart are identified in real time based on the encoded features, and the category and position information of each symbol instance are output. This achieves real-time dynamic streaming symbol segmentation and recognition of online handwritten flowcharts, improving the real-time performance and accuracy of online handwritten flowchart recognition.
[0047] In some embodiments, extracting the geometric representation of strokes from online stroke data to obtain a stroke feature sequence includes: (1) Preprocess the online stroke data by resampling, shifting and height normalization to obtain the processed online stroke data; wherein, the online stroke data includes stroke height information and the resampled stroke sequence.
[0048] Figure 2 This is one of the flowcharts illustrating the stroke feature sequence extraction method provided by the present invention. Figure 2 In the illustrated embodiment, the online stroke data is first resampled using linear interpolation without changing the length of the original strokes, thereby constructing a normalized stroke set with a uniform time step. .
[0049] in, For uniform time steps, It represents the total number of strokes in the flowchart.
[0050] Next, each stroke in the standardized stroke set is translated. Specifically, by subtracting the coordinates of the starting point from the coordinates of each stroke, the stroke is moved to the origin to eliminate the influence of positional information, thus obtaining the stroke set. ; Finally, by dividing the coordinates of each stroke by its height, we obtain a set of strokes with normalized height and originating from the origin. ,in That is, the processed online stroke data.
[0051] (2) A fixed-length feature vector is extracted from the processed online stroke data based on a fixed-length feature extraction network. The fixed-length feature extraction network is obtained by training the encoder-decoder architecture by using the sample online stroke data as the training samples of the Transformer encoder, the sample fixed-length feature vector as the input of the Maskformer decoder network, and the sample stroke coordinate sequence as the output of the Maskformer decoder. The encoder-decoder architecture is trained by minimizing the difference between the sample stroke coordinate sequence and the resampled stroke sequence. The sample stroke coordinate sequence is obtained by reconstructing the sample fixed-length feature vector.
[0052] exist Figure 2 In the embodiment shown, the preprocessed strokes The sequence is padded to a uniform length and then input into the encoder. In the process, geometric features of fixed-length strokes are extracted. As shown in the following formula: ; in, This represents the learnable parameters of a neural network.
[0053] In this embodiment, the encoder The Transformer network architecture is adopted, which is widely used in various sequence-to-sequence tasks due to its powerful sequence modeling capabilities and parallel computing advantages. In particular, this embodiment introduces a padding mask when computing the self-attention mechanism. As shown in the following formula, where Indicates the first step in the flowchart The original length of each stroke sequence.
[0054] ; In this embodiment, the padding mask is used to block out the padding portion in the sequence, ensuring that the model is not affected by invalid data when calculating attention weights; and for different trajectory points of the same stroke, their information should be able to access and reference each other in order to capture the local structure and continuity features inside the stroke, so no look-ahead mask is introduced.
[0055] In this embodiment, the decoder A multilayer perceptron structure is adopted. The output can be explicitly modeled as a multimodal probability distribution as shown in the following equation, which represents the three parameters of the Gaussian mixture model, as shown in the following equation: ; in, It is in A proportion value randomly generated within the range, Indicates the time step in the handwriting sequence; With encoder Output stroke features splicing, as a decoder The input, and thus, the decoder When generating strokes, it can focus on both the current time step and utilize the global context; the decoder Output The amount of components is determined by It represents the center position of the Gaussian component. Define the diagonal elements of the component covariance matrix (assuming each dimension is independent). Represents the mixing coefficient vector (using the Softmax function to ensure...) )composition, The preset total number of components, where Corresponding to a two-dimensional stroke coordinate space.
[0056] In this embodiment, the training objective of the fixed-length feature extraction network is achieved by maximizing the log-likelihood.
[0057] Specifically, assuming the target data is composed of a mixture of multiple Gaussian distributions, and by optimizing the network parameters... To maximize the likelihood probability of generating the target data; this target data can be formally represented as follows: ; in, This represents a Gaussian distribution; by designing target data in the above form, the network model can better capture the distribution characteristics of the data and generate reconstructed strokes.
[0058] In this embodiment, to increase the smoothness and temporal continuity of strokes, and to improve the model's generalization ability and robustness, the loss can be calculated using a decoder. Input random parameters Mapped to the original actual strokes After determining the length, a weighted sum (linear interpolation) is performed on adjacent points to generate a new random point sequence. And calculate the loss based on the sequence.
[0059] In this embodiment, via decoder The set of output parameters yields the reconstructed trajectory. Essentially, it is the calculation of the conditional expectation of a mixture Gaussian distribution, as shown in the following formula: .
[0060] Given a mixing coefficient Under the given conditions, the optimal stroke trajectory estimate in the sense of minimum mean square error is obtained by calculating the expected value. The result is then multiplied by the original stroke height and the stroke starting point is added. This allows us to reconstruct the coordinate sequence of the original strokes.
[0061] (3) Based on the fixed-length feature vector and stroke height information, the geometric representation is obtained, and the stroke feature sequence is obtained according to the geometric representation corresponding to different stroke numbers in the online stroke data.
[0062] In this embodiment, during the data preprocessing stage, the height information of the strokes is filtered out for high normalization purposes. To fully utilize the geometric features of the strokes, the height information of the strokes needs to be reassembled into the feature vector during subsequent feature extraction to restore its complete representation of the stroke's geometric information, as shown in the following formula: ; in, This is the concatenation operator.
[0063] In this embodiment, the stroke feature sequence is represented as follows: ,in, The number of strokes in the flowchart is used; furthermore, automatic stroke features are determined by measuring the difference between the strokes in the original flowchart and the strokes in the reconstructed flowchart. The quality.
[0064] Figure 3 This is the second flowchart illustrating the stroke feature sequence extraction method provided by this invention. Figure 3 In the illustrated embodiment, a fixed-length feature extraction network trained using the encoder-decoder architecture described above is used to reconstruct the strokes of the flowchart, that is, to reconstruct the stroke sequence in the flowchart. S 1. S 1. ... S n As input, a fixed-length geometric feature representation of each stroke is extracted by the encoder. v 1. v 2. ... v n Subsequently, the extracted features are reconstructed into the original stroke sequence using a Gaussian mixture model through a decoder, generating the reconstructed flowchart strokes. , … Finally, the automatic stroke features are determined by measuring the differences between the strokes of the original flowchart and the reconstructed flowchart. The quality.
[0065] The dynamic online handwritten flowchart recognition method provided in this invention preprocesses the online stroke data by resampling, shifting, and height normalizing it. Then, a fixed-length feature vector is extracted from the processed online stroke data using a fixed-length feature extraction network. Geometric representations are obtained based on these fixed-length feature vectors and stroke height information. Finally, stroke feature sequences are obtained according to the geometric representations corresponding to different numbers of strokes in the online stroke data. This method maintains the advantages of unsupervised training and provides better stroke geometric feature representations for stroke classification and symbol segmentation, accurately capturing the inherent geometric information of the strokes.
[0066] In some embodiments, the target encoding operation is implemented through a Transformer encoder, which includes a multi-head attention layer. The stroke feature sequence is subjected to streaming inference using an attention masking mechanism through the target encoding operation to obtain encoded features, including: relative temporal encoding of the stroke feature sequence to obtain a temporal encoding result; relative spatial encoding of the stroke feature sequence to obtain a spatial encoding result; multi-head attention calculation based on the temporal and spatial encoding results, and using an attention mask in the multi-head attention layer to control the visibility range of each stroke feature, and using a padding mask to block invalid stroke features, to obtain the multi-head attention calculation result; the multi-head attention calculation result is then subjected to layer normalization and input into a feedforward neural network, and the output of the feedforward neural network is subjected to layer normalization again to obtain the encoded features.
[0067] In this embodiment, the contextual relationship of strokes is introduced by encoding relative time and space.
[0068] Specifically, a relative time encoding with a pruning function is used. For temporal distances within a certain range, a unique code is used to represent that temporal distance; for stroke pairs outside the distance range, the relationship between them is weak, so they are not distinguished; for stroke pairs... and ,use A learnable vector To represent relative timing positions, specific encoding It can be expressed by the following formula: ; ; Since spatial relationships are important contextual information, this embodiment selects the centroid of the stroke to represent the stroke position and uses a polar coordinate system to represent the point position. Compared with the Cartesian coordinate system, polar coordinates can divide the area near the origin into a small area and the area far from the origin into a large area.
[0069] In this embodiment, the growth trend of the region area is equal in all directions of the polar coordinate system; specifically, this embodiment uses... Establish a polar coordinate system for the origin ( and Indicates strokes and (center of mass).
[0070] Specifically, for angle encoding, Divided into equal parts Each of the angular intervals corresponds to an angular code. For distance encoding, the maximum distance is divided into equal parts. Segments, similar to time encoding, use Distance encoding To indicate; express and Angles in polar coordinates express and The normalized distance; the calculation formula is as follows: ; ; ; ; ; in, For strokes and The relative angle encoding between them For strokes and Relative distance encoding between them Indicates strokes and The normalized distance from the center of mass. This is the final relative spatial encoding.
[0071] In this embodiment, since the encoder needs to batch process multiple flowcharts at once, the embedding vectors of flowcharts with fewer strokes need to be padded with zero vectors so that their length is consistent with the flowchart with the longest stroke in the dataset.
[0072] Specifically, assuming the flowchart with the most strokes in a batch has the following stroke count... For a character with a number of strokes... Flowchart (where If ), then add to the end of its embedding vector. There are zero vectors; to avoid the padding affecting model training, the padding positions are usually processed with a padding mask when calculating self-attention. Similar to the attention mask, it needs to be directly added to the attention weight matrix of the Transformer.
[0073] It should be noted that the Transformer network, due to its parallel processing capabilities and global context modeling advantages, can be used as the encoder in this embodiment for temporal and spatial encoding.
[0074] Figure 4 This is a schematic diagram illustrating the working mechanism of the attention mask provided by the present invention. Figure 4In the illustrated embodiment, when creating the mask matrix using the Transformer encoder, the input stroke sequence is first divided into non-overlapping blocks of size 3, and then the matrix is constructed according to the following rules: (1) Strokes within the same block can be seen from each other. For example, when calculating strokes... When the encoder outputs, all frames belonging to the same block, including All of these should be included in attention calculation. You can see the next two strokes, and The future cannot be foreseen. Therefore, the average number of forward-looking strokes for each stroke is half the block size; (2) If the two strokes are in different blocks, the left one cannot see the right one when calculating attention; for example, Cannot be used This information. In this way, the block boundaries can strictly limit the receptive field, avoiding a linear increase in the receptive field as the model deepens; (3) If two strokes are in different blocks and their distance is less than the history window size, then the stroke on the right can see the one on the left. This method causes the receptive field on the left to increase as the model deepens; setting the history window size to 3, so in the third layer, You can see But you can't see it. .
[0075] After calculating the multi-head attention calculation result using the above method, the multi-head attention calculation result is subjected to layer normalization and input into the feedforward neural network. Finally, the output of the feedforward neural network is subjected to layer normalization again to obtain the encoded features.
[0076] Figure 5 This is a schematic diagram of the streaming inference process provided by the present invention. Figure 5 In the illustrated embodiment, L× represents an L-layer Transformer network. Stroke features are input to a normalization layer (LN), and the normalized stroke features are then input to a multi-head attention layer (MHA) for relative temporal and relative spatial encoding. During the encoding process, streaming inference is achieved through attention masks (ensuring that strokes within the same block are mutually visible, and only allowing access to previous blocks within the history window). Invalid strokes are masked through padding masks (such as zero padding). The encoded output is normalized again through the LN layer, and the processing result is input to a feedforward network (FFN). Finally, the outputs of the L feedforward networks are concatenated to obtain the final encoded features.
[0077] The dynamic online handwritten flowchart recognition method provided in this invention performs relative temporal encoding and relative spatial encoding on the stroke feature sequence. Multi-head attention calculation is performed based on the temporal and spatial encoding results. In the multi-head attention layer, attention masks are used to control the visibility range of each stroke feature, and padding masks are used to block invalid stroke features, resulting in the multi-head attention calculation results. After layer normalization, the multi-head attention calculation results are input into a feedforward neural network to obtain the corresponding encoded features. By simulating the streaming stroke input process through a stroke mask mechanism, the following information of incomplete strokes can be blocked, allowing the model to infer based solely on current and historical stroke points. This satisfies the low-latency requirement of real-time interaction while ensuring the causality of the recognition process.
[0078] In some embodiments, the real-time identification of symbol instances in the online handwriting flowchart based on encoded features, and the output of category and location information for each symbol instance, includes: generating a query embedding vector based on the encoded features and multiple positional encoded vectors using a Maskformer decoder; processing the query embedding vector using a multilayer perceptron to obtain a mask embedding vector, and processing the encoded features using a multilayer perceptron to obtain a stroke embedding vector; merging the stroke embedding vector and the mask embedding vector to obtain a predicted binary mask; predicting the symbol category based on the mask embedding vector using a target classifier to obtain category information; and performing symbol segmentation mask prediction on the predicted binary mask to obtain location information; wherein the target classifier is determined based on a linear classifier and a Softmax function.
[0079] In this embodiment, since MaskFormer has the advantage of instance segmentation compared to Transformer network, MaskFormer can be used as the decoder.
[0080] In this embodiment, the decoder section uses a standard Maskformer decoder, based on the encoder output features. and A learnable position code is used to generate a query vector for computation; the position code is generated using alternating sine and cosine functions, as shown in the following formula: ; ; ; in, This represents normalized location information, where the location index... This indicates the total number of strokes. Indicates the dimension of the query vector. Indicates the frequency base.
[0081] In this embodiment, the query embedding vector is obtained in the following way: Figure 6 This is a flowchart illustrating the symbol instance recognition method provided by the present invention. Figure 6 In the illustrated embodiment, the Transformer network aggregates contextual information from the encoder through a cross-attention mechanism, dynamically adjusts the focus of each query, and finally outputs... Each query embedding vector has a dimension of . These dimensions are Each query embedding vector encodes global information about the corresponding flowchart predicted by MaskFormer; similar to the standard MaskFormer network, this decoder generates predictions for symbol instance classification and symbol instance segmentation in parallel.
[0082] exist Figure 6 In the illustrated embodiment, during the symbol category prediction stage, the output of the Transformer network is... Each of the embedding vectors is passed through a multilayer perceptron to obtain a dimension of... The vector is used to obtain the class prediction probability for each query through a linear classifier and the Softmax function. , express 3D probability simplex Indicates the number of symbol categories.
[0083] It should be noted that the linear classifier additionally predicted a "no object" category. In the original MaskFormer network These represent embedding vectors that do not match any actual region, and often contain the image's background; in this embodiment, since the online handwriting flowchart does not distinguish between foreground and background... This represents an embedding vector that indicates a segmentation error or an incomplete symbol, i.e., a symbol that does not match any actual symbol.
[0084] exist Figure 6 In the illustrated embodiment, during the symbol segmentation mask prediction stage, the output of the Transformer network is... Each embedding vector is passed through a multilayer perceptron to obtain a mask embedding vector. Additionally, the encoder output Stroke-level embedding is obtained through a multilayer perceptron. Finally, the mask is embedded into the vector. and stroke-level embedding Performing a dot product operation and then using the sigmoid function yields the sign-separated binary mask. As shown in the following formula: ; in, Indicates the first Each prediction mask embedding vector, i.e. . Representing the The predicted binary mask for each stroke, i.e. Indicates taking the matrix All rows List, This represents the transpose of a vector.
[0085] In this embodiment, during the encoding stage, a Transformer encoder is used to process temporal stroke features through a masked self-attention layer, preserving long-range dependencies of completed strokes while avoiding the leakage of unknown stroke information. During the decoding stage, the Maskformer decoder predicts symbol categories and segments binary masks in parallel based on the encoded temporal features through learnable position queries. This design enables the model to simulate the following real-world writing scenario: when a user is drawing a symbol, the model makes real-time inferences based only on the already input strokes, without being affected by subsequent unwritten strokes.
[0086] The dynamic online handwritten flowchart recognition method provided in this invention generates a query embedding vector based on encoded features and multiple positional encoded vectors using a Maskformer decoder; processes the query embedding vector using a multilayer perceptron to obtain a mask embedding vector; processes the encoded features using a multilayer perceptron to obtain a stroke embedding vector; merges the stroke embedding vector and the mask embedding vector to obtain a predicted binary mask; predicts the symbol category based on the mask embedding vector using a target classifier to obtain category information; and performs symbol segmentation mask prediction on the predicted binary mask to obtain position information. This time-aware masking strategy enables streaming processing for flowchart recognition and allows for real-time application of the online handwritten flowchart recognition algorithm on mobile devices and embedded systems, providing a real-time writing experience for scenarios such as human-computer interaction, intelligent education, and office automation.
[0087] In some embodiments, the multiple location encoding vectors include query vectors and additional query vectors; the target classifier is obtained through the following steps: the target classifier is trained on the DETR architecture using sample radix masks and sample mask embedding vectors as training samples, joint loss as the loss function, and bipartite graph matching mechanism according to one-to-one matching and one-to-many matching methods; wherein, the one-to-one matching method is achieved by optimally matching K query vectors with m real symbols using the Hungarian algorithm; the one-to-many matching method is achieved by copying m real symbols n times and optimally matching K query vectors with m×n symbols using the Hungarian algorithm; the joint loss is determined based on the one-to-one matching loss, the one-to-many matching loss function, and the auxiliary loss.
[0088] It should be noted that when using the DETR (Detection Transformer) model to implement a one-to-one matching mechanism, there is a problem of positive sample sparsity. This results in only a small number of queries being assigned as positive samples. This sample allocation strategy significantly reduces the proportion of effective positive samples during model training, making the supervision signal too sparse and thus making it difficult to fully optimize the model parameters. Especially in the early stage of training, due to the extreme imbalance between positive and negative samples, a large number of queries are classified as negative samples, making it difficult for the model to obtain enough gradient information to effectively learn the key features of the object detection task.
[0089] To address this issue, this embodiment introduces an improved hybrid matching strategy to obtain more sufficient supervision signals to improve performance during each iteration of the neural network. Specifically, a bipartite graph matching mechanism based on the Hungarian algorithm is used to achieve a strict one-to-one and one-to-many correspondence between predicted bounding boxes and labeled boxes, that is, the DETR architecture is iteratively trained through set prediction.
[0090] In this embodiment, the Hungarian algorithm is the core mechanism for achieving optimal matching between the predicted object query and the true value during the training process. This algorithm uses a bipartite graph matching strategy to uniquely assign each true value to a query while minimizing the matching cost matrix. In this embodiment, the improved hybrid matching strategy can be applied to the field of online handwritten flowchart symbol recognition; for example, when initializing the query vector, by additionally introducing... There are query vectors, totaling Query vectors; Previous Each query can be viewed as a group, used to perform a one-to-one match; later Each query can be viewed as a group for performing one-to-many matching. One-to-many matching refers to assigning multiple queries to the same ground truth during the training phase to increase positive sample supervision signals and improve model learning efficiency.
[0091] In this embodiment, only one-to-one matching queries are needed during streaming inference, while one-to-many matching queries are only used to obtain more supervision signals. To prevent information leakage between the two sets of query vectors during training, this embodiment introduces a query mask. As shown in the following formula: ; Among them, when querying and query When they belong to the same one-to-one matching group or one-to-many matching group That is, it allows self-attention computation; when querying and query When they belong to different groups This makes the attention weight at the corresponding position after Softmax zero, completely blocking the flow of information between query groups.
[0092] It should be noted that when the Transformer decoder calculates the attention weight matrix, in addition to the query mask, an attention mask must also be added to mask irrelevant strokes, and a padding mask must be added to mask the impact of the padding zero vectors on training; during loss calculation, queries in one-to-one matching groups are matched using the Hungarian algorithm. Group truth values, while queries in one-to-many matching groups require more supervisory signals, therefore... Group truth copying The samples are then matched using the Hungarian algorithm.
[0093] Since the improved strategy of introducing hybrid matching (i.e., introducing a one-to-many matching method) into the Hungarian algorithm can provide more supervision signals to the DETR architecture to improve model performance, the DETR architecture can achieve a strict one-to-one correspondence between predicted bounding boxes and labeled boxes by using the bipartite graph matching mechanism based on the Hungarian algorithm during the training of the target classifier.
[0094] In this embodiment, the one-to-one matching loss and the one-to-many matching loss are determined by the category prediction cost, the binary mask loss cost, and the focus loss cost.
[0095] In this embodiment, the cost matrix of a one-to-one matching group The cost of class prediction is defined by the following formula. Binary mask loss cost and focus loss cost It is defined in three parts and expressed by the following formula: ; in, , and These are the weighting coefficients for each part; category prediction cost. Defined by the following formula: in, The model is for the true category. Predicted probability The negative sign indicates that the higher the probability, the lower the matching cost. For indicator functions, when the category is not... The value is 1; Binary mask loss cost Defined by the following formula: in, The binary mask representing the model's prediction. It is a real binary mask; Dice loss; focus loss cost It can be expressed by the following formula: ; in, The focus loss is detailed in the following formula. Cost matrix for one-to-many matching groups. It is defined by the following formula, where is the weight coefficient of each part. , and The cost matrix parameters are shared with the one-to-one matching group.
[0096] ; In this embodiment, a one-to-one matching group The query uses Hungarian matching and Matching one true value to calculate the loss; one-to-many matching groups The query uses Hungarian matching and Match each true value to calculate the loss.
[0097] Below, we will introduce the calculation methods for one-to-one matching loss and one-to-many matching loss according to different task categories: (1) For symbol classification tasks, use the weighted cross-entropy loss function. As shown in the following formula: ; in, It is the total number of symbols; yes The true label of each sample; The model is for the first The predicted probability of each sample is calculated by the softmax function from the output of the symbol classification. It is the first The weighting coefficients for each sample are the reciprocal of the proportion of the number of symbols belonging to that category in the training samples. It is worth noting that for the "no object" category... Since its proportion is much larger than that of the symbol, its weight is uniformly set to 0.1 during training.
[0098] Symbol segmentation mask prediction consists of two parts: Dice loss. and focus loss; among which, It is used to measure the similarity between masks; the smaller the value, the higher the similarity. To address the class imbalance problem, higher weights are assigned to difficult samples; the above losses are shown in the following formulas: ; ; ; ; ; Among them, middle, Represents the predicted binary mask. It is a true binary mask; The smoothing coefficient is set to... Summation operation It sums up all positions of the binary mask; It is the average number of symbols in each flowchart in the training data; Indicates the category balance weight; Represents a dynamic probability term; Represents the binary cross-entropy function; where, It is the positive sample weight constant. It is the adjustment constant for easy and difficult samples; in focus loss In the middle, category balance weight Adjusting the loss ratio between positive and negative samples resolves the class imbalance problem; dynamic probability term. Adjustment constant for easy and difficult samples It can automatically reduce the weight of easily classified samples, allowing the model to focus on difficult samples. The two work together to improve the segmentation performance of small targets and sparse regions while maintaining accuracy.
[0099] For the one-to-one matching method, the final loss function The weighted cross-entropy loss function Dice loss Focus loss The weighted sum of the three is shown in the following formula: ; in, , and These are the weight coefficients of each part, which are shared with the weight parameters of each cost matrix during Hungarian matching.
[0100] In this embodiment, the loss function for one-to-many matching The weighted cross-entropy loss function Dice loss Focus loss The weighted sum of the three, only it needs to be... Group truth copying After matching, the loss is calculated using the Hungarian algorithm.
[0101] In this embodiment, to effectively address the issues of sparse positive samples, slow convergence, unstable matching, and insufficient optimization of deep networks during model training, an auxiliary loss is introduced. , In each layer of the Transformer decoder, classification loss and symbol segmentation mask prediction loss are calculated for one-to-one matching and one-to-many matching. This multi-layer supervision mechanism can enhance gradient backpropagation and improve model performance.
[0102] In this embodiment, joint loss One-to-one matching loss function One-to-many matching loss function and auxiliary losses The sum of the three is shown in the following formula: ; During reasoning, this embodiment only uses the previous... One-to-one matching of each query is performed using the Hungarian algorithm for symbol recognition, eliminating the need for post-processing steps.
[0103] The dynamic online handwritten flowchart recognition method provided in this invention uses sample radix masks and sample mask embedding vectors as training samples, joint loss as the loss function, and a bipartite graph matching mechanism to train the DETR architecture to obtain a target classifier according to one-to-one matching and one-to-many matching methods, thereby ensuring the inference efficiency of the DETR architecture and improving the inference accuracy.
[0104] Figure 7 This is the second schematic diagram of the dynamic online handwritten flowchart recognition method provided by the present invention. Figure 7In the illustrated embodiment, online stroke data (corresponding to the original flowchart) is extracted using unsupervised stroke feature extraction to obtain a stroke feature sequence. This sequence is then encoded using a Transformer encoder to obtain encoded features. Next, on one hand, a query embedding vector is generated using the Transformer encoder based on the query vector, an additional query vector, and the encoded features. This query embedding vector is then processed by a multilayer perceptron to obtain a mask embedding vector, which is used for symbol class prediction. On the other hand, the encoded features are processed by another multilayer perceptron to obtain a stroke embedding vector. This stroke embedding vector is then fused with the aforementioned mask embedding vector to obtain a predicted binary mask, which is then used for dynamic symbol segmentation mask prediction. Finally, the recognized flowchart is obtained based on the symbol class prediction result (category information) and the dynamic symbol segmentation mask prediction result (position information).
[0105] The following describes the dynamic online handwritten flowchart recognition device provided by the present invention. The dynamic online handwritten flowchart recognition device described below can be referred to in correspondence with the dynamic online handwritten flowchart recognition method described above.
[0106] Figure 8 This is a schematic diagram of the structure of the dynamic online handwritten flowchart recognition device provided by the present invention, as shown below. Figure 8 As shown, the dynamic online handwritten flowchart recognition device includes: a feature extraction module 810, a streaming reasoning module 820, and a recognition module 830.
[0107] The feature extraction module 810 is used to extract the geometric representation of strokes from online stroke data to obtain a stroke feature sequence; The streaming inference module 820 is used to perform streaming inference on the stroke feature sequence through target encoding operation using an attention masking mechanism to obtain encoded features. The target encoding operation includes dividing the stroke feature sequence into non-overlapping blocks, setting each stroke feature within a block to be visible to each other, setting the current block to only access the previous block in the history window, and masking information of future blocks. The recognition module 830 is used to identify symbol instances in the online handwritten flowchart in real time based on the encoding features, and output the category information and position information of each symbol instance.
[0108] The dynamic online handwritten flowchart recognition device provided in this invention extracts the geometric representation of strokes from online stroke data to obtain a stroke feature sequence. It then uses an attention masking mechanism to perform streaming inference on the stroke feature sequence through target encoding operations to obtain encoded features. Specifically, the stroke feature sequence is segmented into non-overlapping blocks, each stroke feature within a block is made mutually visible, and the current block only accesses preceding blocks within the history window, while future block information is masked. Finally, based on the encoded features, symbol instances in the online handwritten flowchart are identified in real time, and the category and position information of each symbol instance are output. This achieves real-time dynamic streaming symbol segmentation and recognition of online handwritten flowcharts, improving the real-time performance and accuracy of online handwritten flowchart recognition.
[0109] Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9 As shown, the electronic device may include a processor 910, a communications interface 920, a memory 930, and a communication bus 940. The processor 910, communications interface 920, and memory 930 communicate with each other via the communication bus 940. The processor 910 can call logical instructions in the memory 930 to execute a dynamic online handwritten flowchart recognition method. This method includes: extracting the geometric representation of strokes from online stroke data to obtain a stroke feature sequence; performing streaming inference on the stroke feature sequence using an attention masking mechanism through target encoding operations to obtain encoded features; the target encoding operations include segmenting the stroke feature sequence into non-overlapping blocks, setting each stroke feature within a block to be mutually visible, and setting the current block to only access previous blocks within the history window, while masking future block information; and recognizing symbol instances in the online handwritten flowchart in real time based on the encoded features, outputting the category and position information of each symbol instance.
[0110] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0111] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the dynamic online handwritten flowchart recognition method provided by the above methods. The method includes: extracting the geometric representation of strokes from online stroke data to obtain a stroke feature sequence; performing streaming inference on the stroke feature sequence through a target encoding operation using an attention masking mechanism to obtain encoded features; the target encoding operation includes dividing the stroke feature sequence into non-overlapping blocks, setting each stroke feature within a block to be mutually visible, and setting the current block to only access the preceding blocks in the history window and masking future block information; and identifying symbol instances in the online handwritten flowchart in real time based on the encoded features, and outputting the category information and position information of each symbol instance.
[0112] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the dynamic online handwritten flowchart recognition method provided by the above methods. The method includes: extracting the geometric representation of strokes from online stroke data to obtain a stroke feature sequence; performing streaming inference on the stroke feature sequence through a target encoding operation using an attention masking mechanism to obtain encoded features; the target encoding operation includes dividing the stroke feature sequence into non-overlapping blocks, setting each stroke feature within a block to be mutually visible, and setting the current block to only access previous blocks within the history window and masking future block information; and identifying symbol instances in the online handwritten flowchart in real time based on the encoded features, and outputting the category information and position information of each symbol instance.
[0113] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0114] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for recognizing dynamically online handwritten flowcharts, characterized in that, include: Geometric representations of strokes are extracted from online stroke data to obtain stroke feature sequences; An attention masking mechanism is used to perform streaming inference on the stroke feature sequence through a target encoding operation to obtain encoded features. The target encoding operation includes dividing the stroke feature sequence into non-overlapping blocks, setting each stroke feature within a block to be mutually visible, setting the current block to only access the preceding blocks in the history window, and masking information of future blocks. Based on the encoded features, symbol instances in the online handwriting flowchart are identified in real time, and the category and location information of each symbol instance are output.
2. The dynamic online handwritten flowchart recognition method according to claim 1, characterized in that, The extraction of geometric representations of strokes from online stroke data to obtain stroke feature sequences includes: The online stroke data is preprocessed by resampling, shifting, and height normalization to obtain processed online stroke data; wherein, the online stroke data includes stroke height information and the resampled stroke sequence; A fixed-length feature extraction network extracts fixed-length feature vectors from the processed online stroke data. This network uses sample online stroke data as training samples for a Transformer encoder, the sample fixed-length feature vectors as input to a Maskformer decoder network, and the sample stroke coordinate sequence as the output of the Maskformer decoder. The encoder-decoder architecture is trained by minimizing the difference between the sample stroke coordinate sequence and the resampled stroke sequence. The sample stroke coordinate sequence is reconstructed from the sample fixed-length feature vectors. The geometric representation is obtained based on the fixed-length feature vector and the stroke height information, and the stroke feature sequence is obtained according to the geometric representation corresponding to different numbers of strokes in the online stroke data.
3. The dynamic online handwritten flowchart recognition method according to claim 1, characterized in that, The target encoding operation is implemented through a Transformer encoder, which includes a multi-head attention layer; The attention masking mechanism is used to perform streaming inference on the stroke feature sequence through target encoding operations to obtain encoded features including: The stroke feature sequence is subjected to relative temporal encoding to obtain a temporal encoding result; the stroke feature sequence is subjected to relative spatial encoding to obtain a spatial encoding result. Multi-head attention calculation is performed based on the temporal and spatial encoding results. In the multi-head attention layer, the attention mask is used to control the visibility range of each stroke feature, and a filling mask is used to block invalid stroke features to obtain the multi-head attention calculation result. The multi-head attention calculation result is processed by layer normalization and then input into the feedforward neural network. The output of the feedforward neural network is then processed by layer normalization again to obtain the encoded features.
4. The dynamic online handwritten flowchart recognition method according to claim 1, characterized in that, The step of identifying symbol instances in the online handwriting flowchart in real time based on the encoded features, and outputting the category and location information of each symbol instance, includes: Based on the Maskformer decoder, a query embedding vector is generated according to the encoded features and multiple positional encoded vectors; The query embedding vector is processed using a multilayer perceptron to obtain a mask embedding vector, and the encoded features are processed using a multilayer perceptron to obtain a stroke embedding vector. The stroke embedding vector and the mask embedding vector are merged to obtain the predicted binary mask; The target classifier predicts the symbol category based on the mask embedding vector to obtain the category information; the predicted binary mask is then used for symbol segmentation mask prediction to obtain the position information; wherein, the target classifier is determined based on a linear classifier and a Softmax function.
5. The dynamic online handwritten flowchart recognition method according to claim 4, characterized in that, The plurality of location encoding vectors includes a query vector and an additional query vector; The target classifier is obtained through the following steps: The target classifier is trained using sample radix masks and sample mask embedding vectors as training samples, joint loss as the loss function, and the DETR architecture is trained using a bipartite graph matching mechanism based on one-to-one matching and one-to-many matching methods. The one-to-one matching method is achieved by using the Hungarian algorithm to optimally match K query vectors with m real symbols; the one-to-many matching method is achieved by copying the m real symbols n times and using the Hungarian algorithm to optimally match the K query vectors with m×n symbols; the joint loss is determined based on the one-to-one matching loss function, the one-to-many matching loss function, and the auxiliary loss.
6. The dynamic online handwritten flowchart recognition method according to claim 5, characterized in that, The one-to-one matching loss and the one-to-many matching loss are determined by the category prediction cost, the binary mask loss cost, and the focus loss cost.
7. A dynamic online handwritten flowchart recognition device, characterized in that, include: The feature extraction module is used to extract the geometric representation of strokes from online stroke data to obtain stroke feature sequences; The streaming inference module is used to perform streaming inference on the stroke feature sequence through target encoding operation using an attention masking mechanism to obtain encoded features; the target encoding operation includes dividing the stroke feature sequence into non-overlapping blocks, setting each stroke feature within a block to be mutually visible, setting the current block to only access the preceding blocks in the history window, and masking information of future blocks; The recognition module is used to identify symbol instances in the online handwriting flowchart in real time based on the encoded features, and output the category information and position information of each symbol instance.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the dynamic online handwritten flowchart recognition method as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the dynamic online handwritten flowchart recognition method as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the dynamic online handwritten flowchart recognition method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Online hand-drawn table cell type identification method and device
CN117953523A
Streaming voice conversion method based on block masking
CN119132321A
KR20250038564A