Online handwritten signature character segmentation method, system, storage medium and electronic device

The end-to-end segmentation of online handwritten characters through deep learning technology solves the problem of low user feedback and recognition accuracy in the prior art, and achieves higher segmentation accuracy and recognition accuracy, which is suitable for high-demand application scenarios and reduces costs.

CN114283417BActive Publication Date: 2025-05-09CHONGQING AOXIONG INFORMATION TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111540191.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-05-09
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The existing online handwritten character recognition technology requires user feedback information, the information granularity is the stroke order level, the recognition accuracy is low, the applicable scenarios are limited, and the acquisition cost is high.

Method used

Deep learning technology is used to form an end-to-end solution for online handwritten characters segmentation, input the acquired sequence of sample point data (such as coordinates, pressure values, time, etc.) to directly obtain the character segmentation result.

Benefits of technology

It improves segmentation accuracy and signature recognition accuracy, and is suitable for application scenarios with high requirements for identification accuracy, such as judicial appraisal of handwriting, reducing acquisition costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283417B_ABST
    Figure CN114283417B_ABST
Patent Text Reader

Abstract

The present invention seeks protection for an online handwritten signature character segmentation method. An end-to-end solution for online handwritten character segmentation is formed using deep learning technology. The collected signature stroke data is preprocessed to obtain a stroke feature sequence of paired data to form a data set. The data set is input into a basic prediction model for training. The output of the basic prediction model is a probability distribution vector of the same length as the stroke feature sequence, which is the label prediction probability corresponding to the signature stroke feature sequence position. The index value corresponding to the maximum value is taken, and the index value is mapped to a label according to the mapping relationship. Then, the point index corresponding to each signature is obtained according to the label. By inputting the sampling point data sequence obtained by the relevant writing device, the segmentation result of the character can be directly obtained. This provides basic support for the construction of the character library and the subsequent character content recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of online handwritten character recognition, and in particular to a method for segmenting online handwritten characters. Background Art

[0002] Application number CN202010590507.2, invention name "A handwritten formula recognition method based on an end-to-end network model" discloses a handwritten formula recognition method based on an end-to-end network model, collects a handwritten formula recognition picture data set, preprocesses the handwritten formula recognition picture data set; establishes an encoder-decoder model, the encoder-decoder model includes a feature extraction layer, an encoder encoder and a decoder decoder; uses the preprocessed handwritten formula recognition picture data set to train the encoder-decoder model to obtain the encoder-decoder model after training; uses the encoder-decoder model after training to recognize the handwritten formula picture to be recognized, and obtains the recognition result of the handwritten formula picture to be recognized. The present invention uses different perceptions to extract the size features of the handwritten formula picture, completes the word vector encoding and the docking and decoding operations with the label for different features, optimizes the network hyperparameters, avoids the problem of partial text missing due to the low resolution of the feature map after multiple convolutions, can better recognize handwritten formulas, and has a high recognition accuracy. Application number CN202110225996.6, invention name "An online handwritten formula recognition method and device based on user feedback information", discloses an online handwritten formula recognition method and device based on user feedback information, which introduces user participation such as deletion operations, stroke-filling operations and / or structural movement operations into the existing recognition methods. With the help of the idea of ​​human-computer hybrid intelligence, user feedback information is integrated into the different stages of the "character segmentation-character recognition-structural analysis" recognition method, and an interactive technology suitable for user writing and error correction is designed. The present invention designs an interactive means suitable for sketch recognition, avoiding various problems encountered by formula recognition methods based on image processing, and providing basic guarantees for users to modify strokes with typos or ambiguities, structural errors in formulas, etc., thereby improving the formula recognition rate and meeting user needs. Patent application number: CN202110488219.0, invention name "Processing method, device, system and storage medium based on smart pen handwritten images", provides a processing method, device, system and storage medium based on smart pen handwritten images, which are used to improve the access efficiency of smart pen handwritten images based on online education.The processing method based on smart pen handwriting images includes: using a preset stroke order recognition neural network model to perform character segmentation, multi-layer stroke order feature extraction, character recognition and character information generation on a sequence of smart pen handwriting images based on online education in sequence to obtain stroke order recognition information; filtering the stroke order recognition information and the sequence of smart pen handwriting images based on online education to obtain preprocessed data; obtaining the file size of the preprocessed data, and comparing and analyzing the file size with a preset threshold to obtain a comparison analysis result; according to the comparison analysis result, converting the preprocessed data into a binary format and / or slicing the preprocessed data to obtain data to be stored; storing the data to be stored in a preset database based on distributed file storage.

[0003] In the process of recognizing handwritten input formulas, the technology disclosed in the above-mentioned prior art partially uses offline handwriting data for formula recognition. Compared with the information that can be used by online handwriting data, information that is helpful for recognition, such as pressure and speed, is not taken into account, which restricts the recognition effect of handwritten electronic signatures. In addition, the recognition effect is further improved through some interactive means with users (filling in strokes, deleting). Although the performance of the recognition effect is improved, it is necessary for users to feedback information during the recognition process, which increases user operations. The granularity of the input information is at the stroke order level, and the recognition accuracy is low, which limits its applicability to scenarios. It is not suitable for scenarios such as forensic identification and handwritten content recognition where the cost of obtaining stroke order information is high. Summary of the invention

[0004] The present invention aims to address the problems in the prior art of online writing character processing, such as the need for user feedback information on the recognition effect, the information granularity being at the stroke order level, the low recognition accuracy, the limited applicable scenarios, and the high acquisition cost.

[0005] The technical solution of the present invention to solve the above technical problems is to use deep learning technology to form an end-to-end solution for online handwritten character segmentation. A sequence of sampling point data (such as coordinates, pressure values, time and other information) obtained by a related writing device (such as a signature plate) is input to directly obtain the segmentation results of the characters. This provides basic support for the construction of the character library and subsequent character content recognition. Compared with the segmentation method based only on image input, this technical solution takes into account information such as time, pressure and writing order, and further improves the accuracy of segmentation.

[0006] Specifically, an online handwritten signature character segmentation method is provided, wherein the collected signature stroke data is preprocessed to obtain a stroke feature sequence including stroke data features and paired data of labels to form a data set, and the data set is divided into a training set and a verification set to be input into a basic prediction model for training; the preprocessed signature data features are input into the basic prediction model, and the output of the basic prediction model is a probability distribution vector of the same length as the stroke feature sequence, which is the label prediction probability corresponding to the position of the signature stroke feature sequence, and the index value corresponding to the maximum value is taken, and the index value is mapped to a label according to the mapping relationship, and then the point index corresponding to each signature is obtained according to the label, and the point to which each signature belongs is obtained, and different characters are segmented from the signature stroke feature sequence according to the point. That is, the purpose of segmenting different characters from a sequence is achieved.

[0007] A further preferred solution is to perform data normalization, downsampling, padding, label data conversion, data collection, and data splitting processing on the collected signature stroke data.

[0008] In a further preferred embodiment, data normalization includes: normalizing the signature stroke data according to the formula: range max =2·max((X mean -X min ), (Y mean -Y min ))Calculate the scaling factor range max , according to the formula: Normalize the horizontal coordinate (X) and vertical coordinate (Y) of the stroke data, where X min With Y min Respectively represent the minimum value of the horizontal and vertical coordinate sequences of an online handwriting data, X mean and Y mean Respectively represent the average values ​​of the corresponding sequences of the horizontal and vertical axes, X norm and Y norm Respectively represent the normalized horizontal and vertical coordinate sequences, range max range max =2·max((X mean -X min ), (Y mean -Y min ))Convert the time of the stroke data into a fixed unit of duration to perform time normalization processing.

[0009] A further preferred solution is that the downsampling is specifically to determine the stroke breakpoints through the pen lifting action represented by the writing pressure value change information, retain the pressure value of each breakpoint and the two points before and after it, and obtain the stroke pressure value sequence after downsampling by sampling the remaining unselected stroke data points at intervals.

[0010] In a further preferred embodiment, the filling and label data conversion operation is to perform a filling operation on the stroke pressure value sequence after downsampling, fill a portion greater than a predetermined length of the target sequence with a value of 0, and convert the label value into a BIO format for filling.

[0011] The BIO format is: the first point at the beginning of the sequence points of the stroke feature of a character is represented by B, and the subsequent component sequence points are represented by I. The next character is also represented by B at the beginning of the first sequence point, and the subsequent component sequence points are also represented by I. This is repeated until the valid sequence points of the signature strokes end. If there is a filling operation in the stroke data sequence, the label at the corresponding position is set to O.

[0012] Through the above-mentioned preprocessing operations such as data normalization, downsampling, padding, label data conversion, etc., paired data including stroke data features and labels are obtained to form a data set.

[0013] In a further preferred embodiment, the downsampling includes: extracting strokes, removing adjacent points, removing straight line points, and homogenizing line density, specifically including: extracting strokes by determining the start and end of each stroke through the pressure value of the pen lift point; removing strokes that satisfy the formula within the same stroke:

[0014] The i-th pixel (x i ,y i ), x i is the horizontal coordinate of the i-th pixel in the stroke, y i is the ordinate of the i-th pixel, α is the proportional coefficient, T dist is the distance threshold; for a straight stroke or a straight part of a stroke, retain its two end points; according to the formula:

[0015] Determine the line density threshold, set a sliding window, and discard the points in the window whose line density meets the line density threshold. β is the density ratio coefficient.

[0016] The data set is split into a training set and a validation set according to a certain ratio and input into the basic prediction model for training. The model includes a bidirectional long short-term memory network (Bi-lstm), a fully connected layer (Dense) and a conditional random field (CRF). Among them, Bi-Lstm is responsible for inputting the encoding information, and the signature stroke feature sequence is input from the model input layer as the encoding information. The Dense layer converts the encoder output dimension into the label dimension, and CRF is responsible for obtaining the label output. The output of the basic prediction model is a probability distribution vector of the same length as the input feature sequence, which is the label prediction probability corresponding to the position of the signature stroke data feature sequence. The index value corresponding to the maximum value is taken, and the index value is mapped to the label according to the mapping relationship, and then the point index corresponding to each signature is obtained according to the label.

[0017] The present invention also proposes an online handwritten signature character segmentation system, including a preprocessing module, a data set, a basic prediction model, and a character segmentation module. The preprocessing module preprocesses the collected signature stroke data to obtain a stroke feature sequence of paired data including stroke data features and labels to form a data set. The data set is divided into a training set and a verification set and input into a basic prediction model for training to obtain a prediction model. The preprocessed signature data features are input into the prediction model for prediction processing, and the output is a probability distribution vector of the same length as the stroke feature sequence, which is the label prediction probability corresponding to the position of the signature stroke feature sequence. The character segmentation module takes the index value corresponding to the maximum value of the label prediction probability, maps the index value to a label according to a mapping relationship, and then obtains the point index corresponding to each signature according to the label, obtains the point to which each signature belongs, and segments different characters from the signature data according to the point.

[0018] A further preferred solution is that the basic prediction model includes a bidirectional long short-term memory network Bi-lstm, a fully connected layer Dense and a conditional random CRF, wherein the Bi-Lstm is responsible for inputting encoding information, and the signature stroke writing feature sequence is input from the model input layer as encoding information, the Dense layer converts the encoder output dimension into a label dimension, and the CRF is responsible for outputting the label.

[0019] The present invention also proposes a computer-readable storage medium on which a computer program is stored. The program is executed by a processor to implement the online handwritten signature character segmentation method described in any one of claims 1 to 6.

[0020] The present invention also proposes an electronic device, comprising: one or more processors; a memory; and one or more applications, which are stored in the memory and configured to be loaded and run by the one or more processors to execute the online handwritten signature character segmentation method described in any one of claims 1 to 6.

[0021] The present invention uses deep learning technology to form an end-to-end solution for online handwritten character segmentation. By inputting a sequence of acquired sampling point data (such as coordinates, pressure values, time and other information), the character segmentation results can be directly obtained, thereby providing basic support for character library construction and subsequent character content recognition. Compared with the segmentation method based only on image input, the stroke data features, labels and other features formed by information such as time, pressure and writing order are considered, and the probability vector is obtained through a prediction model to determine the mapped labels, which further improves the segmentation accuracy and the signature recognition accuracy. It can be applied to application scenarios such as handwriting forensic identification that have high requirements for recognition accuracy, and reduces the acquisition cost and other problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1Schematic diagram of the signature character segmentation process of the present invention;

[0023] Figure 2 Schematic diagram of the basic prediction model structure;

[0024] Figure 3 Schematic diagram of the prediction model training process. DETAILED DESCRIPTION

[0025] The implementation of the present invention is described in detail below with reference to the accompanying drawings and specific examples.

[0026] like Figure 1 The figure shows the flow chart of signature prediction of the present invention. The prepared data set is input into the model for training, the trained model is used as the prediction model, the signature stroke features and label waiting prediction data are preprocessed, and the prediction result is input into the prediction model for prediction output.

[0027] Data preprocessing specifically includes preprocessing handwritten signature handwriting feature data, including normalization, downsampling, padding, label data conversion, data collection, data splitting, etc., to obtain a stroke feature sequence including stroke data features and label paired data, that is, a set of input and output pairs of different samples. The input is the x, y, p, t and other features of the electronic signature, and the output is a label vector corresponding to each position of the sequence. When stored, it can be a serialized Pickle file, etc. The preprocessed data is encoded, and the encoding output is inversely operated to obtain the predicted label, and further obtain the signature point. The preprocessing output constitutes a data set, and the data set is divided into a training set and a validation set and input into the basic prediction model for training.

[0028] Data Normalization

[0029] 1) Normalize the collected signature stroke data, including normalizing the stroke coordinates, pressure, and time. The following example illustrates a method for normalizing the coordinate vector (x, y), namely, coordinate normalization, according to the formula: range max =2·max((X mean -X min ), (Y mean -Y min ))Calculate the scaling factor range max , according to the formula: Normalize the horizontal axis (X) and the vertical axis (Y), where X min With Y min Respectively represent the minimum value of the horizontal and vertical coordinate sequences of an online handwriting data, X mean and Y meanThey represent the average values ​​of the corresponding sequences (the average value can be obtained by summing up all the horizontal coordinate values ​​and then dividing by the sequence length), X norm and Y norm Respectively represent the normalized horizontal and vertical coordinate sequences.

[0030] Normalization of pressure value (P) can be done by min-max normalization, and special values ​​(such as the pressure value representing the pen lifting action) are not included in the normalization. Special values ​​can retain their original values ​​or be modified to other fixed values ​​that are not within the normal value representation range.

[0031] The time (T) is normalized and converted into a duration representation of a fixed unit (it can be converted into milliseconds or other units of time through difference calculation or other transformation).

[0032] Other methods of the present invention may also be used to normalize the electronic signature handwriting feature sequence.

[0033] In order to improve the speed of model application and further improve the accuracy, it is necessary to downsample the input data. Downsampling can be done by the following methods (other conventional sampling methods in this field can also be used):

[0034] Stroke extraction. The start and end of each stroke can be determined through the pen lift point information of the pressure value. Specifically, a "pen down-move-pen lift" process is a stroke. Subsequent downsampling operations are all performed within the same stroke, and the same process is performed for each stroke. The so-called "stroke" refers to a set of data points formed by a user performing a "pen down-move-pen lift" process on an online writing device.

[0035] Remove adjacent points. There are many redundant points in the input data, one of which is that the distances between them are relatively close. Removing these points will basically not affect the change of the glyph. In the same stroke, those that meet the following conditions can be removed (the first and last points are retained and do not participate in the downsampling and point removal process). i is the horizontal coordinate of the i-th point in the stroke, y i is the ordinate of the i-th point, X is the vector formed by the abscissas of all points in the input data, Y is the vector formed by the ordinates of all points, α is the proportional coefficient, and the recommended value range is 0.024±0.01, T dist is the distance threshold, which is calculated as follows:

[0036]

[0037] T dist =α·min(X mean , Y mean )

[0038] Remove straight line points. For a straight stroke or a straight part of a stroke, retaining its two end points can effectively retain the information of the straight line segment. Specifically, all points in the stroke that satisfy the following formula can be discarded. In the formula, Δx i =x i+1 -x i ,

[0039] Δy i =y i+1 -y i , T cos is the cosine threshold.

[0040] Homogenize line density. For character segmentation tasks, the integrity of some glyphs can be sacrificed to improve the overall uniformity of the writing sampling points, thereby improving the segmentation accuracy. Homogenize line density by calculating line density through a sliding window. Points in the window that meet the calculation formula are discarded. The window size s w The value is generally between 3 and 15. The first and last points are not involved in this process. For points near the stroke boundary, we can fall back to the window calculation. For example, the line density of the second point of the stroke can be calculated by setting the window size to 3. In the formula, Δx i =x i -x i-1 ,Δy i =y i -y i-1 , the corner mark i refers to the i-th point in the window, T density is the line density threshold,

[0041]

[0042]

[0043] Among them, β is the density proportionality coefficient, which is 0.2±0.1, and X and Y are vectors composed of all horizontal and vertical coordinates respectively.

[0044] After the above-mentioned downsampling, the length of the stroke data sequence obtained will still fluctuate within a small range. At this time, a target length of the padding sequence that is greater than the predetermined length of the target sequence (such as 50 or other values ​​suitable for the training model) can be specified, and its input can be padded with 0 values, and the label value can be filled with a label representing the padding. If the original sequence length becomes 150 after downsampling, and the predetermined length is set to 200 (training requires that the length of all sample sequences be kept equal), then all the handwriting features (such as x, y, p, t) from 151 to 200 are set to 0, and this operation is called padding. Because this task has a corresponding label for each position, and the labels corresponding to these padding positions are O (it can also be any other reasonable representation, which can be any value different from the starting and middle points).

[0045] Label data conversion

[0046] 1) Convert the tag value to BIO format. The so-called BIO format is to use B to represent the first point of the sequence point of a word, and I to represent the subsequent sequence points of the components. The first sequence point of the next word is also represented by B, and the subsequent sequence points are also represented by I. Repeat this process until the valid sequence points of the signature strokes end. If there is a filling operation in the stroke data sequence, the tag at the corresponding position is set to O. (The label format is not limited to this format, and can be other label formats that can reasonably represent the characteristic points of the signature)

[0047] 2) Map the above BIO tags to digital representations, such as 0, 1, 2, etc.

[0048] Data collection, obtains paired data including stroke data features and labels.

[0049] Through the above-mentioned preprocessing operations such as data normalization, downsampling, padding, label data conversion, etc., paired data including the (input stroke data features, labels) structure can be obtained. The collected signature stroke data features include X, Y, P, T, etc., and the labels are in the form of BIIIIII...BIIIIII...OO... etc.

[0050] Data splitting, split the data set (paired data) into training set and validation set according to a certain ratio (split only at the sample level). In general, the validation set ratio is lower than the training set ratio. An example is as follows: For example, the data set: {(I1,,O1),(I2,,O2),(I3,O3),(I4,,O4),……(In,,On)}; after splitting, the training set: {(I1,,O1),(I5,O5),(I11,O11)……}, the validation set: {(I2,,O2),(I3,,O3),(I4,O4)……}. Among them, I represents input, O represents output, and the subscript represents the sample number.

[0051] Build and train a basic prediction model including an encoder

[0052] like Figure 2 The basic prediction model structure of the present invention is shown, including a bidirectional long short-term memory network (Bi-lstm), a fully connected layer (Dense) and a conditional random field layer (CRF). Among them, Bi-Lstm is responsible for inputting coding information, and inputting the signature stroke writing feature sequence from the model input layer as coding information (a series of data sequences containing signature stroke coordinates, time, and pressure features); the Dense layer converts the encoder output dimension into a label dimension; and the CRF is responsible for obtaining the label output.

[0053] The following example uses a single sample input as an example. Multiple inputs parallelize the process.

[0054] A matrix in the form of (input sequence length) X (number of features) (X represents multiplication). Bi-LSTM is a bidirectional (Bi) long short-term memory network (LSTM), which can be regarded as the concatenation of two LSTMs, one iterating the input from front to back, and the other iterating the input from back to front. The two directions can make up for a certain long-distance forgetting problem. For the input at each position, a corresponding output will be generated (a vector, the dimension is related to the number of units set by LSTM, for example, the number of units is set to 128, and the output vector dimension of Bi-LSTM here is 256.), and the output of this layer is a matrix of sequence length X256 (taking the number of units as 128 as an example). Dense: This layer is responsible for reducing the 256-dimensional vector at each position to the same number of output label categories (such as 3 dimensions corresponding to three BIO labels), and the output of this layer is a matrix of sequence length X3. CRF: Conditional Random Field layer, which can consider both the input and the current output to give the final output (more suitable for situations where different positions of the output labels have dependencies). This layer is responsible for calculating the loss function and decoding output during training (generally using the Viterbi algorithm to find the optimal path). The decoding output of this layer is a vector of sequence length dimension, and each value represents the label index of the corresponding position (this index is converted to a BIO label through a fixed mapping relationship).

[0055] The encoder of the basic pre-trained model of this embodiment can be composed of Bi-LSTM, but in principle any model that can perform sequence encoding can be used as an encoder, such as a convolutional neural network or a Transformer and their combination forms can be used as the encoder of the present invention.

[0056] The output of the basic prediction model is a probability distribution vector of the same length as the input feature sequence (the dimension is determined by the type of label. If the BIO format label is used, the dimension is 3), which represents the label prediction probability corresponding to each sequence position. For example, for the nth position, the label here may be a type of BIO. For example: the probability vector output at the nth position is [0.3, 0.98, 0.1]. If this vector corresponds to the label BIO, then this vector indicates that the model predicts that the probability of this position being the B label is the highest. Take the index value corresponding to the maximum value, map the index value to the label based on the mapping relationship, and then obtain the point index corresponding to each signature according to the label. Training hyperparameter configuration. The training model can be regarded as a relatively complex mapping function. The goal of training is to iteratively optimize (adjust) the parameters so that the gap between the model prediction value and the true value gradually decreases, so that the parameters in the model converge to a suitable degree, such as the error between the prediction value and the true value is within an acceptable range. The model obtained at this time can be used to predict on unknown data and obtain the prediction result. Generally, forward propagation is used, the loss function is used to calculate the loss value, and back propagation is used to adjust the model parameters according to the gradient information of back propagation. The early stopping strategy is generally used to terminate the training.

[0057] like Figure 3 The figure shows a schematic diagram of the model training process.

[0058] The loss function is recommended to use negative log-likelihood, but other reasonable loss functions are not limited to use, which needs to be determined according to the specific data. The loss function is used to calculate the difference between the predicted value and the true value, which is used for subsequent back propagation and gradient calculation.

[0059] Furthermore, in order to optimize the segmentation accuracy, a character width ratio regularization term can be added. This loss term can be added to the standard crf loss function. The calculation method is as follows: max is the width of the widest segmentation word in the current prediction result, w min is the width of the narrowest segmented word, γ is the proportionality coefficient, which is generally set between 0.001 and 0.05, and ∈ is the compensation coefficient, which is set to 1. -5 ~1 -3 .

[0060]

[0061] The optimizer is a parameter adjustment algorithm that allows the model to converge after continuous iteration of the parameter adjustment values ​​given by the algorithm. One of the simplest optimizers is stochastic gradient descent (SGD). The optimizer recommended by the present invention is rmsprop, but it can also be other optimizers, such as adam, adgrad, etc.

[0062] 1) The training batch size batch_size can be determined based on the actual training machine resources.

[0063] 2) Early stopping strategy: If there is no decrease in the validation set after n rounds, stop training and restore the parameters to the optimal value. The value of n is set according to the flatness of the hyperparameter space, generally ranging from 3 to 20.

[0064] The signature data features after preprocessing are input into the prediction model to obtain the model output. The model output is inversely operated to obtain its predicted label, and further obtain the point to which the word belongs. The inverse operation can be specifically to obtain the label of each point through model prediction, and the labels of each position are combined to form a result sequence, such as: BIIIIIIIIBIIIIIIIIIBIIIIIIIIIOOOOOO, at this time, the first B (inclusive) and the point before the nearest B or O all belong to the first word, and the same is true for the second word. In this way, the point to which each word belongs can be obtained, which completes the purpose of segmenting different characters from a sequence.

[0065] The prediction result can be directly obtained by inputting the original label sequence data, such as the sequence data composed of features such as X, y, p, t, etc. For a sample, it is a matrix of sequence length X feature number. It can also process sequence inputs with edge length (Bi-lstm model), and it only takes about 100ms to get the result for a single input.

Claims

1. An online handwritten signature character segmentation method, characterized in that: The collected signature stroke data is preprocessed to obtain a stroke feature sequence including stroke data features and label paired data to form a data set, and the data set is input into a basic prediction model for training to obtain a prediction model; the preprocessed signature data features are input into the prediction model, and the output of the prediction model is a probability distribution vector of the same length as the stroke feature sequence, which is the label prediction probability corresponding to the position of the signature stroke feature sequence, and the index value corresponding to the maximum value is taken, and the index value is mapped to a label according to the mapping relationship, and then the point index corresponding to each signature is obtained according to the label, and the point to which each signature belongs is obtained, and different characters are segmented according to the point; The preprocessing includes: normalizing the collected signature stroke data to obtain a horizontal coordinate X norm and the ordinate Y norm Sequence, downsample the signature stroke data, fill the signature stroke data and convert the label data.

2. The method according to claim 1, characterized in that Normalize the signature stroke data according to the formula: range max =2·max((X mean -X min ), (Y mean -Y min ))Calculate the scaling factor range max , according to the formula: Normalize the horizontal coordinate X and vertical coordinate Y of the stroke data, where X min With Y min Respectively represent the minimum value of the horizontal coordinate sequence and the vertical coordinate sequence of a handwriting data, X mean and Y mean Respectively represent the average values ​​of the corresponding sequences of the horizontal and vertical axes, X norm and Y norm Respectively represent the normalized horizontal and vertical coordinate sequences.

3. The method according to claim 1, characterized in that The downsampling of the signature stroke data is to determine the stroke breakpoints through the pen lifting action represented by the writing pressure value change information, retain the pressure value of each breakpoint and the two points before and after it, and obtain the downsampled stroke pressure value sequence based on the interval sampling of the remaining unselected stroke data points.

4. The method according to claim 3, characterized in that The downsampling includes: extracting strokes, removing adjacent points, removing straight line points, and uniformizing line density, specifically including: extracting strokes by determining the start and end of each stroke through the pressure value of the pen lift point; removing strokes that meet the formula within the same stroke: The i-th pixel (x i ,y i ), x i is the horizontal coordinate of the i-th pixel in the stroke, y i is the ordinate of the i-th pixel, α is the proportional coefficient, which is 0.024±0.01, T dist is the distance threshold; for a straight stroke or a straight part of a stroke, retain its two end points; according to the formula: Determine the line density threshold, set a sliding window, and discard the points in the window whose line density meets the line density threshold. β is the density ratio coefficient, and its value is 0.2±0.

1.

5. The method according to claim 3, characterized in that: The filling and label data conversion of the signature stroke data is to fill the stroke pressure value sequence after downsampling, fill the part greater than the predetermined length of the target sequence with 0 value, and convert the label value into BIO format for filling. The BIO format is: the first point at the beginning of the component sequence points of a stroke feature of a certain character is represented by B, and the subsequent component sequence points are represented by I. The next character is also represented by B at the beginning of the first sequence point, and the subsequent component sequence points are also represented by I. This is repeated until the valid sequence points of the signature stroke end. If there is a filling operation in the stroke data sequence, the label at the corresponding position is set to O.

6. The method according to any one of claims 1 to 5, characterized in that: The basic prediction model includes a bidirectional long short-term memory network Bi-lstm, a fully connected layer Dense and a conditional random field layer CRF. The Bi-Lstm inputs the encoding information, and the signature stroke writing feature sequence is input from the model input layer as the encoding information. The Dense layer converts the encoder output dimension into the label dimension. The CRF is responsible for calculating the loss function loss and the decoding output. The loss function includes a regularization term: The decoding output is a vector of sequence length dimension, each value represents the label index of the corresponding position, where w max is the width of the widest segmentation word in the current prediction result, w min is the width of the narrowest segmented word, and γ is the proportional coefficient, which is set to 0.001 to 0.

05.

7. An online handwritten signature character segmentation system, characterized in that: It includes a preprocessing module, a data set, a basic prediction model, and a character segmentation module. The preprocessing module preprocesses the collected signature stroke data to obtain a stroke feature sequence including stroke data features and paired data of labels to form a data set. The data set is input into the basic prediction model for training to obtain a prediction model. The preprocessed signature data features are input into the prediction model for prediction processing, and the output is a probability distribution vector of the same length as the stroke feature sequence, which is the label prediction probability corresponding to the position of the signature stroke feature sequence. The character segmentation module takes the index value corresponding to the maximum value of the label prediction probability, maps the index value to the label according to the mapping relationship, and then obtains the point index corresponding to each signature according to the label, obtains the point to which each signature belongs, and separates different characters from the signature data according to the point. The preprocessing includes: normalizing the collected signature stroke data to obtain a horizontal coordinate X norm and the ordinate Y norm Sequence, downsample the signature stroke data, fill the signature stroke data and convert the label data.

8. The system according to claim 7, characterized in that The basic prediction model includes a bidirectional long short-term memory network Bi-lstm, a fully connected layer Dense and a conditional random layer CRF. Among them, Bi-Lstm is responsible for inputting encoding information, and the signature stroke writing feature sequence is input from the model input layer as encoding information. The Dense layer converts the encoder output dimension into the label dimension. CRF is responsible for calculating the loss function loss and decoding output. The decoding output is a vector of sequence length dimension, and each value represents the label index of the corresponding position.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the program is executed by a processor to implement the online handwritten signature character segmentation method described in any one of claims 1 to 6.

10. An electronic device, characterized in that: include: one or more processors; Memory; One or more application programs are stored in the memory and are configured to be loaded and run by the one or more processors so as to execute the online handwritten signature character segmentation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • A method and apparatus for online handwritten formula recognition based on user feedback information

    CN112926567B

  • Processing method, device and system of handwritten image of smart pen, and storage medium

    CN113158961A

  • Handwritten formula recognition method based on end-to-end network model

    CN111738169A

  • Electronic handwritten signature recognition method based on OCSVM

    CN111950331A