An online character segmentation and recognition system, method, device and storage medium

Through the dual-channel stroke feature extraction and progressive multi-task learning module of PMLNet network, the problem of overlap and interleaving of characters in electronic signatures is solved, efficient character slicing and recognition effects are achieved, and the model's slicing accuracy and speed are improved.

CN115273106BActive Publication Date: 2025-07-29CHONGQING AOXIONG INFORMATION TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202211049688.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2025-07-29
Estimated Expiration
2042-08-30

AI Technical Summary

Technical Problem

The prior art is difficult to solve the problem of text overlap and interweaving in electronic signature recognition, the over-segmentation method is difficult to recognize, and the segmentation method cannot provide segmentation results without segmentation method, which cannot meet the character word cutting requirements in actual scenarios.

Method used

The online character word-cut recognition method based on the progressive multi-task learning network PMLNet is adopted. Through the dual-channel stroke feature extraction module, the superimposed encoding module and the progressive multi-task learning module, the spatial and semantic features of the strokes are extracted respectively, the long-distance dependence between different strokes is modeled, and the classification features and word-cut features are explicitly interacted, and the accuracy of word-cutting and recognition tasks is optimized.

Benefits of technology

It improves the word slicing and recognition accuracy of electronic signatures, reduces the complexity and calculation time of model, and improves the accuracy and speed of character slicing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273106B_ABST
    Figure CN115273106B_ABST
Patent Text Reader

Abstract

The present invention discloses an online character segmentation and recognition method, which relates to the recognition of electronic signature characters. Aiming at the problem that the existing over-segmentation-based method has poor character segmentation effect on online characters, a Progressive Multi-Task Learning Network (PMLNet) is proposed for online character segmentation and recognition. Two channels are used to extract the spatial features and semantic features of strokes respectively, model the long-distance dependence relationship between different strokes, enhance the feature quality, adopt a progressive learning strategy, explicitly interact the classification features and segmentation features, and improve the accuracy of both the segmentation and recognition tasks by making full use of the relationship between the segmentation and recognition tasks. It is widely used in high-quality electronic signature recognition and authentication scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer information processing technology, and specifically to an electronic character recognition technology. Background Art

[0002] Electronic Chinese is widely used in bank transactions and judicial procedures. Because it contains valuable personal information, Chinese is regarded as the highest standard for personal identification. Different from handwritten Chinese texts, Chinese has diverse styles, lacks semantic information, has a complex font structure, and there are overlaps between characters, resulting in poor Chinese character segmentation accuracy. In recent years, the increasing growth of electronic devices has witnessed the rise of on-line Chinese. Compared with off-line Chinese, on-line Chinese has a unique time attribute, so it can provide richer information for the segmentation task. On-line Chinese character segmentation can support downstream tasks such as the establishment of Chinese character libraries and the verification of single-character Chinese, which is of great significance. However, unfortunately, few people have studied it.

[0003] The method based on over-segmentation includes three parts: character over-segmentation, character classification, and language context modeling. However, the method based on over-segmentation is difficult to solve the problems of Chinese character overlap, interweaving, and illegibility. At the same time, the method based on over-segmentation has a complex structure and the efficiency of the model is limited. In contrast, in recent years, the method based on non-segmentation has gradually become the mainstream of handwritten text recognition. However, the method based on non-segmentation cannot provide segmentation results and cannot meet the requirements of Chinese character segmentation in actual scenarios. Therefore, the method based on non-segmentation is not applicable to the Chinese character segmentation and recognition tasks.

[0004] Publication number: CN113627260A, title: "Method, system and computing device for identifying the stroke order of handwritten Chinese characters". A method for identifying the stroke order of handwritten Chinese characters is disclosed, including: obtaining a sequence of trajectory points of the handwritten Chinese character to be identified, and preprocessing the sequence of trajectory points of the handwritten Chinese character to obtain a sequence of key points of the handwritten Chinese character; extracting stroke features of the handwritten Chinese character based on the sequence of key points of the handwritten Chinese character, where the stroke features include basic features and semantic features; and using a trained stroke order recognition model to process the stroke features of the handwritten Chinese character and the stroke features of the corresponding template Chinese character together to obtain a stroke order recognition result. Obtaining a sequence of trajectory points of the handwritten Chinese character to be identified, and preprocessing the sequence of trajectory points of the handwritten Chinese character to obtain the sequence of key points of the handwritten Chinese character; extracting the stroke features of the handwritten Chinese character based on the sequence of key points of the handwritten Chinese character, where the stroke features include basic features and semantic features; and using a trained stroke order recognition model to process the stroke features of the handwritten Chinese character and the stroke features of the corresponding template Chinese character together to obtain a stroke order recognition result. Publication number: CN111931710A, title: "An on-line handwritten character recognition method, device, electronic device and storage medium". A character recognition system and method based on the combination of a neural network and an attention mechanism. A convolutional neural network feature extraction module is used for the spatial features of a character image; inputting the spatial features extracted by the convolutional neural network into a bidirectional long short-term memory network module, and the bidirectional long short-term memory network can extract the sequence features of the character; performing semantic encoding on the extracted feature vectors, and then allocating attention weights to the feature vectors through the attention mechanism to make the attention focus on the feature vectors with higher weights; the decoding part of the model is implemented by a nested long short-term memory network. Using the features extracted by the attention and the prediction information of the previous moment as the input of the nested long short-term memory network. The purpose of using long short-term memory networks both before and after is to maintain the temporal characteristics of the feature vectors and make the model pay attention to the position points that change continuously over time; the present invention can more accurately detect the text area in a natural scene, and has a good detection effect on small target characters and texts with a small inclination angle.

[0005] Publication number: CN111259880A, Title: "A Text Recognition Method for Electric Power Operation Tickets Based on Convolutional Neural Network", discloses a text recognition method for electric power operation tickets based on convolutional neural network, which relates to a text recognition method. At present, the text recognition of electric power operation tickets is not clear. A convolutional neural network model with only 3 convolutional layers, no pooling layer, and no fully connected layer is constructed, and a non-linear mapping function is trained to improve the peak signal-to-noise ratio of the image; a handwriting feature calculation method is used to obtain the imaginary stroke feature, path Chinese feature, and 8-direction feature of the text in the electric power operation ticket image respectively; an integrated convolutional neural network model with 6 convolutional layers, 5 pooling layers, and 1 fully connected layer is constructed, and combined with the imaginary stroke feature, path Chinese feature, and 8-direction feature, a text recognition model is trained; this technical solution utilizes the superior spatial feature learning ability of convolutional neural network, and uses an image enhancement method and a text recognition method based on convolutional neural network to improve the accuracy of text recognition of electric power operation ticket images. Publication number: "CN109389091B", Title: "Text Recognition System and Method Combining Neural Network and Attention Mechanism", an on-line handwritten text recognition method, the on-line handwritten text recognition method includes: obtaining a handwritten trajectory and determining the stroke coordinate sequence of each handwritten stroke in the handwritten trajectory; performing single-character segmentation on the handwritten trajectory according to the stroke coordinate sequence of the handwritten stroke to obtain a single-character handwritten trajectory; determining the feature vector of the single-character handwritten trajectory according to the feature points of each handwritten stroke in the single-character handwritten trajectory; inputting the feature vector of the single-character handwritten trajectory into a machine learning classification model to obtain the text content corresponding to the handwritten trajectory. The above patent documents still have difficulty in solving the problems of overlapping, interweaving, and illegibility of Chinese characters during single-character segmentation based on the trajectory, and do not effectively utilize recognition information to improve the accuracy and effect of intimate recognition. Summary of the Invention

[0006] Aiming at the problems existing in the prior art in electronic signature recognition, such as the over-segmentation method being difficult to solve the problems of overlapping, interweaving, and illegibility of characters, and the non-segmentation method being unable to provide a segmentation result and unable to meet the demand for character segmentation in actual scenarios, the present invention proposes an on-line character segmentation and recognition method based on the Progressive Multi-Task Learning Network (PMLNet), which predicts the segmentation and recognition results of characters in an end-to-end manner.

[0007] The technical solution of the present invention to solve the above technical problems is to propose an online character segmentation and recognition system based on PMLNet, including: a dual-channel stroke feature extraction module, a stacked encoding module, and a progressive multi-task learning module. The dual-channel stroke feature extraction module is used to encode the stroke position information and trajectory information of the electronic signature into a vector of a fixed dimension, and multiple strokes form a stroke sequence feature. The stacked encoding module is used to model the long-distance dependence relationship between different strokes and encode the stroke sequence feature into a character sequence feature. The progressive multi-task learning module includes parallel segmentation branches and recognition branches. The two branches independently predict to obtain the intermediate results of character segmentation classification and recognition classification, and then send the two intermediate classification results into the other branch for interaction to optimize the character segmentation result and recognition result.

[0008] Further preferably, the dual-channel stroke feature extraction module includes two channels. Channel 1 processes the signature stroke trajectory information, and Channel 2 processes the stroke position information. The obtained information features are normalized, the pen tip movement trajectory is divided into strokes, the feature vector represented by the stroke trajectory contour is obtained, the stroke position feature is obtained, and the spatial feature and semantic feature of the signature are obtained based on the feature vector represented by the stroke trajectory contour and the position feature.

[0009] Further preferably, the stacked encoding module is composed of multiple Transformer encoders stacked together. Each Transformer encoder is composed of multiple multi-head self-attention layers and fully connected layers cascaded at intervals. Residual connections and normalization are used between two Transformer encoders, and position encoding is added to the original input to utilize the order of the sequence signal. The progressive multi-task learning module uses the PMLNet network to perform progressive and interactive electronic signature segmentation and recognition. The PMLNet network includes: an OCR classifier C ocr 、a segmentation classifier C seg 、a statistical graph module. The OCR classifier includes a fully connected layer and a softmax layer, which are used to predict the character category to which each stroke belongs. The segmentation classification includes a fully connected layer and a conditional random field layer, which are used to predict each stroke category, including character start, character middle, and others. The statistical graph module simulates two-category Gaussian distributions, and balances the segmentation and classification intermediate results according to the loss value of the PMLNet network and the KL divergence to obtain high-quality segmentation results and classification results.

[0010] Further preferably, the dual-channel stroke feature extraction module includes two different channels. Channel 1 uses all the feature vectors of the strokes as inputs, uses a one-dimensional convolutional layer to extract the spatial features of the strokes, and uses a max pooling layer and a fully connected layer to merge the point-by-point features into stroke-by-stroke features to obtain the stroke trajectory features. Channel 2 uses the position features in the stroke feature vector As input, three cascaded fully connected layers are used to extract the semantic features of the strokes according to the stroke position features, and the spatial features and semantic features are concatenated.

[0011] Further preferably, the two channels in the dual-channel stroke feature extraction module are the sub-network 1 (channel 1) and sub-network 2 (channel 2) of two parallel branches. The sub-network 1 is composed of a cascaded convolutional layer, max pooling layer, and fully connected layer to extract the spatial features of the electronic signature. The sub-network 2 is composed of three cascaded fully connected layers to extract the semantic features of the electronic signature.

[0012] Further preferably, the further division of the pen tip movement writing trajectory into strokes includes that the preprocessing part calculates the normalized coordinate values of the Chinese trajectory point coordinates (X, Y). Call the formula: Calculate the stroke turning angle θ. According to the pen-up / pen-down information of the stroke feature points, the turning angle change determines the starting point of the stroke. When the pen-up / pen-down information of a certain point of the stroke is 0 and the turning angle change is greater than a predetermined angle, this point is the starting point of the stroke. Determine the number of strokes according to the starting point of the stroke, and divide the electronic signature into strokes, where, represents the derivative with respect to time, represents the derivative with respect to time.

[0013] Further preferably, the dual-channel stroke feature extraction module extracts the feature vector sequence of the direction vector, state vector, and position vector contained in the electronic signature stroke, concatenates the above feature vector sequences to obtain the linear difference of the stroke segment to maintain the electronic signature stroke trajectory contour, and represents the electronic signature feature vector according to the set maximum number of strokes and linear difference.

[0014] Further preferably, according to the formula: Divide the 8-direction set into two subsets and For any point t ij in the stroke trajectory, the corresponding turning angle θ ij is respectively decomposed into two adjacent directions d ij and d′ ij , according to the formula:

[0015] a ij =|cos(θ ij ) - sin(θ ij )|

[0016]

[0017] Calculate the two adjacent directions d ij and d′ij The corresponding value a ij , a' ij , according to the formula:

[0018]

[0019]

[0020] v s (t ij ) = [1(s ij = 0), 1(s ij = 1), 1(s ij = 2)] T

[0021]

[0022] Calculate the two directional feature vectors v ij of any stroke point t d (t ij ), v d′ (t ij ), the state vector v s (t ij ), and the position vector v c (t ij ), where s ij is the start and end pen information of the stroke point t ij , is the normalized coordinate value, 1(·) represents the indicator function, e d , e d′ is the standard basis vector.

[0023] Further preferably, the position features of the start point and end point of the electronic signature stroke are input to the second fully connected layer of channel 2, and the position features of the middle point of the electronic signature stroke are input to the first fully connected layer of channel 2.

[0024] Further preferably, the negative log-likelihood function of the sequence label Y seg predicted based on the conditional random field is used to determine the minimization of the character cutting prediction loss The cross-entropy function of the classification label Y ocr and the classification prediction result is used to determine the minimization of the classification prediction loss The PMLNet network loss value is the weighted sum of the character cutting prediction loss and the classification prediction loss. Specifically, according to the formula Calculate the PMLNet network loss value L total , where, represents the set of all losses, C represents the set of weights corresponding to the losses, L'seg and L′ ocr indicating minimizing the loss of intermediate value prediction for character segmentation and classification, KL seg and KL ocr is to minimize the KL divergence of character segmentation and classification.

[0025] In a second aspect, the present invention also proposes an online character segmentation and recognition method based on PMLNet, including: extracting stroke sampling points to obtain the pen tip movement trajectory and position features, performing normalization processing, dividing the movement writing trajectory into signature strokes, encoding the electronic signature stroke position information and trajectory information into a vector with a fixed dimension, multiple strokes forming a stroke sequence feature, and obtaining the spatial feature and semantic feature of the signature according to the stroke trajectory information and position information; modeling the long-distance dependence relationship between different strokes, and encoding the stroke sequence feature into a character sequence feature; performing single-character segmentation and recognition on the electronic signature by parallel character segmentation branches and recognition branches according to the electronic signature spatial feature, semantic feature, and sequence feature, and interacting the results of the character segmentation branch and the recognition branch to optimize and obtain high-quality electronic signature character segmentation results and recognition results.

[0026] Further preferably, the dividing the movement writing trajectory into signature strokes further includes calculating the normalized coordinate values of the electronic signature trajectory point coordinates (X, Y) calling the formula: calculating the Chinese stroke turning angle θ, determining the starting point of the stroke according to the pen-up / pen-down information of the stroke feature points and the turning angle change. When the pen-up / pen-down information of a certain point of the stroke is 0 and the turning angle change is greater than a predetermined angle, this point is the starting point of the stroke, determining the number of strokes according to the starting point of the stroke, and dividing the signature into strokes, where represents the derivative with respect to time, represents the derivative with respect to time.

[0027] Further preferably, the obtaining the character feature vector further includes: extracting the feature vector sequence of the direction vector, state vector, and position vector included in the character stroke, concatenating the above feature vector sequences to obtain the linear difference of the stroke segment to maintain the electronic signature stroke trajectory contour, and representing the character feature vector according to the set maximum number of strokes and the linear difference.

[0028] Further preferably, the obtaining the spatial feature and semantic feature of the signature according to the stroke trajectory information and position information further includes: Channel 1 uses all the feature vectors of the character stroke as input, and uses a one-dimensional convolutional layer to extract the spatial feature of the stroke; uses a max-pooling layer and a fully connected layer to merge the point-by-point features into stroke-by-stroke features to obtain the stroke trajectory feature; Channel 2 uses the position feature in the signature stroke feature vector As the input, three cascaded fully-connected layers are used to extract the semantic features of the strokes based on the stroke position features, and the spatial features and semantic features are cascaded.

[0029] Further preferably, the position features of the starting point and the ending point of the signature stroke are input to the second fully-connected layer of channel 2, and the position features of the middle point of the signature stroke are input to the first fully-connected layer of channel 2.

[0030] Further preferably, according to the formula: The 8-direction set is divided into two subsets and Any point t ij in the stroke trajectory ij The corresponding turning angle θ ij is decomposed into two adjacent directions d ij and d'

[0031] a ij = |cos(θ ij ) - sin(θ ij )|

[0032]

[0033] Calculate the corresponding values a ij and a' ij of two adjacent directions d ij and d' ij , according to the formula:

[0034]

[0035]

[0036] v s (t ij ) = [1(s ij = 0), 1(s ij = 1), 1(s ij = 2)] T

[0037]

[0038] Calculate the two direction feature vectors v ij of any stroke point t d (t ij ), v d′ (t ij ), the state vector v s (t ij ) and the position vector v c (tij ), where s ij is the starting and ending stroke information of stroke point t ij , is the normalized coordinate value, 1(·) represents the indicator function, e d , e d′ is the standard basis vector.

[0039] Further preferably, the character segmentation result and recognition result prediction further include: using the PMLNet network progressively and interactively for character segmentation and recognition, for predicting the category to which each stroke belongs and each stroke category; simulating two-category Gaussian distributions through statistical graphs, and balancing the segmentation, classification intermediate results according to the PMLNet network loss value and KL divergence, to obtain high-quality segmentation results and classification results.

[0040] Further preferably, based on the negative log-likelihood function of the sequence label Y seg predicted by the conditional random field to determine the minimization of the segmentation prediction loss determined by the cross-entropy function of the classification label Y ocr and the classification prediction result to determine the minimization of the classification prediction loss The PMLNet network loss value is the weighted sum of the segmentation prediction loss and the classification prediction loss. Specifically, according to the formula calculate the PMLNet network loss value L total , where represents the set of all losses, C represents the set of weights corresponding to the losses, L′ seg and L′ ocr represent the minimization of the segmentation and classification intermediate value prediction losses, KL seg and KL ocr are the KL divergences for minimizing segmentation and classification.

[0041] In a third aspect, the present invention also proposes an electronic device, including: one or more processors, a memory, one or more applications, which are stored in the memory and configured to be loaded and run by the one or more processors to execute the above-mentioned online character segmentation and recognition method based on PMLNet.

[0042] In a fourth aspect, the present invention also proposes a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned online character segmentation and recognition method based on PMLNet.

[0043] Aiming at the poor character cutting effect of existing over-segmentation-based methods, which are difficult to solve problems such as character overlap, interweaving, and illegibility, the present invention proposes an online character cutting and recognition system and method based on the Progressive Multi-Task Learning Network (PMLNet). Two channels are used to extract the spatial features and semantic features of strokes respectively, model the long-distance dependence relationship between different strokes, enhance the feature quality, and adopt a progressive learning strategy to explicitly interact the classification features and cutting features, improving the accuracy of both the cutting and recognition tasks by making full use of the relationship between the cutting and recognition tasks. Since characters may generate many sequence points due to slow writing speed, the method based on sequence point classification has a large number of sequence points, making it difficult for the model to process with long-term memory and consuming a lot of time. Splitting a single character into multiple simple strokes overcomes the problem of complex stroke representation and is convenient for encoding. After being segmented into stroke granularity, the quantity is significantly reduced, which is beneficial to improving the segmentation accuracy and calculation speed. Brief Description of the Drawings

[0044] Figure 1 Schematic structural diagram of a Chinese character cutting and recognition system based on the PMLNet network structure;

[0045] Figure 2 Schematic structural diagram of two continuously stacked Transformer encoders;

[0046] Figure 3 Schematic diagram of a stroke sequence signal label. Detailed Embodiments

[0047] To facilitate a clear understanding of the present invention, make the technical problems to be solved, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below in conjunction with the drawings and specific embodiments. In the following description, providing specific details such as specific configurations and components is only to help a comprehensive understanding of the embodiments of the present invention. Therefore, those skilled in the art should clearly understand that various changes and modifications can be made to the embodiments described here without departing from the scope and spirit of the present invention. Additionally, for the sake of clarity and conciseness, descriptions of known functions and structures are omitted. It should be understood that the embodiments are only for illustrating the present invention and not for limiting the protection scope of the present invention.

[0048] One implementation method includes that the initially extracted features include 8 directional features, and these features are input into the subsequent dual channels to obtain the stroke-by-stroke abstract features representing spatial information and semantic information respectively. The Transformer stacked encoding module is composed of multiple Transformers stacked together and is used to model the long-distance dependency relationships between different strokes. The progressive multi-task learning module has two parallel branches and can simultaneously predict the character segmentation result and recognition result. In the progressive multi-task learning module, a progressive learning strategy is introduced to further improve the accuracy of character classification and segmentation by making full use of the relationship between the classification task and the segmentation task. Specifically, it includes an on-line character segmentation and recognition system based on PMLNet, a dual-channel stroke feature extraction module, a Transformer stacked encoding module, and a progressive multi-task learning module. First, manual features are extracted, which may include 8 directional features, and these manual features are input into the subsequent dual channels to obtain the stroke-by-stroke abstract features representing spatial information and semantic information respectively. The Transformer stacked encoding module is composed of multiple Transformers stacked together and is used to model the long-distance dependency relationships between different strokes. The progressive multi-task learning module has two parallel branches and simultaneously predicts the character segmentation result and recognition result. In the progressive multi-task learning module, a progressive learning strategy is introduced to further improve the accuracy of character classification and segmentation by making full use of the relationship between the classification task and the segmentation task.

[0049] Figure 1 The figure shows a schematic diagram of the structure of a segmentation and recognition system based on the PMLNet network structure. It includes: a dual-channel stroke feature extraction module, a Transformer stacked encoding module, and a progressive multi-task learning module. The dual-channel stroke feature extraction module combines traditional knowledge and a deep learning model to extract the spatial features and semantic features of a signature based on the stroke features of the signature; the Transformer stacked encoding module takes all stroke features as input and extracts the sequence features of characters. The progressive multi-task learning module is used to predict the segmentation result and recognition result. PMLNet is a multi-task network that can utilize the advantages of character recognition to improve the segmentation performance. Most traditional methods can only perform a single task (i.e., the segmentation task or the recognition task).

[0050] Specifically, the dual-channel stroke feature extraction module consists of a preprocessing part, a manual feature extraction part, and a dual-channel sub-network. The preprocessing module performs manual feature extraction and normalization on the original trajectory features formed by the signature pen tip, divides the pen tip movement writing trajectory into different strokes, and the manual feature extraction module extracts the features of the stroke sampling points to assist the subsequent feature extraction at the network level. The output stroke sampling point features are respectively input into the dual-channel sub-networks (sub-network 1, sub-network 2), and the two sub-network channels respectively extract the spatial features and semantic features of the signature. Among them, sub-network 1 includes: a convolutional layer, a max pooling layer, and a fully connected layer, which extracts the spatial features of the character. Sub-network 2 includes: three cascaded fully connected layers, which extracts the semantic features of the character. The dual-channel sub-network consists of two branches, extracts the stroke features, and improves the character segmentation performance by adding an identification branch and introducing a progressive learning strategy. The Transformer stacking encoding module consists of multiple Transformers stacked together, so as to model the long-distance dependencies between different strokes, enhance the feature quality and then input it into the progressive multi-task learning module to segment and identify the characters, and predict the character segmentation result and recognition result. Specifically as follows:

[0051] The dual-channel stroke feature extraction module preprocesses the character. The following examples specifically illustrate the feature extraction and normalization of the character trajectory.

[0052] The electronic device can capture the character coordinates and pen-down / pen-up information as the character trajectory (X, Y, S) at a fixed time sampling rate. Among them, (X, Y) represents the two-dimensional coordinates of the character in the coordinate system, and S represents the pen-down / pen-up information. S = 0 represents pen-down, S = 1 represents writing, and S = 2 represents pen-up. Generally, characters (especially text) are written horizontally. Without changing the word length, we normalize X and Y based on the length of the coordinates (X, Y) under the premise of a fixed aspect ratio. For example, the following method can be used to achieve normalization. Call the formula:

[0053]

[0054] Calculate the normalized coordinate values of the character trajectory points (X, Y) Among them, min(X), min(Y) are the minimum values of the character stroke coordinates, and max(Y) is the maximum value of the character stroke coordinates.

[0055] Calculate the normalized coordinate values The derivative with respect to time Call the formula:

[0056]

[0057] Calculate the angle θ of the stroke trajectory tangent as the turning angle, where denotes The derivative with respect to time, denotes the derivative with respect to time, according to the formula:

[0058]

[0059] Calculate the gradient of the turning angle θ ∈ represents the time step.

[0060] Based on the pen-up / pen-down information of the stroke feature points, the turning angle change determines the starting point of the stroke.

[0061] The starting point of each stroke must satisfy one of the following two conditions:

[0062] (a) Point t i0 is the pen-down point, and the pen-up / pen-down information S i0 = 0 (5)

[0063] (b) Point t i0 has a sudden turning angle change

[0064] Therefore, by detecting the stroke feature points, when the pen-up / pen-down information of a certain point is 0 and the turning angle change is greater than a predetermined angle (such as ) this point is the starting point of the stroke. Determine the number of strokes based on the starting point of the stroke and divide the signature into strokes.

[0065] The original character T is divided into M strokes: T = {T i | i = 0, 1,..., M - 1} (7)

[0066] Each stroke can be divided into N points: T i = {t ij | j = 0, 1,..., N i - 1} (8)

[0067] where, T i is the i-th stroke of the character T, t ij is the j-th point of the i-th stroke T i , M is the total number of strokes, and N i is the number of points of the i-th stroke.

[0068] In this patent, strokes are segmented using starting points and corner - turning change conditions, and a single character is split into multiple simple strokes, overcoming the problem of complex stroke representation and facilitating coding. Since the writing speed of characters is slow, many sequence points may be generated. Based on the method of classifying sequence points, due to the large number of sequence points, it is difficult for the model to process with long - term memory and it takes a long time. After segmentation into stroke granularity, the quantity is significantly reduced, which is beneficial to improving the segmentation accuracy and calculation speed. The following is the process of feature extraction for each stroke:

[0069] The manual feature extraction module extracts the feature vector sequence of two - direction feature vectors, state vectors, and position vectors contained in the character strokes; the two - direction feature vectors, state feature vectors, and position feature vectors are concatenated to obtain the linear difference of the stroke segment to maintain the text stroke trajectory contour, and the character is represented according to the set maximum number of strokes and the linear difference.

[0070] The 8 evenly - divided direction sets can be divided into two subsets and

[0071]

[0072]

[0073] For any point t ij corresponding to the turning angle θ ij in the stroke trajectory, it is decomposed into two adjacent directions d ij and d′ ij on the two subsets respectively. The specific decomposition method is as follows: according to the formula:

[0074]

[0075]

[0076] d and d′ respectively represent the angular projection values of the turning angle θ ij for each element on the two subsets and . The turning angle θ ij corresponding to any point t ij is decomposed into two adjacent directions d ij and d′ ij .

[0077] According to the formula:

[0078] a ij =|cos(θ ij ) - sin(θ ij )| (13)

[0079]

[0080] Calculate the direction transformation values a ij and a' ij corresponding to two adjacent directions d ij and d'. ij ,

[0081] Map the points in the stroke trajectory to two direction feature vectors where represents a set of 4-dimensional vectors. A state vector and a position vector For any stroke point t ij , according to the formula:

[0082]

[0083]

[0084] v s (t ij ) = [1(s ij = 0), 1(s ij = 1), 1(s ij = 2)] T (17)

[0085]

[0086] Calculate the two direction feature vectors v ij of any stroke point t d (t ij ), v d′ (t ij ), the state vector v s (t ij ) and the position vector v c (t ij ). Where s ij is the start / end pen information of the stroke point t ij , is the normalized coordinate value, 1(·) represents the indicator function, e d , e d′ is the standard basis vector, and

[0087] The direction feature vectors v d and v d′ are one-hot encodings of the decomposed vectors, and the state vector v s is the one-hot encoding of the pen-up / pen-down information S; the position vector v c is the normalized (x, y) coordinate. Therefore, each stroke T iIt can be represented as a sequence of feature vectors including a direction feature vector, a state vector, and a position vector.

[0088] To ensure the stability of the subsequent network, we set the number of sampling points for each stroke to a fixed value (e.g., it can be set to 20 or other conventional numbers according to experience). For strokes with fewer sampling points than the fixed value, we fill their tails with 0s. For strokes with more sampling points than the fixed value, we linearly resample them to the fixed value N (20) points. The finally obtained sequence of feature vectors is expressed as: V d′ (T i ) ∈ R N×O , V s (T i ) ∈ R N×P , V c (T i ) ∈ R N×L . Among them, N (usually taken as 20) represents the number of sampling points for each stroke, O (usually taken as 4) represents the maximum number of characters in the signature, P (usually taken as 3) is the length of the state vector and the three states of lifting the pen, moving the pen, and putting down the pen, and L (usually taken as 2) is the length of the position vector v c of (x, y) coordinates, and the number of sampling points can be set arbitrarily according to the sampling requirements.

[0089] It should be noted that linear interpolation can still maintain the trajectory contour, which is very important for character segmentation and recognition. Although this may lose the dynamic contour that is very important for character authentication.

[0090] According to the formula:

[0091] Obtaining the linear interpolation of the stroke segment T i can be cascaded into the feature vector

[0092] For the recognition of the character sequence, in order to facilitate the model processing, the maximum number of strokes needs to be preset. For signatures, almost all characters are less than 4 characters, and the strokes of any ordinary characters do not exceed dozens of strokes. If we set the maximum number of strokes M to 256, characters with fewer strokes will be filled with zeros. Therefore, the signature T can be represented by the feature vector . It represents the signature stroke feature vector V(T) composed of 256 strokes and the feature vector of each stroke. These feature vectors extracted by the feature extraction module are used as the input of the subsequent dual-channel sub-network.

[0093] The dual-channel stroke feature extraction module contains two different channels. One channel is used to process the stroke trajectory, and the other channel is used to process the stroke position. Specifically, the previous channel (channel 1) takes all the feature vectors of the signature strokes As input, a one-dimensional convolutional layer is used to extract the spatial features of the strokes, and a max-pooling layer and a fully-connected layer are used to merge the point-by-point features into stroke-by-stroke features to obtain the stroke trajectory features. The latter channel (channel 2) only processes the position features in the signature stroke feature vector. Three cascaded fully-connected layers are used to extract the stroke position features. It should be noted that the position features of the starting and ending points of the stroke are directly input into the second fully-connected layer, and only the position features corresponding to the middle points are input into the first fully-connected layer. Due to the large receptive field of the fully-connected layer, channel 2 extracts the semantic features of the strokes through three levels of cascaded fully-connected layers. The spatial features and semantic features of the text strokes obtained by channel 1 and channel 2 are concatenated. Finally, the spatial features and semantic features of each stroke are concatenated into one feature and input into the subsequent stacked encoding module for further feature extraction.

[0094] One implementation of the stacked encoding module is Transformer. Considering the many advantages of Transformer, we adopt the setting where the internal dimension of the fully-connected layer is set to 10 and the output dimension is set to 64.

[0095] The Transformer stacked encoding module. Transformer uses the full attention mechanism to model the long-range dependencies of the sequence in parallel, which is superior to the recurrent neural network and long short-term memory network that require sequential calculations. Transformer has become the basic structure in many research works. Therefore, we use the Transformer encoder to model the long-range dependencies between different strokes and enhance the feature quality.

[0096] Figure 2 Two consecutive stacked Transformer encoders are shown. The Transformer encoder is composed of a multi-head self-attention layer and a fully-connected layer cascaded at the sub-layer interval. Residual connections and normalization are used in both sub-layers. Position encoding is added to the original input to utilize the order of the stroke sequence signal. After concatenating the spatial features and semantic features of the strokes, they are input into the Transformer encoder. The position encoding of the Transformer encoder is the sine encoding of the stroke order. Since the Transformer encoder is a bidirectional encoding structure, it can overcome the long-range position dependence problem. The Transformer encoder outputs features of the same stroke length. Based on this feature and the classification loss, the character segmentation task can be achieved, and the classification results are the probabilities of B / I / O, representing the start of the character, the middle of the character, and others, respectively. Based on this feature and another classification loss, the character recognition task can be achieved, and the classification result is the probability of the character encoding.

[0097] Progressive multi-task learning module. AsFigure 1 As shown in Figure 1 , the PMLNet network includes an OCR recognition classifier C ocr , a character segmentation classifier C seg , an SM seg layer, and an SM ocr layer. The effect of character segmentation on left-right structured text characters is not good. The importance of character recognition leads us to turn to multi-task learning of character segmentation and recognition, trying to use the character recognition task to assist the character segmentation task. We find that character segmentation and recognition are two mutually dependent tasks. Therefore, the present invention uses the PMLNet network to perform character segmentation and recognition progressively and interactively. However, the method and system proposed in the present invention are not only applicable to Chinese character recognition, but also applicable to the recognition, identification, and classification of other characters.

[0098] Specifically, the two classifiers first independently predict to obtain the intermediate results Y′ seg of signature segmentation classification and the intermediate results Y′ ocr of signature recognition classification. These two intermediate results contain important information that can assist each other. Then, the intermediate classification results of segmentation and recognition are sent to the other classifier for interaction to optimize the character segmentation result and recognition result. For example, if the character recognition result Y′ seg on a certain stroke shows that the probability of a certain character recognition is very high, this information is also helpful for strengthening the character segmentation of this stroke.

[0099] Assume that the output stroke feature of the stacked encoding is H. The mixed information flows of the intermediate result Y′ seg of character segmentation and the intermediate result Y′ ocr of character recognition are respectively adjusted by the SM seg and SM ocr layer statistical graphs, so as to form better feature representations T seg and T ocr . Combining the intermediate results obtained by the segmentation classifier and the recognition classifier, higher-quality segmentation results and recognition results are obtained. Among them, the recognition classifier C ocr includes a fully connected layer and a conventional classifier softmax layer, which are used to predict the probability of the character category to which each stroke belongs. The segmentation classifier C seg includes a fully connected layer and a conditional random field layer, which are used to predict the stroke attribute category (BIO category) to which each stroke belongs. Among them, B represents that this stroke is the starting stroke of a character, I represents that this stroke is the middle stroke of a character, and 0 represents that this stroke does not belong to any character.

[0100] Figure 3Shown is the stroke sequence signal label. As can be seen from the figure, "Deng Jie" is split into 13 strokes. "seg" represents the segmentation label, which consists of the starting letter "B" and the middle letter "I". "OCR" represents the recognition label, which consists of the character code and "0". The 7th stroke and the 13th stroke represent the "Deng" and "Jie" labels respectively. The present invention correctly distinguishes these labels through a deep neural network training model.

[0101] In the progressive multi-task learning module, SM seg layer, SM ocr layer is used to fuse the output stroke features h of the stacked encoding, the intermediate segmentation result y' seg , and the intermediate recognition result y' ocr , to obtain a better segmentation feature t seg , t OCR . SM seg layer, SM ocr layer adopts the KL divergence variational approximation method. This method balances the importance of different information by minimizing the KL divergence KL(p(t seg |h, y' seg , y' ocr )) || r(t seg |y' seg , y' ocr ), and is used to represent the sufficient statistics T seg from Y' ocr , T seg , T OCR . Among them, r(t seg |y' seg , y' ocr ) is the variational approximation of p(t seg |h, y' seg , y' ocr ).

[0102] Determination of the loss function. Both the character segmentation and recognition tasks include three types of target losses: intermediate result loss, KL divergence loss, and final result loss. Among them, minimizing the intermediate value prediction losses L' seg and L' ocr is used to improve the performance of early character segmentation and classification. According to the formula:

[0103] L' seg = NLL(Y seg |H) (20)

[0104] L' ocr = CE(Y ocr , Y' ocr ) (21)

[0105] Calculate the intermediate prediction losses L' seg and L'ocr , L′ seg is the sequence label Y predicted based on the conditional random field on the basis of the feature map H output by the original superimposed coding module seg The negative log-likelihood function NLL, L′ ocr is the classification label Y ocr and the classification prediction result Y′ ocr The cross-entropy function CE of

[0106] The KL divergence loss is based on the formula:

[0107] KL seg = KL(p(t seg |h, y′ seg , y′ ocr ))||r(t seg |y′ seg , y′ ocr )) (22)

[0108] KL ocr = KL(p(t ocr |h, y′ seg , y′ ocr )||r(t ocr |y′ seg , y′ ocr )) (23)

[0109] Obtain the KL divergence cut that minimizes the KL divergence loss KL seg and identify the KL divergence loss KL ocr , which is used to improve the representation ability of T seg and T ocr

[0110] Determine the final segmentation loss and the final recognition loss value. According to the formula:

[0111]

[0112]

[0113] Obtain the minimized character segmentation classification loss and the recognition classification loss for improving the accuracy of the final prediction result. is the negative log-likelihood function of the sequence label Y predicted based on the conditional random field on the basis of counting T seg seg is the classification label Y ocr and the classification prediction result ​​​The cross-entropy function, whereby the total loss of the PMLNet network of the present invention is the weighted sum of all losses, and its expression is as follows:

[0114]

[0115] Among them, represents the set of all losses, and C represents the set of weights corresponding to all losses, and its expression is as follows:

[0116]

[0117] It should be noted that the result of character recognition is stroke by stroke. During the training of the PMLNet network, the last stroke of each character needs to participate in the calculation of the classification loss. In the test stage, the classification result of the last stroke of each character in the segmentation result is used as the character recognition result. This ensures the consistency of the character segmentation and classification tasks.

[0118] On our character dataset, the method of the present invention has better results than the existing character segmentation and OCR methods. The character segmentation accuracy reaches 96.9%, and the errors are mainly caused by connected strokes. The character recognition accuracy reaches 87.5%.

[0119] To facilitate the understanding of the technical solution and effect of the invention, as well as the difficulty of character recognition, the embodiments of the present invention are specifically described and analyzed by taking the recognition, segmentation, and classification of Chinese characters as an example. However, the method and system proposed by the present invention are not only applicable to Chinese character recognition, but also applicable to the recognition, identification, and classification of other characters.

Claims

1. An on-line character splitting and recognition system, characterized in that, Including: A dual-channel stroke feature extraction module, a superimposed encoding module, and a progressive multi-task learning module. The dual-channel stroke feature extraction module is used to encode the electronic signature stroke position information and trajectory information into a vector with a fixed dimension, and multiple strokes form a stroke sequence feature; The superimposed coding module is composed of multiple stacked Transformer encoders, which are used to model the long-distance dependencies between different strokes and encode the stroke sequence features into sequence features; the progressive multi-task learning module includes parallel character cutting branches and recognition branches. The two branches independently predict to obtain the intermediate results of character cutting classification and recognition classification, and then send the two intermediate classification results into the other branch for interaction to optimize and obtain high-quality character cutting results and recognition results. Specifically, the PMLNet network is used for progressive and interactive electronic signature character cutting and recognition. The PMLNet network includes: a recognition classifier , a character cutting classifier , and a statistical graph module. The recognition classifier contains a fully connected layer and a softmax layer, which are used to predict the character category to which each stroke belongs; the character cutting classifier contains a fully connected layer and a conditional random field layer, which are used to predict each stroke category; the statistical graph module simulates two-category Gaussian distributions, and balances the character cutting and classification intermediate results according to the loss value and KL divergence of the PMLNet network to obtain high-quality character cutting results and classification results.

2. The system according to claim 1, wherein The dual-channel stroke feature extraction module contains two channels. Channel 1 processes the signature stroke trajectory information, and Channel 2 processes the stroke position information. The obtained information features are normalized. The pen tip movement trajectory is divided into strokes, and the feature vector represented by the stroke trajectory contour is obtained. The stroke position feature is obtained. Based on the feature vector represented by the stroke trajectory contour and the position feature, the spatial feature and semantic feature of the signature are obtained, and a vector with a fixed dimension is obtained.

3. The system according to claim 1, wherein Each Transformer encoder in the superimposed encoding module is composed of multiple multi-head self-attention layers and fully connected layers cascaded at intervals. Residual connections and normalization are used between the two Transformer encoders. Position encoding is added to the original input to utilize the order of the sequence signals.

4. The system according to any one of claims 1 to 3, characterized in that The two channels in the dual-channel stroke feature extraction module are sub-network 1 and sub-network 2 with two parallel branches. Sub-network 1 is composed of a convolutional layer, a max pooling layer, and a fully connected layer cascaded to extract the electronic signature spatial feature. Sub-network 2 is composed of three serially connected fully connected layers to extract the electronic signature semantic feature.

5. The system according to claim 4, wherein The sub-network 1 takes all the feature vectors of the stroke as input, uses a one-dimensional convolutional layer to extract the spatial features of the stroke, and uses a max pooling layer and a fully connected layer to merge the point-by-point features into stroke-by-stroke features to obtain the stroke trajectory features; The sub-network 2 uses the position feature in the stroke feature vector as input, and uses three cascaded fully-connected layers to extract the semantic features of the strokes according to the stroke position features, and concatenates the spatial features and the semantic features.

6. The system according to claim 2, wherein The further division of the pen tip movement trajectory into strokes includes that the preprocessing part calculates the normalized coordinate values ( , ) of the character trajectory point coordinates ( ), and calls the formula: to calculate the stroke rotation angle . According to the pen-up / pen-down information of the stroke feature points and the rotation angle change, the starting point of the stroke is determined. When the pen-up / pen-down information of a certain point of the stroke is 0 and the rotation angle change is greater than a predetermined angle, this point is the starting point of the stroke. The number of strokes is determined according to the starting point of the stroke, and the electronic signature is divided into strokes. Among them, represents the derivative with respect to time, represents the derivative with respect to time.

7. The system according to any one of claims 1-3 and 6, characterized in that, The dual-channel stroke feature extraction module extracts the feature vector sequence of the direction vector, state vector, and position vector contained in the electronic signature stroke, concatenates the above feature vector sequences to obtain the linear difference of the stroke segment to maintain the signature stroke trajectory contour, and represents the text feature vector according to the set maximum number of strokes and linear difference.

8. The system according to claim 7, wherein According to the formula: , Divide the eight-direction set into two subsets and , for any point t in the stroke trajectory ij The corresponding turning angle Decompose it into two adjacent directions on the two subsets respectively and , according to the formula: Calculate two adjacent directions and corresponding values , , according to the formula: ; ; ; ; Calculate any stroke point of the two directional feature vectors , , state vector and position vector , where is the start / end information of the stroke point , is the normalized coordinate value represents the indicator function is the standard basis vector; , is the direction transformation value corresponding to two adjacent directions and , d and d' respectively represent the angular projection values of the rotation angle for each element in the two subsets and .

9. The system according to claim 4, characterized in that, The position features of the start point and end point of the electronic signature stroke are input to the second fully connected layer of sub-network 2, and the position features of the middle point of the electronic signature stroke are input to the first fully connected layer of sub-network 2.

10. The system according to any one of claims 1-3 and 6, characterized in that, Sequence Tagging Based on Conditional Random Field Prediction The negative log-likelihood function is used to determine the word segmentation prediction loss , and the cross-entropy function of the classification label and the classification prediction result is used to determine the classification prediction loss , and the PMLNet network loss value is the weighted sum of the word segmentation prediction loss and the classification prediction loss. Specifically, according to the formula calculate the PMLNet network loss value , where represents the set of all losses , C represents the set of weights corresponding to the losses and represent minimizing the intermediate value prediction loss of word segmentation and classification and is to minimize the KL divergence of word segmentation and classification 11. An online character cutting recognition method based on the progressive multi-task learning network PMLNet, characterized in that, Including: Extract stroke sampling points to obtain the pen tip movement trajectory and position features, perform normalization processing, divide the movement writing trajectory into signature strokes, encode the electronic signature stroke position information and trajectory information into a vector with a fixed dimension, and multiple strokes form a stroke sequence feature. Obtain the spatial feature and semantic feature of the signature according to the stroke trajectory information and position information; The superimposed encoding module composed of multiple stacked Transformer encoders models the long-range dependence relationship between different strokes and encodes the stroke sequence into a sequence feature; According to the electronic signature spatial feature, semantic feature, and sequence feature, the electronic signature is cut into single characters and recognized by parallel cutting character branches and recognition branches. The intermediate results of the cutting character branch and the recognition branch are interacted to optimize and obtain the cutting character result and recognition result of the high-quality electronic signature. Specifically, the PMLNet network is used for progressive and interactive cutting and recognition. The recognition classifier predicts the character category to which each stroke belongs, and the cutting character classifier predicts each stroke category. The statistical graph module simulates the Gaussian distribution of two categories, and balances the cutting character and classification intermediate results according to the PMLNet network loss value and KL divergence to obtain high-quality cutting character results and classification results.

12. The method according to claim 11, wherein Said dividing the motion writing trajectory into signature strokes further includes calculating the normalized coordinate values ( , ) of the electronic signature trajectory points ( ), and calling the formula: to calculate the stroke rotation angle . According to the pen-up / pen-down information of the stroke feature points and the rotation angle change, the starting point of the stroke is determined. When the pen-up / pen-down information of a certain point of the stroke is 0 and the rotation angle change is greater than a predetermined angle, this point is the starting point of the stroke. According to the starting point of the stroke, the number of strokes is determined, and the signature is divided into strokes. Among them, represents the derivative with respect to time, represents the derivative with respect to time.

13. The method according to claim 11, wherein Extract the feature vector sequence containing the direction vector, state vector, and position vector in the stroke. Concatenate the above feature vector sequences to obtain the linear difference of the stroke segment, and maintain the outline of the electronic signature stroke trajectory. Represent the feature vector according to the set maximum number of strokes and linear difference.

14. The method according to any one of claims 11-13, characterized in that The obtaining of the spatial features and semantic features of the signature based on the stroke trajectory information and position information further includes: all the feature vectors of the text strokes Input channel 1, using a one-dimensional convolutional layer to extract the spatial features of the strokes; using a max pooling layer and a fully connected layer to merge the point-by-point features into stroke-by-stroke features to obtain stroke trajectory features; the position features in the signature stroke feature vectors Input channel 2, using three cascaded fully connected layers to extract the semantic features of the strokes according to the stroke position features, and cascading the spatial features and semantic features.

15. The method according to any one of claims 11-13, characterized in that The position features of the starting point and ending point of the signature stroke are input to the second fully connected layer of channel 2, and the position features of the intermediate point of the signature stroke are input to the first fully connected layer of channel 2.

16. The method according to claim 13, wherein According to the formula: , Divide the eight-direction set into two subsets and , for any point t in the stroke trajectory ij The corresponding turning angle Decompose it into two adjacent directions on the two subsets respectively and , according to the formula: ; ; Calculate the values corresponding to two adjacent directions and respectively , , According to the formula: ; ; ; ; Calculate arbitrary stroke points of two directional feature vectors , , state vector and position vector , where is the start / end pen information of the stroke point , is the normalized coordinate value represents the indicator function is the standard basis vector; , is the direction transformation value corresponding to two adjacent directions and , d and d′ respectively represent the angle projection values of the corner angles in two subsets and of each element.

17. The method according to claim 16, characterized in that, Sequence tagging based on conditional random field prediction The negative log-likelihood function is used to determine the word segmentation prediction loss , and the cross-entropy function of the classification label and the classification prediction result is used to determine the classification prediction loss , and the loss value of the PMLNet network is the weighted sum of the word segmentation prediction loss and the classification prediction loss. Specifically, the loss value of the PMLNet network is calculated according to the formula : , Among them, represents the set of all losses, , C represents the set of weights corresponding to the losses, and represent minimizing the word segmentation prediction loss and the classification prediction loss, and is to minimize the KL divergence of word segmentation and classification.

18. An electronic device, characterized in that, Including: One or more processors, a memory, and one or more applications, which are stored in the memory and configured to be loaded and run by the one or more processors to execute the online character cutting recognition method based on the progressive multi-task learning network PMLNet according to one of claims 11-17.

19. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, it implements the online character cutting recognition method based on the progressive multi-task learning network PMLNet according to one of claims 11-17.

Citation Information

Patent Citations

  • A character recognition system and method based on the combination of neural networks and attention mechanisms

    CN109389091B

  • Electric power operation ticket character recognition method based on convolutional neural network

    CN111259880A

  • Online handwritten character recognition method and device, electronic equipment and storage medium

    CN111931710A

  • Method and system for recognizing stroke order of handwritten Chinese character and calculation equipment

    CN113627260A

  • Real-time identification method for on-line handwriting sentences

    CN101853126A