A Handwritten Handwriting Recognition System Based on a Continuous Convolution SPP Network
Through the handwriting recognition system based on continuous convolution SPP network, the problem of handwriting recognition in the prior art relying on fixed templates and professional equipment is solved, and efficient and accurate recognition effect is achieved, which is suitable for normalized evaluation.
Patent Information
- Application Number
- CN202210469762.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-04-30
AI Technical Summary
The existing handwriting recognition technology relies on a fixed polar coordinate expansion algorithm, requiring subjects to draw on a prescribed writing template, and requires professional equipment to collect dynamic handwriting features, resulting in large recognition errors and is not suitable for normalized evaluation.
A handwriting recognition system based on a continuous convolution SPP network is adopted, including an input unit, a preprocessing unit, an intelligent recognition unit and an output unit. The hand-painted images are processed smoothly, cropped and binarized, and the continuous convolution SPP network model is used for pre-training and analysis to identify the features of Archimedes spirals. Without professional equipment, images are directly collected using ordinary paper and pen and scanning equipment.
It improves the accuracy and efficiency of handwriting recognition, reduces labor costs, is suitable for normalized evaluation, and retains the original characteristics, reducing the recognition costs.
Smart Images

Figure CN115035536B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition, and particularly relates to a handwritten handwriting recognition system based on a continuous convolutional SPP network. Background Art
[0002] Tremor is described as an involuntary, rhythmic, and oscillatory movement of one or more parts of the body. Tremors are generally divided into resting tremors and action tremors. Resting tremors occur when the body part is in a stationary position and there is no voluntary muscle activity. However, action tremors occur with voluntary muscle contractions. The most common forms of tremors are physiological tremors, essential tremors, and Parkinsonian tremors, mainly seen in the elderly. Mental stress or voluntary behavior can increase the amplitude and frequency of Parkinsonian tremors. Mental stress also increases essential tremors, but the most important point is that essential tremors occur while maintaining a position (posture) against gravity and may occur during muscle contractions against gravity.
[0003] As a complex activity involving both motor and perceptual components, handwritten handwriting is an effective marker of hand tremors. Handwriting analysis has developed into a clinical tool in neurology, and a simple pen test can be completed in just a few seconds. By observing the patient's writing process and analyzing the form and content of the hand-drawn samples, professionals can analyze and classify the corresponding hand tremor types. The spiral line, as a pattern without a usage scenario and language restriction, is the preferred pattern for handwritten handwriting analysis.
[0004] In recent years, in order to make the auxiliary discrimination of tremor types based on handwritten handwriting more convenient and fast, researchers have introduced machine learning and deep learning technologies. Although relevant research has achieved certain results and established a variety of handwritten datasets, there are still the following problems: First, the current spiral line feature classification method usually uses a fixed polar coordinate expansion algorithm, resulting in most handwritten datasets requiring subjects to draw on a specified writing template; Second, the acquisition of handwriting dynamic features can only rely on professional data acquisition equipment, and the acquisition process is strict and cumbersome, and most require designing a dedicated data processing system to complete feature recognition and sample classification, so it is not suitable for routine evaluation; Third, most classification networks for hand-drawn images will perform scaling or random cropping operations on the images at the input end, and for hand-drawn images with shape characteristics, these operations will change the original features, causing certain recognition errors. Summary of the Invention
[0005] In view of this, the present invention provides a handwritten handwriting recognition system based on a continuous convolutional SPP network, including an input unit, a preprocessing unit, an intelligent recognition unit, an output unit, a user terminal, and a cloud, wherein,
[0006] The input unit, the preprocessing unit, the intelligent recognition unit, and the output unit are connected in sequence. The output of the user terminal is connected to the input unit, the output of the output unit is connected to the cloud and the user terminal, and the cloud is also connected to the preprocessing unit and the intelligent recognition unit;
[0007] The input of the input unit is hand-drawn image data; the preprocessing unit performs smoothing denoising, binarization, and cropping; the intelligent recognition unit pre-trains, analyzes, and performs model voting on the continuous convolutional SPP network model to obtain a recognition result; the output unit sends the recognition result to the user terminal and the cloud respectively.
[0008] Preferably, the input unit processes the hand-drawn image data into an Archimedes spiral image.
[0009] Preferably, the preprocessing unit includes a filtering module, a binarization module, and a partitioning module. Among them, the filtering module performs smoothing denoising on the image, the binarization module binarizes the filtered image according to the color range of the handwriting, and the partitioning module crops the original spiral image into four parts: upper left, lower left, upper right, and lower right.
[0010] Preferably, the binarization formula of the binarization module is:
[0011]
[0012] In the formula, F represents the original image, and N i (F) represents the updated image, and R i (F), G i (F), B i (F) respectively represent the values of the red, green, and blue channels of the i-th pixel in F.
[0013] Preferably, the continuous convolutional SPP network model includes an input layer, a continuous convolutional layer, a pooling layer, an SPP spatial pyramid pooling layer, and a fully connected layer.
[0014] Preferably, the intelligent recognition unit labels the obtained Archimedes spiral image, and different categories use numbers 0, 1, 2,..., n as labels;
[0015] The labeled images are divided into a training set, a validation set, and a test set according to a certain ratio;
[0016] Build a continuous convolutional SPP network model;
[0017] Put the labeled training set and validation set into the input end, perform multi-size training to obtain an output result, and train different models according to each local feature;
[0018] Calculate the Loss value through the loss function, and update the network weights through backpropagation. The loss function uses the cross-entropy loss function, and its calculation formula is:
[0019] H(p,q)=-∑ i p i log2q i
[0020] Among them, p represents the true value of the image, that is, the probability distribution of the original label, and q represents the model prediction value, that is, the predicted label probability distribution;
[0021] Adjust the number of iterations epoch according to the size of the training set, and repeat the training until the model converges.
[0022] Preferably, the continuous convolutional SPP network model includes an input layer, 2 continuous convolutional modules, an SPP spatial pyramid pooling layer, and a fully connected layer. Among them, the 2 continuous convolutional modules include module A and module B, and the last layer of module A is a max pooling layer; in module A, i layers of continuous convolutional layers are set, and the value of i changes according to the complexity of the hand-drawn image. In module B, a fixed three-layer continuous convolutional structure is built. The SPP spatial pyramid pooling layer is connected after module B. The SPP spatial pyramid pooling layer consists of three parallel pooling layers with sizes of 4×4, 2×2, and 1×1 respectively, and extends the image from the previous layer to the input of the fully connected layer with a fixed length; regardless of the size of the input image, the operation process of the fully connected layer remains unchanged, and the ReLU activation function is used for both the convolutional layer and the pooling layer kernels.
[0023] Preferably, the training of the continuous convolutional SPP network model by the intelligent recognition unit includes:
[0024] S301: Randomly divide the labeled data in the database into a training set, a validation set, and a test set according to a ratio of 8:1:1. The input layer of the network scales the training set to multiple sizes: 180×180, 224×224, 360×360, with the unit of pixels. When the original image is smaller than the unified size, the bilinear interpolation method is used to enlarge the image, and the formula is:
[0025] P(x,y)=P(x1,y1)(1-x)(1-y)+P(x2,y1)x(1-y)
[0026] +P(x1,y2)(1-x)y+P(x2,y2)xy
[0027] In the formula, (x1,y1), (x2,y1), (x1,y2), and (x2,y2) represent the four nearest neighbor points of the point to be interpolated, and P(x,y) represents the pixel value of the point with the abscissa x and the ordinate y in the pixel coordinate system;
[0028] S302: The multi - sized images of the training set are input into the network to obtain the corresponding output results. Calculate the loss value according to the cross - entropy loss function. For multi - classification, the formula is:
[0029]
[0030] In the formula, M represents the total number of categories; N represents the number of samples in a training batch; p(x ij ) indicates whether the category is the same as the category of sample i. If it is the same, it is 1; if it is different, it is 0; q(x ij ) represents the predicted probability that the observed sample i belongs to category j;
[0031] S303: Initialize the learning rate lr = 0.00001. During the training process, use backpropagation for parameter update, and use the Adam algorithm to optimize the model. Calculate the backpropagation algorithm according to the following three situations:
[0032] 1) When the current layer is a fully - connected layer, δ i,l =(W l+1 ) T δ i,l+1 ⊙σ′(z i,l )
[0033] 2) When the current layer is a convolutional layer, δ i,l =δ i,l+1 *rot180(W l+1 )⊙σ′(z i,l )
[0034] 3) When the current layer is a pooling layer, δ i,l =upsample(δ i,l+1 )⊙σ′(z i,l )
[0035] In the backpropagation calculation, δ i,l represents the gradient of the l - th layer, δ i,l+1 represents the gradient of the (l + 1) - th layer, W l+1 represents the weight of the (l + 1) - th layer, z i,l is the intermediate output result of the i - th image of the l - th layer. The ReLU activation function is represented by σ(z);
[0036] S304: Take epoch = 2000 as the benchmark. Increase epoch according to the expansion of the training set, and repeat the training until the model converges.
[0037] Preferably, in the intelligent recognition unit, the discrimination probabilities for the input hand - drawn images are different. Take the highest probability output for each category as the discrimination result.
[0038] Preferably, the intelligent recognition unit sends the results of each model vote, the visual feature analysis, and the final discrimination conclusion to the user side through the output unit.
[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0040] (1) Traditional means for evaluating handwritten scripts mainly rely on the subjective judgment of neurologists or experts. The present invention uses deep learning methods to identify the characteristics of Archimedes spiral lines, effectively saving labor costs while improving the recognition accuracy.
[0041] (2) The present invention includes a continuous convolutional SPP network architecture, which can not only extract the tremor characteristics of spiral line handwriting, but also extract the changes in the pitch between turns and the shape characteristics of the spiral line, and classify various types of tremor hand-drawn graphs at the same time.
[0042] (3) The present invention only requires the subject to provide handwritten script data and does not require professional dynamic acquisition equipment, greatly saving the cost of recognition.
[0043] (4) The input end of the present invention does not need to perform scaling and cropping operations on the hand-drawn graph, retaining all the characteristics of the original spiral line and ensuring the high efficiency and accuracy of deep learning. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to make the objectives, technical solutions, and beneficial effects of the present invention clearer, the following drawings are provided for illustration:
[0045] Figure 1 It is a structural block diagram of a handwritten script recognition system based on a continuous convolutional SPP network according to an embodiment of the present invention;
[0046] Figure 2 It is a flow chart of an intelligent recognition unit of a handwritten script recognition system based on a continuous convolutional SPP network according to an embodiment of the present invention;
[0047] Figure 3 It is a schematic diagram of the structure of a continuous convolutional SPP network of a handwritten script recognition system based on a continuous convolutional SPP network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The preferred embodiments of the present invention will be described in detail below in conjunction with the drawings.
[0049] See Figure 1, shown is a handwritten handwriting recognition system based on a continuous convolutional SPP network of the present invention. The input unit 10 collects the hand-drawn Archimedes spiral image of the subject, the preprocessing unit 20 preprocesses the hand-drawn image, the intelligent recognition unit 30 analyzes the image features, and the recognition conclusion is obtained by integrating the model voting results; the output unit 40 feeds back the conclusion to the user terminal 50; the recognition conclusion is also uploaded to the cloud 60 for model optimization.
[0050] During the acquisition process of the input unit 10, the subject can use A4 paper and an ordinary black signature pen for drawing, without external drawing interference and copying templates. After the drawing is completed, the image can be scanned and uploaded using scanning devices such as mobile phones and iPads.
[0051] In the preprocessing unit 20, different drawing pens will result in Archimedes spiral lines of different colors, and at the same time, the scanning device will also affect the image quality. The preprocessing unit 20 specifically includes:
[0052] S201: Crop the area containing only the spiral line according to the bounding box. The size of the scanned and cropped original image is 1135 pixels × 1135 pixels. Therefore, the ratio of pixels to the actual distance is 7.567 pixels / mm;
[0053] S202: Smooth the image using a 3*3 mean filter;
[0054] S203: Use a 3*3 median filter to remove salt-and-pepper noise and random noise in the image;
[0055] S204: Since the colors of different images in the dataset are different, we binarize the image, and the spiral line is the foreground in the binary image. The mathematical formula for this step is:
[0056]
[0057] In the formula, F represents the original image, N i (F) represents the updated image, R i (F), G i (F), B i (F) respectively represent the red, green, and blue channel values of the i-th pixel in F;
[0058] S205: The binarized image is cropped into four local feature maps of the upper left, upper right, lower left, and lower right, plus the uncropped image, a total of five types of datasets, and data augmentation is performed by rotation and other methods.
[0059] In the intelligent recognition unit 30, the continuous convolutional SPP network model constructed by the present invention is used. The specific structure of the model is as Figure 3As shown in the figure. The continuous convolutional SPP network as a whole consists of an input layer, two continuous convolutional modules (Module A and Module B), an SPP spatial pyramid pooling layer, and a fully connected layer. The last layer of Module A is a max pooling layer, and the ReLU activation function is used for both the convolutional layer and the pooling layer kernels. In Module A, in order to identify the shape and spacing features of the spiral dataset, we designed i consecutive convolutional layers, and the value of i can be changed according to the complexity of the hand-drawn graph ( Figure 2 in Figure 2 , i = 2), so the continuous convolutional SPP network designed in the present invention can also be used in the recognition of other types of hand-drawn images; in the design of Module B, in order to extract more subtle features in the spiral image and reduce the dimension of the image, we built a fixed three-layer continuous convolutional structure.
[0060] The SPP layer is connected after Module B. The SPP layer consists of three parallel pooling layers with sizes of 4×4, 2×2, and 1×1 respectively, and extends the image from the previous layer into the input of the fully connected layer with a fixed length. Regardless of the size of the image input to the network, the operation process of the fully connected layer is the same.
[0061] In the intelligent recognition unit 30, five models with different parameters are trained according to the five groups of preprocessed data sent by the processing unit. See Figure 2 , and the specific steps for training the continuous convolutional SPP network model are as follows:
[0062] S301: The labeled data in the database is randomly divided into a training set, a validation set, and a test set according to a ratio of 8:1:1. The input layer of the network scales the training set to multiple sizes: 180×180, 224×224, 360×360 (unit: pixel). When the original image is smaller than the unified size, the bilinear interpolation method is used to enlarge the image. The formula is:
[0063] P(x,y) = P(x1,y1)(1 - x)(1 - y) + P(x2,y1)x(1 - y)
[0064] + P(x1,y2)(1 - x)y + P(x2,y2)xy
[0065] In the formula, (x1,y1), (x2,y1), (x1,y2), and (x2,y2) represent the four nearest neighbor points of the point to be interpolated, and P(x,y) represents the pixel value of the point with abscissa x and ordinate y in the pixel coordinate system;
[0066] S302: The multi-sized images of the training set are input into the network to obtain the corresponding output results. The loss value is calculated according to the cross-entropy loss function. The formula for multi-classification is:
[0067]
[0068] Wherein, M represents the total number of categories; N represents the number of samples in a training batch; p(x ij ) indicates whether the category is the same as the category of sample i. If it is the same, it is 1; if it is different, it is 0; q(x ij ) represents the predicted probability that the observed sample i belongs to category j;
[0069] S303: Initialize the learning rate lr = 0.00001. During the training process, backpropagation is used for parameter update, and the Adam algorithm is used to optimize the model. The backpropagation algorithm is calculated according to the following three situations:
[0070] 1) When the current layer is a fully connected layer, δ i,l =(W l+1 ) T δ i,l+1 ⊙σ′(z i,l )
[0071] 2) When the current layer is a convolutional layer, δ i,l =δ i,l+1 *rot180(W l+1 )⊙σ′(z i,l )
[0072] 3) When the current layer is a pooling layer, δ i,l =upsample(δ i,l+1 )⊙σ′(z i,l )
[0073] In the backpropagation calculation, δ i,l represents the gradient of the l-th layer, δ i,l+1 represents the gradient of the (l + 1)-th layer, W l+1 represents the weight of the (l + 1)-th layer, z i,l is the intermediate output result of the i-th image of the l-th layer, and the ReLU activation function is represented by σ(z);
[0074] S304: Take epoch = 2000 as the benchmark, increase epoch according to the expansion of the training set, and repeat the training until the model converges;
[0075] In the intelligent recognition unit 30, the discrimination probabilities of each model for the input hand-drawn image are different. The highest probability output for each category is taken as the discrimination result; the voting results, visual feature analysis, and final discrimination conclusions of each model are organized into a document and sent to the user terminal 50 through the output unit 40;
[0076] The previously drawn images will be sent by the system to the cloud 60 for storage, and the system will retrain and optimize the model according to the latest data set.
[0077] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A handwritten handwriting recognition system based on a continuous convolutional SPP network, characterized in that, It includes an input unit, a preprocessing unit, an intelligent recognition unit, an output unit, a client and a cloud. Among them, the input unit, the preprocessing unit, the intelligent recognition unit and the output unit are connected in sequence. The output of the client is connected to the input unit, and the output of the output unit is connected to the cloud and the client. The cloud is also connected to the preprocessing unit and the intelligent recognition unit; the input of the input unit is hand-drawn image data; the preprocessing unit performs smoothing denoising, binarization and cropping; the intelligent recognition unit pre-trains, analyzes and performs model voting on the continuous convolutional SPP network model to obtain a recognition result; the output unit sends the recognition result to the client and the cloud respectively; the continuous convolutional SPP network model includes an input layer, 2 continuous convolutional modules, an SPP spatial pyramid pooling layer and a fully connected layer. Among them, the 2 continuous convolutional modules include module A and module B, and the last layer of module A is a max pooling layer; in module A, i layers of continuous convolutional layers are set, and the value of i changes according to the complexity of the hand-drawn image. In module B, a fixed three-layer continuous convolutional structure is built. The SPP spatial pyramid pooling layer is connected after module B. The SPP spatial pyramid pooling layer consists of three parallel pooling layers with sizes of 4×4, 2×2, and 1×1 respectively, and extends the image from the previous layer into a fully connected layer input with a fixed length; regardless of the size of the input image, the operation process of the fully connected layer remains unchanged, and the ReLU activation function is used for both the convolutional layer and the pooling layer kernels; the training of the continuous convolutional SPP network model by the intelligent recognition unit includes: S301: The labeled data in the database is randomly divided into a training set, a validation set and a test set according to the ratio of 8:1:
1. The input layer of the network scales the training set to multiple sizes: 180×180, 224×224, 360×360, with the unit of pixel. When the original image is smaller than the unified size, the bilinear interpolation method is used to enlarge the image, and the formula is: P(x,y)=P(x1,y1)(1-x)(1-y)+P(x2,y1)x(1-y) +P(x1,y2)(1-x)y+P(x2,y2)xy In the formula, (x1,y1), (x2,y1), (x1,y2) and (x2,y2) represent the four nearest points of the points to be interpolated, and P(x,y) represents the pixel value of the point with the abscissa x and the ordinate y in the pixel coordinate system; S302: The multi-size images of the training set are input into the network to obtain the corresponding output results; the loss value is calculated according to the cross-entropy loss function. The formula for multi-classification is: where M represents the total number of categories; N represents the number of samples in a training batch; p(x ij ) indicates whether the category of this sample is the same as that of sample i, with the same being 1 and different being 0; q(x ij ) represents the predicted probability that the observed sample i belongs to category j; S303: Initialize the learning rate lr = 0.00001. During the training process, backpropagation is used for parameter update, and the Adam algorithm is used to optimize the model. The backpropagation algorithm is calculated according to the following three situations: 1) The current layer is a fully connected layer, δ i,l =(W l+1 ) T δ i,l+1 ⊙σ′(z i,l ) 2) The current layer is a convolutional layer, δ i,l = δ i,l+1 * rot180(W l+1 ) ⊙ σ′(z i,l ) 3) The current layer is a pooling layer, δ i,l = upsample(δ i,l+1 ) ⊙ σ′(z i,l ) In backpropagation calculation, δ i,l represents the gradient of the l-th layer, and δ i,l+1 represents the gradient of the (l + 1)-th layer. W l+1 represents the weights of the (l + 1)-th layer, and z i,l is the intermediate output result of the i-th image of the l-th layer. The ReLU activation function is denoted as σ(z); S304: Take epoch = 2000 as the benchmark, increase epoch according to the expansion of the training set, and repeat the training until the model converges.
2. The system according to claim 1, wherein The input unit processes the hand-drawn image data into an Archimedes spiral image.
3. The system according to claim 1, wherein The preprocessing unit includes a filtering module, a binarization module, and a partitioning module. Among them, the filtering module performs smoothing and denoising on the image, the binarization module binarizes the filtered image according to the color range of the handwriting, and the partitioning module crops the original spiral image into four parts: upper left, lower left, upper right, and lower right.
4. The system according to claim 3, characterized in that, The binarization formula of the binarization module is: Wherein, F represents the original image, and N i (F) represents the updated image, and R i (F), G i (F), B i (F) respectively represent the values of the red, green, and blue channels of the i-th pixel in F.
5. The system according to claim 1, wherein The continuous convolutional SPP network model includes an input layer, a continuous convolutional layer, a pooling layer, an SPP spatial pyramid pooling layer, and a fully connected layer.
6. The system according to claim 5, wherein The intelligent recognition unit labels the obtained Archimedes spiral image, and uses numbers 0, 1, 2, …, n as labels for different categories; The labeled images are divided into a training set, a validation set, and a test set according to a certain ratio; Build a continuous convolutional SPP network model; Put the labeled training set and validation set into the input end, perform multi-size training to obtain the output result, and train different models according to each local feature; Calculate the Loss value through the loss function, and update the network weights by backpropagation. The loss function uses the cross-entropy loss function, and its calculation formula is: H(p,q) = -∑ i p i log2q i Among them, p represents the true value of the image, that is, the probability distribution of the original label, and q represents the model prediction value, that is, the predicted label probability distribution; Adjust the number of iterations epoch according to the size of the training set, and repeat the training until the model converges.
7. The system according to claim 1, characterized in that In the intelligent recognition unit, the discrimination probabilities of the input hand-drawn images are different, and the highest probability output for each category is taken as the discrimination result.
8. The system according to claim 1, characterized in that, The intelligent recognition unit sends the voting results of each model, the visual feature analysis, and the final discrimination conclusion to the user end through the output unit.
Citation Information
Patent Citations
Handwritten Dongba recognition method based on convolutional neural network
CN111291696A
Electric power operation ticket character recognition method based on convolutional neural network
WO2021139175A1