Handwritten character recognition model training method, handwritten character recognition method and handwritten character recognition device
By preprocessing handwritten text data and training a convolutional neural network based on attention mechanism, the problem of insufficient feature extraction ability of handwritten text recognition methods in the prior art under complex fonts and writing styles is solved, and the accuracy and robustness of recognition are improved.
Patent Information
- Application Number
- CN202411955191.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-27
AI Technical Summary
The existing handwritten text recognition method based on deep learning has limited feature extraction capabilities when dealing with complex handwritten fonts and writing styles, and it is difficult to dig up effective feature information, resulting in low recognition accuracy and poor robustness.
By pre-processing the first text data by image scaling, grayscale and denoising, the pre-processed data set is obtained, and the first convolutional neural network is iteratively trained as the training sample. The convolutional neural network designed by the attention mechanism is used to update the weights through the backpropagation algorithm to obtain the handwritten text recognition model.
It improves the accuracy and robustness of handwritten text model recognition, and can more effectively deal with complex handwritten fonts and writing styles to meet the needs of practical applications.
Smart Images

Figure CN120047959A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and in particular, to a method for training a handwritten text recognition model, a method for recognizing handwritten text, and a device therefor. Background Art
[0002] The recognition of handwritten characters and numbers has always been one of the research hotspots in the fields of computer vision and artificial intelligence; with the development of deep learning technology, significant progress has been made in the methods for recognizing handwritten characters and numbers.
[0003] Currently, the methods for recognizing characters and numbers in handwritten documents based on deep learning have been widely applied to various scenarios, such as postal express delivery, financial bill recognition, handwritten digit input, etc.
[0004] In related technologies, traditional methods for recognizing handwritten characters and numbers usually require a large amount of feature engineering and manually designed feature extraction algorithms. Especially for complex handwritten fonts and writing styles, the feature representation ability of manually designed features is limited, and the neural networks adopted cannot obtain more valuable character or number information, resulting in low accuracy and poor robustness of model recognition.
[0005] Therefore, it is necessary to further research and improve the methods for recognizing characters and numbers in handwritten documents based on deep learning to improve the accuracy, speed, and robustness of recognition to meet the requirements of practical applications. Summary of the Invention
[0006] The present invention provides a method for training a handwritten text recognition model, a method for recognizing handwritten text, and a device therefor, which are used to solve the defect that in the prior art, when dealing with complex handwritten fonts and writing styles, the feature representation ability of the extracted features is limited, and it is difficult to mine effective feature information by using a neural network, resulting in low accuracy of handwritten text model recognition, and improve the accuracy and robustness of handwritten text model recognition.
[0007] The present invention provides a method for training a handwritten text recognition model, including: Performing preprocessing of image scaling, grayscale conversion, and denoising on first text data to obtain a preprocessed data set; wherein, the first text data includes at least one of a text data set in a publicly available database and user handwritten text data; Iteratively training a first convolutional neural network with the preprocessed data set as training samples, using the text writing features as input features, and updating the weights of the first convolutional neural network through the backpropagation algorithm, and obtaining a handwritten text recognition model when the first convolutional neural network converges; wherein, the first convolutional neural network is designed by an LSTM network based on an attention mechanism, a LeNet architecture, and an AlexNet architecture.
[0008] According to a method for training a handwritten character recognition model provided by the present invention, after obtaining the handwritten character recognition model, the method further includes: Constructing a second convolutional neural network based on the convolutional layer and the target fully connected layer of the first convolutional neural network; Using the preprocessed second text data as a training sample to iteratively train the second convolutional neural network, and fine-tuning the learning rate of the second convolutional neural network. When the second convolutional neural network converges, a new handwritten character recognition model is obtained.
[0009] According to a method for training a handwritten character recognition model provided by the present invention, after obtaining the new handwritten character recognition model, the method further includes: Calculating a target evaluation index corresponding to the new handwritten character recognition model with the preprocessed third sample as a test sample to obtain an evaluation result, and optimizing the new handwritten character recognition model according to the evaluation result; wherein, the target evaluation index includes at least one of accuracy, recall rate, F1 value, and average edit distance.
[0010] According to a method for training a handwritten character recognition model provided by the present invention, the character writing features include at least one of stroke direction and trajectory, connectivity features, scale and proportion, local shape features, and morphological features.
[0011] The present invention also provides a handwritten character recognition method, including: Obtaining a handwritten character to be recognized; Processing the handwritten character to be recognized based on the handwritten character recognition model to obtain a handwritten character recognition result; wherein, the handwritten character recognition model is trained by the handwritten character recognition model training method.
[0012] The present invention also provides a handwritten character recognition model training device, including: An image preprocessing module for preprocessing the first text data by image scaling, grayscale conversion, and denoising to obtain a preprocessed data set; wherein, the first text data includes at least one of a text data set in a publicly available database and user handwritten text data; A training module for iteratively training a first convolutional neural network with the preprocessed data set as a training sample, using the character writing features as input features, and updating the weights of the first convolutional neural network through the backpropagation algorithm. When the first convolutional neural network converges, a handwritten character recognition model is obtained; wherein, the first convolutional neural network is designed based on an LSTM network with an attention mechanism, a LeNet architecture, and an AlexNet architecture.
[0013] The present invention also provides a handwritten text recognition device, including: A handwritten text acquisition module for acquiring handwritten text to be recognized; A recognition module for processing the handwritten text to be recognized based on a handwritten text recognition model to obtain a handwritten text recognition result; wherein, the handwritten text recognition model is trained by the handwritten text recognition model training method.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the handwritten text recognition model training method or the handwritten text recognition method described in any one of the above is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the handwritten text recognition model training method or the handwritten text recognition method described in any one of the above is implemented.
[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the handwritten text recognition model training method or the handwritten text recognition method described in any one of the above is implemented.
[0017] The handwritten text recognition model training method, the handwritten text recognition method, and the device provided by the present invention perform preprocessing on the first text data, including image scaling, grayscale conversion, and denoising, and use the preprocessed data set as a training sample to iteratively train the first convolutional neural network. Using the text writing features as input features and updating the weights of the first convolutional neural network through the backpropagation algorithm, and obtaining a handwritten text recognition model when the first convolutional neural network converges, which improves the robustness of the handwritten text model recognition and further improves the accuracy of handwritten text recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 is one of the flow schematic diagrams of the handwritten text recognition model training method provided by the present invention.
[0020] Figure 2 is the structural schematic diagram of the composition of the text writing features provided by the present invention.
[0021] Figure 3 It is the second flow schematic diagram of the handwritten text recognition model training method provided by the present invention.
[0022] Figure 4 It is the flow schematic diagram of the handwritten text recognition method provided by the present invention.
[0023] Figure 5 It is the structural schematic diagram of the handwritten text recognition model training device provided by the present invention.
[0024] Figure 6 It is the structural schematic diagram of the handwritten text recognition device provided by the present invention.
[0025] Figure 7 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed Embodiments
[0026] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts shall fall within the protection scope of the present invention.
[0027] The following combines Figures 1-6 to describe the handwritten text recognition model training method, handwritten text recognition method and device of the present invention.
[0028] Figure 1 It is one of the flow schematic diagrams of the handwritten text recognition model training method provided by the present invention. As Figure 1 shown, the handwritten text recognition model training method includes the following steps: Step 110: Perform preprocessing of image scaling, grayscale conversion and denoising on the first text data to obtain a preprocessed data set; wherein, the first text data includes at least one of the text data set in the publicly available database and the user's handwritten text data.
[0029] In this step, the first text data may be the text data set in the publicly available database. For example, the first text data is the MNIST data set, the ASCII character data set and the Chinese data set; the first text data may also be the user's handwritten characters and digital materials, such as handwritten fonts with different writing styles; the first text data may also include a mixed data set of publicly available data and user's handwritten characters.
[0030] In this embodiment, preprocessing the collected first text data includes operations such as image enhancement (such as rotation, scaling, cropping, etc.), contrast adjustment, normalization, etc., to improve the generalization ability of the model. At the same time, ensure that the image sizes are unified to facilitate model training.
[0031] Specifically, before training the model with the first text data, it is necessary to preprocess the first text data, which helps to reduce the data calculation amount and improve the data quality. Among them, image scaling algorithms such as neighborhood interpolation, bilinear interpolation, bicubic interpolation, or supersampling can be used to scale the first text data; the weighted average method, maximum value method, or average value method can be used to grayscale the scaled first text data; methods such as mean filtering, median filtering, Gaussian filtering, or Wiener filtering can be used to denoise the grayscale image. It should be noted that this embodiment does not specifically limit the order of the above preprocessing operations.
[0032] In this embodiment, after preprocessing the first text data, the data set is divided into a training set, a validation set, and a test set according to a certain ratio (for example, 6:2:2); among them, the training set is used to update the model parameters, the validation set is used to adjust the hyperparameters and control the training process, and the test set is used to evaluate the model performance.
[0033] Step 120: Iteratively train the first convolutional neural network with the preprocessed data set as the training samples, use the text writing features as the input features, and update the weights of the first convolutional neural network through the backpropagation algorithm. When the first convolutional neural network converges, a handwritten text recognition model is obtained; among them, the first convolutional neural network is designed through an LSTM network based on the attention mechanism, the LeNet architecture, and the AlexNet architecture.
[0034] In this embodiment, the first convolutional neural network is composed of a convolutional neural network, a bidirectional long short-term memory network, and an LSTM network using the attention mechanism; among them, the CNN network is responsible for extracting the feature sequence of the input image, the BiLSTM network is responsible for capturing the forward and backward dependencies in the feature sequence to form a semantic vector encoding; the LSTM network using the attention mechanism is responsible for decoding.
[0035] Figure 2 is a schematic structural diagram of the composition of the text writing features provided by the present invention. In Figure 2 the embodiment shown, the text writing features include at least one of stroke direction and trajectory, connectivity features, scale and proportion, local shape features, and morphological features.
[0036] Specifically, the text writing features may further include, but are not limited to, stroke direction and trajectory, connectivity features, scale and proportion, local shape features, and morphological features (such as stroke thickness, closedness), etc.
[0037] In this embodiment, the LeNet architecture may be the LeNet-5 architecture; the LeNet-5 architecture includes two convolutional layers (Conv Layer 1, Conv Layer 2), two average pooling layers (Avg Pooling Layer 1, Avg Pooling Layer 2), two fully connected layers (Fully Connected Layer 1, Fully Connected Layer 2), and an output layer (Output Layer).
[0038] Specifically, the settings of each network layer are as follows: Conv Layer 1: 6 5x5 convolutional kernels, stride of 1, padding of 0, activation function of ReLU; Avg Pooling Layer 1: 2x2 average pooling, stride of 2; Conv Layer 2: 16 5x5 convolutional kernels, stride of 1, padding of 0, activation function of ReLU; Avg Pooling Layer 2: 2x2 average pooling, stride of 2; Fully Connected Layer 1: 120 neurons, activation function of ReLU; Fully Connected Layer 2: 84 neurons, activation function of ReLU; Output Layer: 10 neurons, using the softmax activation function.
[0039] In this embodiment, the AlexNet architecture includes five convolutional layers (Conv Layer 1-5), three max pooling layers (Avg Pooling Layer 1-3), two fully connected layers (Fully Connected Layer 1-2), and an output layer (Output Layer).
[0040] Specifically, each network layer is set as follows: Conv Layer 1: 96 convolutional kernels of 11x11, stride of 4, padding of 2, activation function ReLU; Max Pooling Layer 1: Max pooling of 3x3, stride of 2; Conv Layer 2: 256 convolutional kernels of 5x5, stride of 1, padding of 2 (group convolution, 48 in each group), activation function ReLU; MaxPooling Layer 2: Max pooling of 3x3, stride of 2; Conv Layer 3: 384 convolutional kernels of 3x3, stride of 1, padding of 1, activation function ReLU; Conv Layer 4: 384 convolutional kernels of 3x3, stride of 1, padding of 1, activation function ReLU; Conv Layer 5: 256 convolutional kernels of 3x3, stride of 1, padding of 1 (group convolution, 64 in each group), activation function ReLU; Max Pooling Layer 3: Max pooling of 3x3, stride of 2; Fully ConnectedLayer 1: 4096 neurons, activation function ReLU, Dropout; Fully Connected Layer 2: 4096 neurons, activation function ReLU, Dropout; Output Layer: 1000 neurons, using softmax activation function.
[0041] In this embodiment, the shallow feature extraction ability of LeNet-5 can be combined with the deep feature learning ability of AlexNet, that is, the architecture of the first convolutional neural network is as follows: (1) Use the convolutional layers of LeNet-5 to extract preliminary features. These convolutional layers will perform convolution operations using smaller convolutional kernels (such as 5x5) and appropriate strides. (2) Connect the output of the last convolutional layer of LeNet-5 to the input of the deep convolutional layers of AlexNet to increase the depth and feature extraction ability of the model; these convolutional layers will use larger convolutional kernels (such as 11x11) and a larger number of convolutional kernels to capture more complex text features.
[0042] (3) After the convolutional layer, use a pooling layer to reduce the size of the feature map and the number of parameters, while improving the generalization ability of the model; in this embodiment, max pooling or average pooling can be selected, depending on the requirements of the task and the performance of the model. (4) After the pooling layer, use a fully connected layer to combine and integrate the features extracted from the convolutional layer; among them, the fully connected layer will use an appropriate activation function (such as ReLU) to increase the non-linearity of the model and use techniques such as dropout to prevent overfitting.
[0043] (5) The output layer designs an appropriate number of output nodes according to the specific requirements of the task (such as character recognition, word recognition, etc.); for the character recognition task, the output layer can include the number of nodes corresponding to the size of the character set, and use the softmax activation function to calculate the probability distribution of each character.
[0044] In this embodiment, cross-entropy loss and optimizers such as Adam and SGD can be selected to train the first convolutional neural network, and hyperparameters such as the learning rate and batch size can be adjusted according to the training effect of the model and the performance of the validation set. Regularization techniques such as dropout and weight decay are used to prevent the model from overfitting.
[0045] In this embodiment, algorithms such as CTC (Connectionist Temporal Classification) can be used to perform sequence decoding on the model output to improve the accuracy of text recognition.
[0046] The method for training a handwritten text recognition model provided by the present invention preprocesses the first text data by image scaling, grayscaling, and denoising, and iteratively trains the first convolutional neural network with the preprocessed data set as the training sample. Using the text writing features as input features, the weights of the first convolutional neural network are updated through the backpropagation algorithm. When the first convolutional neural network converges, a handwritten text recognition model is obtained, which improves the robustness of the handwritten text model recognition and further improves the accuracy of handwritten text recognition.
[0047] In some embodiments, after obtaining the handwritten text recognition model, the method for training the handwritten text recognition model further includes: constructing a second convolutional neural network based on the convolutional layer and the target fully connected layer of the first convolutional neural network; using the preprocessed second text data as the training sample to iteratively train the second convolutional neural network, and fine-tuning the learning rate of the second convolutional neural network. When the second convolutional neural network converges, a new handwritten text recognition model is obtained.
[0048] In this embodiment, during the transfer learning process, the second text data can be the same as or different from the first text data. The second text data includes an image data set of the required text types. These images should contain clear text, and each image should have a corresponding label (i.e., the text content).
[0049] In this embodiment, the preprocessing operations performed on the collected second text data include: image enhancement (such as rotation, scaling, cropping, etc.), contrast adjustment, normalization, etc. operations to improve the generalization ability of the model. At the same time, ensure that the image sizes are unified to facilitate model training.
[0050] In this embodiment, after preprocessing the second text data, the dataset is also divided into a training set, a validation set, and a test set according to a certain ratio (e.g., 6:2:2). The training set is used to update the model parameters, the validation set is used to adjust the hyperparameters and control the training process, and the test set is used to evaluate the model performance.
[0051] In this embodiment, a pre-trained model that has been trained on a large-scale dataset is selected, for example, the convolutional layer of the above-mentioned handwritten text recognition model; then, according to the requirements of the new task, some specific layers are added after the convolutional layer of the handwritten text recognition model. The specific layers include fully connected layers, convolutional layers, etc., and are used to map the output of the pre-trained model to the output space of the new task.
[0052] It should be noted that in the fine-tuning process of this embodiment, some or all of the shared layers in the pre-trained model (i.e., those layers that do not need to be retrained in the new task) can be selected to freeze, which helps to maintain the stability of the model on the new task.
[0053] In this embodiment, the new training set can be used to train these specific layers. During the training process, hyperparameters such as the learning rate and batch size can be adjusted to optimize the training effect.
[0054] In this embodiment, during the training process, the performance change on the validation set is monitored regularly; for example, if the performance on the validation set starts to decline (i.e., overfitting occurs), the training is stopped or other measures (such as reducing the learning rate, increasing dropout, etc.) are taken to prevent overfitting, and finally a trained new handwritten text recognition model is obtained.
[0055] In this embodiment, after obtaining the new handwritten text recognition model, the handwritten text recognition model training method further includes: calculating the target evaluation index corresponding to the new handwritten text recognition model with the preprocessed third sample as the test sample, obtaining the evaluation result, and tuning the new handwritten text recognition model according to the evaluation result; wherein, the target evaluation index includes at least one of accuracy, recall rate, F1 value, and average edit distance.
[0056] In this embodiment, the third text data can be the same as or different from the first text data and the second text data.
[0057] In this embodiment, after training is completed, the performance of the fine-tuned model can be evaluated on the test set; the evaluation metrics adopted in this embodiment include but are not limited to accuracy, recall rate, F1 score, etc.
[0058] In this embodiment, the model is adjusted according to the above index evaluation results, specifically including adjusting the network structure, hyperparameters, data augmentation strategy, etc., and then retraining and evaluating again until satisfactory performance is obtained.
[0059] The handwritten text recognition model training method provided by the present invention constructs a second convolutional neural network based on the convolutional layer and the target fully connected layer of the first convolutional neural network, iteratively trains the second convolutional neural network with the preprocessed second text data as the training sample, and fine-tunes the learning rate of the second convolutional neural network. When the second convolutional neural network converges, a new handwritten text recognition model is obtained. Training the handwritten text recognition model through transfer learning can reduce the training time and computational resource consumption, help the model better cope with the changes and complexities in the target domain, and thus improve the generalization ability of the model.
[0060] Figure 3 is the second flow chart of the handwritten text recognition model training method provided by the present invention. In Figure 3 the illustrated embodiment, a handwritten text recognition model training method is implemented through the following steps: data collection and preprocessing, constructing a CNN (Convolutional Neural Networks) model combining LeNet-5 and AlexNet, data labeling and partitioning, model training, transfer learning, model evaluation and tuning, and model testing.
[0061] The following describes the handwritten text recognition method provided by the present invention. The handwritten text recognition method described below can be mutually corresponding and referred to the handwritten text recognition model training method described above.
[0062] Figure 4 is the flow chart of the handwritten text recognition method provided by the present invention. As Figure 4 shown, the handwritten text recognition method includes the following steps: Step 410, obtain the handwritten text to be recognized.
[0063] Step 420, process the handwritten text to be recognized based on the handwritten text recognition model to obtain the handwritten text recognition result; wherein, the handwritten text recognition model is trained through the handwritten text recognition model training method.
[0064] In step 410, the handwritten text to be recognized can be the test set in the first text data or the second text data, or other data in the publicly available text database or the latest written text data.
[0065] In step 420, the handwritten text recognition model is trained through the following steps: (1) Perform preprocessing on the first text data, including image scaling, grayscale conversion, and denoising, to obtain the preprocessed data set; wherein, the first text data includes at least one of the text data set in the publicly available database and the user's handwritten text data.
[0066] (2) Use the preprocessed dataset as the training samples to iteratively train the first convolutional neural network. Take the text writing features as the input features, and update the weights of the first convolutional neural network through the backpropagation algorithm. When the first convolutional neural network converges, obtain the handwritten text recognition model. Among them, the first convolutional neural network is designed based on the LeNet architecture and the AlexNet architecture.
[0067] In this embodiment, the training processes of each step are the same as those in the corresponding embodiments of the above step 110 - step 120, and will not be elaborated in this embodiment.
[0068] In some embodiments, after obtaining the handwritten text recognition model, the method for training the handwritten text recognition model further includes: constructing a second convolutional neural network based on the convolutional layer and the target fully connected layer of the first convolutional neural network; using the preprocessed second text data as the training samples to iteratively train the second convolutional neural network, and fine-tuning the learning rate of the second convolutional neural network. When the second convolutional neural network converges, obtain a new handwritten text recognition model.
[0069] In some embodiments, after obtaining the new handwritten text recognition model, the method for training the handwritten text recognition model further includes: using the preprocessed third sample as the test sample to calculate the target evaluation metrics corresponding to the new handwritten text recognition model, obtain the evaluation result, and optimize the new handwritten text recognition model according to the evaluation result. Among them, the target evaluation metrics include at least one of accuracy, recall, F1 value, and average edit distance.
[0070] The handwritten text recognition method provided by the present invention iteratively trains the first convolutional neural network with the preprocessed dataset as the training samples, and uses the handwritten text recognition model trained with the text writing features as the input features to recognize the handwritten text to be recognized, improving the accuracy of handwritten text recognition.
[0071] Next, the handwritten text recognition model training device provided by the present invention will be described. The handwritten text recognition model training device described below can be mutually corresponding and referred to the handwritten text recognition model training method described above.
[0072] Figure 5 is a schematic structural diagram of the handwritten text recognition model training device provided by the present invention, as Figure 5 shown, the handwritten text recognition model training device includes: an image preprocessing module 510 and a training module 520.
[0073] The image preprocessing module 510 is used to perform preprocessing of image scaling, grayscale conversion, and denoising on the first text data to obtain the preprocessed dataset. Among them, the first text data includes at least one of the text datasets in the publicly available database and the user's handwritten text data. A training module 520, configured to iteratively train a first convolutional neural network with the preprocessed data set as training samples, using the text writing features as input features, and updating the weights of the first convolutional neural network through the backpropagation algorithm, and obtaining a handwritten text recognition model when the first convolutional neural network converges; wherein, the first convolutional neural network is designed based on an LSTM network with an attention mechanism, a LeNet architecture, and an AlexNet architecture.
[0074] The handwritten text recognition model training device provided by the present invention preprocesses the first text data by image scaling, grayscaling, and denoising, and iteratively trains the first convolutional neural network with the preprocessed data set as training samples, using the text writing features as input features, and updating the weights of the first convolutional neural network through the backpropagation algorithm, and obtaining a handwritten text recognition model when the first convolutional neural network converges, improving the robustness of the handwritten text model recognition, and further improving the handwritten text recognition accuracy.
[0075] The handwritten text recognition device provided by the present invention will be described below. The handwritten text recognition device described below can be correspondingly referred to the handwritten text recognition method described above.
[0076] Figure 6 is a schematic structural diagram of the handwritten text recognition device provided by the present invention, as Figure 6 shown, the handwritten text recognition device includes: a handwritten text acquisition module 610 and a recognition module 620.
[0077] The handwritten text acquisition module 610 is configured to acquire the handwritten text to be recognized; The recognition module 620 is configured to process the handwritten text to be recognized based on the handwritten text recognition model to obtain a handwritten text recognition result; wherein, the handwritten text recognition model is trained by the handwritten text recognition model training method.
[0078] The handwritten text recognition device provided by the present invention recognizes the handwritten text to be recognized by using the handwritten text recognition model trained with the preprocessed data set as training samples and the text writing features as input features, improving the handwritten text recognition accuracy.
[0079] Figure 7 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 7As shown in the figure, the electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 complete communication with each other through the communication bus 740. The processor 710 may call the logical instructions in the memory 730 to execute the handwritten text recognition model training method, which includes: performing preprocessing of image scaling, grayscale conversion, and denoising on the first text data to obtain a preprocessed data set; where the first text data includes at least one of the text data set in the publicly available database and the user's handwritten text data; iteratively training the first convolutional neural network with the preprocessed data set as the training samples, using the text writing features as the input features, and updating the weights of the first convolutional neural network through the backpropagation algorithm, and obtaining a handwritten text recognition model when the first convolutional neural network converges; where the first convolutional neural network is designed through an LSTM network based on the attention mechanism, a LeNet architecture, and an AlexNet architecture.
[0080] Or execute the handwritten text recognition method, which includes: obtaining the handwritten text to be recognized; processing the handwritten text to be recognized based on the handwritten text recognition model to obtain a handwritten text recognition result; where the handwritten text recognition model is trained through the handwritten text recognition model training method.
[0081] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0082] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the handwritten text recognition model training method provided by each of the above methods. The method includes: performing preprocessing of image scaling, grayscale conversion, and denoising on the first text data to obtain a preprocessed data set; wherein the first text data includes at least one of the text data set in the publicly available database and the user's handwritten text data; using the preprocessed data set as training samples to iteratively train the first convolutional neural network, using the text writing features as input features, and updating the weights of the first convolutional neural network through the backpropagation algorithm, and obtaining a handwritten text recognition model when the first convolutional neural network converges; wherein the first convolutional neural network is designed through an LSTM network based on the attention mechanism, a LeNet architecture, and an AlexNet architecture.
[0083] Or execute a handwritten text recognition method, which includes: obtaining the handwritten text to be recognized; processing the handwritten text to be recognized based on the handwritten text recognition model to obtain a handwritten text recognition result; wherein the handwritten text recognition model is trained through the handwritten text recognition model training method.
[0084] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes the execution of the handwritten text recognition model training method provided by each of the above methods. The method includes: performing preprocessing of image scaling, grayscale conversion, and denoising on the first text data to obtain a preprocessed data set; wherein the first text data includes at least one of the text data set in the publicly available database and the user's handwritten text data; using the preprocessed data set as training samples to iteratively train the first convolutional neural network, using the text writing features as input features, and updating the weights of the first convolutional neural network through the backpropagation algorithm, and obtaining a handwritten text recognition model when the first convolutional neural network converges; wherein the first convolutional neural network is designed through an LSTM network based on the attention mechanism, a LeNet architecture, and an AlexNet architecture.
[0085] Or execute a handwritten text recognition method, which includes: obtaining the handwritten text to be recognized; processing the handwritten text to be recognized based on the handwritten text recognition model to obtain a handwritten text recognition result; wherein the handwritten text recognition model is trained through the handwritten text recognition model training method.
[0086] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.
Claims
1. A handwritten character recognition model training method, characterized in that: include: Performing image scaling, grayscale conversion and denoising preprocessing on the first text data to obtain a preprocessed data set; wherein the first text data includes at least one of a text data set in a public database and user handwritten text data; The first convolutional neural network is iteratively trained using the preprocessed data set as a training sample, using the text writing features as input features, and the weights of the first convolutional neural network are updated through a back-propagation algorithm, and a handwritten text recognition model is obtained when the first convolutional neural network converges; wherein the first convolutional neural network is designed using an LSTM network based on an attention mechanism, a LeNet architecture, and an AlexNet architecture.
2. The handwriting recognition model training method according to claim 1, characterized in that: After obtaining the handwritten character recognition model, the method further includes: Constructing a second convolutional neural network based on the convolutional layer of the first convolutional neural network and the target fully connected layer; The second convolutional neural network is iteratively trained using the preprocessed second text data as a training sample, and the learning rate of the second convolutional neural network is fine-tuned, so as to obtain a new handwritten text recognition model when the second convolutional neural network converges.
3. The handwriting recognition model training method according to claim 2, characterized in that: After obtaining the new handwritten character recognition model, the method further includes: The target evaluation index corresponding to the new handwritten text recognition model is calculated using the preprocessed third sample as a test sample to obtain an evaluation result, and the new handwritten text recognition model is tuned according to the evaluation result; wherein the target evaluation index includes at least one of accuracy, recall rate, F1 value and average edit distance.
4. The handwriting recognition model training method according to claim 1, characterized in that: The character writing features include at least one of stroke direction and trajectory, connectivity features, scale and proportion, local shape features and morphological features.
5. A handwritten text recognition method, characterized in that: include: Get the handwritten text to be recognized; The handwritten characters to be recognized are processed based on a handwritten character recognition model to obtain a handwritten character recognition result; wherein the handwritten character recognition model is trained by the handwritten character recognition model training method according to any one of claims 1 to 4.
6. A handwritten character recognition model training device, characterized in that: include: An image preprocessing module, used to perform image scaling, grayscale conversion and denoising preprocessing on the first text data to obtain a preprocessed data set; wherein the first text data includes at least one of a text data set in a public database and a user's handwritten text data; A training module is used to iteratively train the first convolutional neural network using the preprocessed data set as a training sample, using text writing features as input features, and updating the weights of the first convolutional neural network through a back-propagation algorithm, and obtaining a handwritten text recognition model when the first convolutional neural network converges; wherein the first convolutional neural network is designed using an LSTM network based on an attention mechanism, a LeNet architecture, and an AlexNet architecture.
7. A handwritten character recognition device, characterized in that: include: A handwritten text acquisition module is used to acquire the handwritten text to be recognized; A recognition module is used to process the handwritten characters to be recognized based on a handwritten character recognition model to obtain a handwritten character recognition result; wherein the handwritten character recognition model is trained by the handwritten character recognition model training method according to any one of claims 1 to 4.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.