Handwritten Chinese character handwriting identification system and identification method
Through the handwritten Chinese character handwriting identification system based on the SE_ResNet50 model, the problem of low accuracy in the recognition of handwritten Chinese character handwriting in the prior art is solved, and higher identification accuracy and security are achieved.
Patent Information
- Application Number
- CN202510119750.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-05-27
AI Technical Summary
In actual cases under the prior art, the accuracy of handwriting handwriting recognition in handwritten Chinese characters is low, resulting in safety risks.
A handwritten Chinese character handwriting identification system based on SE_ResNet50 model is adopted, including image acquisition, preprocessing, feature extraction, model training and optimization, and identification modules. Improve the discrimination accuracy through deep feature extraction and transfer learning.
It improves the accuracy and efficiency of handwriting handwriting identification, reduces safety risks, and can more accurately identify the handwriting characteristics of different writers.
Smart Images

Figure CN120048007A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of handwriting identification, and in particular, to a handwritten Chinese character handwriting identification system and an identification method based on the SE_ResNet50 model. Background Art
[0002] According to the handwriting identification technical specification GB / T 37239-2018, handwriting identification is a specialized document inspection technology that compares and examines the handwriting characteristics of the questioned document and the sample handwriting, and examines and identifies the writer of the questioned document or the identity with the sample handwriting. It has wide application value in the fields of information security, identity authentication, forensic identification, forensic science, finance, etc., and belongs to the research field of identity recognition like face recognition, fingerprint recognition, and iris recognition. Traditional handwriting identification methods mainly rely on manual vision and morphological analysis. The feature extraction is subjective, inefficient, and limited in accuracy, and there is no objective quantitative method. However, with the surge in the demand for handwriting identification, the handwriting identification mainly based on humans can no longer meet the growing demand for high-precision handwriting identification.
[0003] The rapid development of artificial intelligence and pattern recognition technologies has greatly promoted the challenging problem of handwriting identification. Among them, the emergence of deep learning technology provides a new solution for handwriting identification, which can automatically learn the complex features in handwriting and improve the identification accuracy and efficiency. However, the diversity, complexity, and individual differences of Chinese character handwriting, as well as the large randomness and lack of standardization in the writing of handwritten characters, pose great challenges to recognition, and the recognition accuracy in actual cases is relatively low. Therefore, a technical solution is needed that, based on a given handwriting material library, inputs a handwriting material, can find the differences between the handwriting materials of different writers, and feedback the sample closest to the input handwriting material to determine the identity information of the writer of the input material. Summary of the Invention
[0004] The present application provides a handwritten Chinese character handwriting identification system based on the SE_ResNet50 model to solve the defect of low recognition accuracy of handwritten Chinese character handwriting in actual cases in the prior art, which leads to security risks.
[0005] To achieve the above technical purpose, the present application proposes a handwritten Chinese character handwriting identification system based on the SE_ResNet50 model:
[0006] It includes an image acquisition module that acquires image data of handwritten Chinese characters through an acquisition device and stores it as a picture format file;
[0007] A preprocessing module that performs preprocessing operations such as grayscale conversion, noise reduction, binarization, normalization, and character segmentation on the acquired image;
[0008] A feature extraction module that performs deep feature extraction on the preprocessed image;
[0009] A model training and optimization module that trains a handwriting identification model using a large amount of handwritten Chinese character sample data;
[0010] An identification module that determines the attribution of the handwriting to be identified.
[0011] An identification method for a handwritten Chinese character handwriting identification system based on claim 1, comprising the following steps:
[0012] S1: Image acquisition is performed through an image acquisition module;
[0013] S2: The image is preprocessed;
[0014] S3: Feature extraction is performed;
[0015] S4: Training and optimization are performed through the model training and optimization module;
[0016] S5: Identification is performed through the identification module.
[0017] Preferably: The step S1 further includes: obtaining image data of handwritten Chinese characters, collecting through a scanner or a high-speed camera device, and the collected images include but are not limited to JPEG, JPG, and PNG formats.
[0018] Preferably: The step S2 further includes: performing preprocessing operations such as grayscale conversion, noise reduction, binarization, normalization, and character segmentation on the collected image. Grayscale conversion converts a color image into a grayscale image. For the collected color image, weighted average method is used for grayscale conversion, and the formula is: Gray = 0.299 * R + 0.587 * G + 0.114 * B, where R, G, and B are the pixel values of the red, green, and blue channels of the image respectively; Noise reduction uses a method combining Gabor filter and XGabor filter to remove noise interference in the image, and median filtering algorithm is used for noise reduction processing. The image is traversed with a 3x3 or 5x5 filtering window, and the pixel values within the window are sorted and the median value is taken as the new value of the central pixel; Binarization converts the image pixel values into 0 and 1 to highlight the handwriting outline; Normalization adjusts the image size to a unified size, and bilinear interpolation method is used to normalize the image to a fixed size to ensure that all images input into the model have the same size specification. For connected handwritten Chinese characters, a character segmentation algorithm based on connected regions is used to segment the Chinese characters into individual characters by analyzing the connectivity of pixels.
[0019] Preferably, the step S3 further includes: performing deep feature extraction on the preprocessed image based on the pre-trained SE_ResNet50 model. The input of the model is the preprocessed handwritten Chinese character image, and the output is a feature vector of a fixed length, which can represent the unique handwriting features of the handwritten Chinese characters. Subsequently, use CNN to extract the local features of the handwritten Chinese characters, and at the same time use RNN to extract the sequence features, and finally effectively extract the static and dynamic features of the handwritten Chinese characters.
[0020] Preferably, the step S4 further includes: training the handwriting identification model based on SE_ResNet50 using a large number of handwritten Chinese character sample data. The training samples include handwritten Chinese character images from different writers, and the corresponding writer identity information is labeled. The cross-entropy loss function is used as the objective function of the model, and the parameters of the model are continuously adjusted through the backpropagation algorithm, so that the model can accurately distinguish the handwriting features of different writers. During the training process, use the stochastic gradient descent and its variant optimization algorithms to accelerate the convergence speed of the model. Use the Adam optimization algorithm, set the initial learning rate to 0.001, the learning rate decay strategy is to decay to 0.9 times the original every 10 epochs, the number of iterations is set to 100 epochs, and the batch size is set to 32. After each epoch ends, calculate the evaluation indicators such as the accuracy rate and recall rate of the model on the validation set. When the accuracy rate of the validation set no longer improves for 5 consecutive epochs, stop training to prevent the model from overfitting, and save the model parameters at this time as the final model. At the same time, introduce the L1 and L2 regularization methods to constrain the model parameters and improve the generalization ability of the model; through the path signature feature and data augmentation technology, enhance the adaptability of the model to different writing styles.
[0021] Preferably, the step S5 further includes: after the handwritten Chinese character image to be identified is preprocessed and feature extracted, input it into the trained SE_ResNet50 model to obtain the feature vector of the image, calculate the similarity with the feature vectors of known writers, judge the matching degree between the handwriting to be identified and the handwriting of known writers, and determine the attribution of the handwriting to be identified.
[0022] Specifically 1. System architecture
[0023] This system mainly includes an image acquisition module, a preprocessing module, a feature extraction module, a model training and optimization module, and an identification module.
[0024] 1.1 Image acquisition module: responsible for obtaining the image data of handwritten Chinese characters, which can be collected through devices such as scanners and high-speed cameras. The collected images include but are not limited to common formats such as JPEG / JPG and PNG, and ensure that the images have sufficient resolution and clarity to fully retain the details such as the collocation ratio, stroke movement, and pen marks of the handwriting;
[0025] 1.2 Preprocessing Module: Perform preprocessing operations on the collected images, such as grayscale conversion, noise reduction, binarization, normalization, and character segmentation. Grayscale conversion transforms the color image into a grayscale image. For the collected color image, the weighted average method is used for grayscale conversion, and the formula is: Gray = 0.299 * R + 0.587 * G + 0.114 * B, where R, G, and B are the pixel values of the red, green, and blue channels of the image respectively, ultimately reducing the data volume. Noise reduction uses a method combining Gabor filters and XGabor filters to remove noise interference in the image, and the median filtering algorithm is used for noise reduction processing. The image is traversed with a 3x3 or 5x5 filtering window, and the pixel values within the window are sorted and the middle value is taken as the new value of the central pixel, effectively removing noise interference such as salt-and-pepper noise. Binarization converts the image pixel values into 0 and 1, highlighting the handwriting outline. Normalization adjusts the image size to a unified size, and the bilinear interpolation method is used to normalize the image to a fixed size, such as 64x64 pixels, ensuring that all images input into the model have the same size specification. For connected handwritten Chinese characters, a character segmentation algorithm based on connected regions is adopted. By analyzing the connectivity of pixels, the Chinese characters are segmented into individual characters, providing accurate character data for subsequent model training and identification to meet the input requirements of the SE_ResNet50 model;
[0026] 1.3 Feature Extraction Module: Perform deep feature extraction on the preprocessed images based on the pre-trained SE_ResNet50 model. The SE_ResNet50 model effectively solves the problems of gradient disappearance and gradient explosion in deep neural networks through the residual structure and can learn the high-level semantic features of handwritten Chinese character images. The input of the model is the preprocessed handwritten Chinese character image, and the output is a feature vector of a fixed length. This feature vector can represent the unique handwriting features of handwritten Chinese characters. Subsequently, CNN is used to extract local features of handwritten Chinese characters, such as strokes and structures, and at the same time, RNN is used to extract sequence features, such as stroke order, ultimately effectively extracting the static and dynamic features of handwritten Chinese characters;
[0027] 1.4 Model Training and Optimization Module: The handwriting identification model based on SE_ResNet50 is trained using a large number of handwritten Chinese character sample data (such as the CASIA-HWDB handwritten Chinese character dataset of the Chinese Academy of Sciences). The training samples include handwritten Chinese character images from different writers, and the corresponding writer identity information is labeled. The cross-entropy loss function is used as the objective function of the model, and the parameters of the model are continuously adjusted through the backpropagation algorithm, enabling the model to accurately distinguish the handwriting features of different writers. During the training process, stochastic gradient descent (SGD) and its variant optimization algorithms (such as Adagrad, Adadelta, Adam, etc.) are used to accelerate the convergence speed of the model. The Adam optimization algorithm is used, with an initial learning rate set to 0.001, a learning rate decay strategy of decaying to 0.9 times the original value every 10 epochs, the number of iterations set to 100 epochs, and the batch size set to 32. After each epoch ends, evaluation metrics such as the accuracy and recall rate of the model are calculated on the validation set. When the validation set accuracy does not improve for 5 consecutive epochs, the training is stopped to prevent overfitting, and the model parameters at this time are saved as the final model. At the same time, L1 and L2 regularization methods are introduced to constrain the model parameters and improve the generalization ability of the model. Through path-signature features and data augmentation techniques (DropStroke), the adaptability of the model to different writing styles is enhanced.
[0028] 1.5 Identification Module: After preprocessing and feature extraction of the handwritten Chinese character image to be identified, it is input into the trained SE_ResNet50 model to obtain the feature vector of the image. Then, by calculating the similarity (such as cosine similarity, Euclidean distance, etc.) with the feature vectors of known writers, the matching degree between the handwriting to be identified and the handwriting of known writers is judged, thereby determining the attribution of the handwriting to be identified.
[0029] Compared with the prior art, the beneficial effects of the present invention are:
[0030] 2.1 Improvement and Optimization of the SE_ResNet50 Model
[0031] As the number of network layers increases, the network structure becomes increasingly complex. If the number of network layers is too large, the problem of gradient disappearance will occur. To solve this problem, scholars proposed the residual structure network ResNet50. To increase the dynamic adaptability to the input and improve the feature discrimination performance, an SE attention mechanism module was introduced to construct the SE_ResNet50 network structure. The SE module is a lightweight channel attention module that can form computational units from arbitrary transformations, making it convenient to load into various existing network model frameworks to improve the model performance. The SE module is divided into two steps: compression (Squeeze) and excitation (Excitation). In the Squeeze step, after passing through the global pooling layer and the fully connected layer, it enters the Excitation step. In the Excitation step, the sigmoid function is used to obtain the weighted feature map. Taking ResNet50 as the main structure, the ResNet50 model can adaptively adjust the fully connected layer of the SE_ResNet50 model. According to the task characteristics of handwritten Chinese character handwriting identification, the number of neurons in the output layer is set to correspond to the number of writers in the training samples to achieve the multi-classification task.
[0032] During the training process of the model, the transfer learning method is adopted. First, the SE_ResNet50 model is pre-trained using a large-scale general image dataset (such as ImageNet) to learn the general image feature representation. Then, fine-tuning training is carried out on the handwritten Chinese character handwriting identification dataset to enable the model to better adapt to the specific task requirements of handwritten Chinese character handwriting identification, accelerate the convergence speed of the model, and improve the model performance.
[0033] 2.2 Data augmentation strategy
[0034] To expand the scale of the training dataset and improve the generalization ability of the model, a variety of data augmentation methods are adopted. These include but are not limited to operations such as random rotation, random cropping, random horizontal flipping, random vertical flipping, and adding noise. After performing these data augmentation operations on the original handwritten Chinese character images, new training samples are generated, thereby increasing the model's learning ability for handwritten Chinese character handwriting features in different poses and different noise environments.
[0035] 2.3 Model evaluation metrics
[0036] Metrics such as accuracy, recall, and F1 value are used to evaluate the performance of the handwriting identification model. During the training process, these metrics are calculated regularly on the validation set to monitor the training effect of the model, and the hyperparameters and training strategies of the model are adjusted according to the evaluation results to ensure good performance of the model on the test set. Description of the drawings
[0037] Figure 1 This is the network structure diagram of the SE_ResNet50 of the present invention;
[0038] Figure 2 This is the architecture diagram of the handwritten Chinese character handwriting identification system of the present invention; Specific implementation manners
[0039] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0040] Embodiment 1:
[0041] Please refer to Figure 1-2 ,
[0042] This application proposes a handwritten Chinese character handwriting identification system based on the SE_ResNet50 model. The method includes:
[0043] This system mainly includes an image acquisition module, a preprocessing module, a feature extraction module, a model training and optimization module, and an identification module.
[0044] 1.1 Image acquisition module: Responsible for obtaining the image data of handwritten Chinese characters, which can be collected by devices such as scanners and high-speed cameras. The collected images include but are not limited to common formats such as JPEG / JPG and PNG, and ensure that the images have sufficient resolution and clarity to fully retain the details such as the collocation ratio, stroke movement, and pen marks of the handwriting;
[0045] 1.2 Preprocessing Module: Perform preprocessing operations on the collected images, such as grayscale conversion, noise reduction, binarization, normalization, and character segmentation. Grayscale conversion converts a color image into a grayscale image. For the collected color image, the weighted average method is used for grayscale conversion, and the formula is: Gray = 0.299 * R + 0.587 * G + 0.114 * B, where R, G, and B are the pixel values of the red, green, and blue channels of the image respectively, ultimately reducing the amount of data. Noise reduction uses a combination of Gabor filters and XGabor filters to remove noise interference in the image, and the median filtering algorithm is used for noise reduction processing. The image is traversed with a 3x3 or 5x5 filtering window, and the pixel values within the window are sorted and the median value is taken as the new value of the central pixel, effectively removing noise interference such as salt-and-pepper noise. Binarization converts the image pixel values into 0 and 1, highlighting the handwriting contour. Normalization adjusts the image size to a unified size, and the bilinear interpolation method is used to normalize the image to a fixed size, such as 64x64 pixels, ensuring that all images input into the model have the same size specification. For connected handwritten Chinese characters, a character segmentation algorithm based on connected regions is adopted. By analyzing the connectivity of pixels, the Chinese characters are segmented into individual characters, providing accurate character data for subsequent model training and identification to meet the input requirements of the SE_ResNet50 model;
[0046] 1.3 Feature Extraction Module: Perform deep feature extraction on the preprocessed images based on the pre-trained SE_ResNet50 model. The SE_ResNet50 model effectively solves the problems of gradient disappearance and gradient explosion in deep neural networks through the residual structure and can learn the high-level semantic features of handwritten Chinese character images. The input of the model is the preprocessed handwritten Chinese character image, and the output is a feature vector of a fixed length. This feature vector can represent the unique handwriting features of handwritten Chinese characters. Subsequently, CNN is used to extract local features of handwritten Chinese characters, such as strokes and structures, and at the same time, RNN is used to extract sequence features, such as stroke order, ultimately effectively extracting the static and dynamic features of handwritten Chinese characters;
[0047] 1.4 Model Training and Optimization Module: The handwriting discrimination model based on SE_ResNet50 is trained using a large number of handwritten Chinese character sample data (such as the CASIA-HWDB handwritten Chinese character dataset of the Chinese Academy of Sciences). The training samples include handwritten Chinese character images from different writers, and the corresponding writer identity information is labeled. The cross-entropy loss function is used as the objective function of the model, and the parameters of the model are continuously adjusted through the backpropagation algorithm, so that the model can accurately distinguish the handwriting features of different writers. During the training process, stochastic gradient descent (SGD) and its variant optimization algorithms (such as Adagrad, Adadelta, Adam, etc.) are used to accelerate the convergence speed of the model. The Adam optimization algorithm is used, the initial learning rate is set to 0.001, the learning rate decay strategy is to decay to 0.9 times the original every 10 epochs, the number of iterations is set to 100 epochs, and the batch size is set to 32. After each epoch ends, evaluation metrics such as the accuracy and recall rate of the model are calculated on the validation set. When the validation set accuracy does not improve for 5 consecutive epochs, the training is stopped to prevent the model from overfitting, and the model parameters at this time are saved as the final model. At the same time, L1 and L2 regularization methods are introduced to constrain the model parameters and improve the generalization ability of the model. Through path-signature and data augmentation technology (DropStroke), the adaptability of the model to different writing styles is enhanced.
[0048] 1.5 Discrimination Module: After the handwritten Chinese character image to be discriminated is preprocessed and feature extracted, it is input into the trained SE_ResNet50 model to obtain the feature vector of the image. Then, by calculating the similarity (such as cosine similarity, Euclidean distance, etc.) with the feature vectors of known writers, the matching degree between the handwriting to be discriminated and the handwriting of known writers is judged, so as to determine the attribution of the handwriting to be discriminated.
[0049] 2. Key Technology Implementation
[0050] 2.1 Improvement and Optimization of SE_ResNet50 Model
[0051] As the number of network layers increases, the network structure becomes increasingly complex. If the number of network layers is too large, the problem of gradient disappearance will occur. To solve this problem, scholars proposed the residual structure network ResNet50. To increase the dynamic adaptability to the input and improve the feature discrimination performance, the SE attention mechanism module was introduced to construct the SE_ResNet50 network structure. The SE module is a lightweight channel attention module that can form computational units from arbitrary transformations and is convenient to load into various existing network model frameworks to improve the model performance. The SE module is divided into two steps: compression (Squeeze) and excitation (Excitation). In the Squeeze step, after passing through the global pooling layer and the fully connected layer, it enters the Excitation step. In the Excitation step, the sigmoid function is used to obtain the weighted feature map. Taking ResNet50 as the main structure, the ResNet50 model can adaptively adjust the fully connected layer of the SE_ResNet50 model. According to the task characteristics of handwritten Chinese character handwriting identification, the number of neurons in the output layer is set to correspond to the number of writers in the training samples to achieve the multi-classification task.
[0052] During the training process of the model, the transfer learning method is adopted. First, the SE_ResNet50 model is pre-trained using a large-scale general image dataset (such as ImageNet) to learn the general image feature representation. Then, fine-tuning training is performed on the handwritten Chinese character handwriting identification dataset to enable the model to better adapt to the specific task requirements of handwritten Chinese character handwriting identification, accelerate the convergence speed of the model, and improve the model performance.
[0053] 2.2 Data augmentation strategy
[0054] To expand the scale of the training dataset and improve the generalization ability of the model, various data augmentation methods are adopted. These include but are not limited to operations such as random rotation, random cropping, random horizontal flipping, random vertical flipping, and adding noise. After performing these data augmentation operations on the original handwritten Chinese character images, new training samples are generated, thereby increasing the model's learning ability for handwritten Chinese character handwriting features in different poses and different noise environments.
[0055] 2.3 Model evaluation metrics
[0056] Metrics such as accuracy, recall, and F1 value are used to evaluate the performance of the handwriting identification model. During the training process, these metrics are calculated on the validation set regularly to monitor the training effect of the model, and the hyperparameters and training strategies of the model are adjusted according to the evaluation results to ensure good performance of the model on the test set.
[0057] 1. Data preparation
[0058] Collect a large number of handwritten Chinese character samples from different writers, covering various factors such as different writing styles, fonts, writing tools, and paper backgrounds, to ensure the diversity and representativeness of the samples. Divide the collected samples into a training set, a validation set, and a test set according to a certain proportion. For example, 70% is used for training, 15% for validation, and 15% for testing.
[0059] Annotate each handwritten Chinese character sample and record the corresponding writer identity information for supervised learning during the model training process.
[0060] 2. System Construction and Training
[0061] Build a handwritten Chinese character handwriting identification system based on the SE_ResNet50 model, including the above-mentioned various modules. Load the pre-trained SE_ResNet50 model weights, and initialize and adjust the fully connected layer of the model to make it suitable for the handwriting identification task.
[0062] Input the training set data into the system. After preprocessing, use the improved SE_ResNet50 model for feature extraction and model training. During the training process, according to the set loss function and optimization algorithm, continuously adjust the parameters of the model to gradually reduce the loss of the model on the training set and gradually improve the evaluation metrics on the validation set. Through multiple iterative trainings until the model reaches a convergent state, that is, the performance on the validation set no longer improves or reaches the preset number of training epochs.
[0063] 3. System Testing and Optimization
[0064] Use the test set data to perform performance testing on the trained handwriting identification system, calculate evaluation metrics such as accuracy, recall rate, and F1 value, and analyze the identification effect and existing problems of the system in different scenarios.
[0065] According to the test results, further optimize and improve the system. For example, if it is found that the handwriting identification accuracy of some writers is low, more sample data of these writers can be collected specifically for supplementary training; if the model has an overfitting phenomenon, methods such as adjusting the regularization parameters, increasing the intensity of data augmentation, or adopting a more complex model structure can be used to improve the generalization ability of the model.
[0066] Through the implementation of the above technical solutions, the handwritten Chinese character handwriting identification system based on the SE_ResNet50 model of the present invention can effectively and accurately identify handwritten Chinese character handwriting, has high accuracy, stability, and reliability, and can be widely applied to various actual scenarios requiring handwriting identification.
[0067] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A handwritten Chinese character identification system, characterized in that: It includes an image acquisition module, which acquires image data of handwritten Chinese characters through an acquisition device and stores it as a picture format file; The preprocessing module performs grayscale conversion, noise reduction, binarization, normalization, and character segmentation preprocessing operations on the collected images; Feature extraction module, which performs deep feature extraction on the preprocessed image; Model training and optimization module, which uses a large amount of handwritten Chinese character sample data to train the handwriting identification model; The identification module determines the ownership of the handwriting to be identified.
2. A method for identifying handwritten Chinese characters based on the handwriting identification system of claim 1, characterized in that: The steps include: S1: image acquisition through the image acquisition module; S2: preprocess the image; S3: perform feature extraction; S4: Training and optimization through the model training and optimization module; S5: Authentication is performed through an authentication module.
3. The identification method of the handwritten Chinese character identification system according to claim 2, characterized in that: The step S1 also includes: acquiring image data of handwritten Chinese characters by means of a scanner or a high-definition camera, wherein the acquired images include but are not limited to JPEG, JPG, and PNG formats.
4. The identification method of the handwritten Chinese character identification system according to claim 3, characterized in that: The step S2 also includes: graying, denoising, binarization, normalization, and character segmentation preprocessing operations on the collected image. Graying converts the color image into a grayscale image. For the collected color image, the weighted average method is used for graying. The formula is: Gray = 0.299*R + 0.587*G + 0.114*B, where R, G, and B are the red, green, and blue channel pixel values of the image respectively; denoising uses a method combining Gabor filter and XGabor filter to remove noise interference in the image. The median filter algorithm is used for noise reduction. A 3x3 or 5x5 filter window is used to traverse the image. The pixel values in the window are sorted and the middle value is taken as the new value of the central pixel. Binarization converts the image pixel values into 0 and 1 to highlight the handwriting outline. Normalization resizes the image to a uniform size and uses bilinear interpolation to normalize the image to a fixed size to ensure that all images input to the model have the same size specifications. For handwritten Chinese characters with connected strokes, a character segmentation algorithm based on connected regions is used to segment the Chinese characters into individual characters by analyzing the connectivity of pixels.
5. The identification method of the handwritten Chinese character identification system according to claim 4, characterized in that: The step S3 also includes: performing deep feature extraction on the preprocessed image based on the pretrained SE_ResNet50 model, the input of the model is the preprocessed handwritten Chinese character image, and the output is a feature vector of fixed length, which can characterize the unique handwriting features of the handwritten Chinese characters, and then using CNN to extract local features of the handwritten Chinese characters, and using RNN to extract sequence features, and finally effectively extracting the static and dynamic features of the handwritten Chinese characters.
6. The identification method of the handwritten Chinese character identification system according to claim 5, characterized in that: The step S4 also includes: using a large amount of handwritten Chinese character sample data to train the handwriting identification model based on SE_ResNet50, the training samples include handwritten Chinese character images from different writers, and the corresponding writer identity information is annotated, a cross entropy loss function is used as the objective function of the model, and the parameters of the model are continuously adjusted through the back propagation algorithm so that the model can accurately distinguish the handwriting features of different writers. During the training process, the random gradient descent and its variant optimization algorithm are used to accelerate the convergence speed of the model, and the Adam optimization algorithm is used, the initial learning rate is set to 0.001, and the learning rate decay The strategy is to decay to 0.9 times of the original value every 10 epochs, the number of iterations is set to 100 epochs, and the batch size is set to 32. After each epoch, the accuracy, recall and other evaluation indicators of the model are calculated on the validation set. When the accuracy of the validation set no longer improves after 5 consecutive epochs, the training is stopped to prevent the model from overfitting, and the model parameters at this time are saved as the final model. At the same time, L1 and L2 regularization methods are introduced to constrain the model parameters to improve the generalization ability of the model; through path signature features and data enhancement technology, the adaptability of the model to different writing styles is enhanced.
7. The identification method of the handwritten Chinese character identification system according to claim 6, characterized in that: The step S5 also includes: after preprocessing and feature extraction of the handwritten Chinese character image to be identified, input it into the trained SE_ResNet50 model to obtain the feature vector of the image, and by calculating the similarity with the feature vector of a known writer, determine the degree of matching between the handwriting to be identified and the handwriting of the known writer, and determine the ownership of the handwriting to be identified.