Handwriting image identification method and system based on artificial intelligence
By generating a hybrid feature dataset and optimizing the training using the Triplet network model, the problem of insufficient accuracy in handwriting identification in existing technologies is solved, achieving efficient and accurate handwriting image identification.
Patent Information
- Application Number
- CN202511040784.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies rely on human experience and limited feature analysis, resulting in insufficient accuracy and reliability of handwriting identification results. They are particularly ineffective when dealing with complex writing styles and environmental interference, and cannot meet the needs of large-scale identification.
An artificial intelligence-based approach is used to collect handwriting images of both known and unidentified samples, generate a hybrid feature dataset, combine traditional handwriting features and deep learning handwriting features, and jointly optimize and train a Triplet network model to generate discriminative feature vectors. This results in a handwriting image identification report that includes feature matching scores and author attribution probabilities.
It improves the accuracy and reliability of handwriting identification, enables efficient handwriting image identification, and provides objective, accurate, and detailed identification evidence.
Smart Images

Figure CN120853183A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image analysis technology, specifically to a handwriting image identification method and system based on artificial intelligence. Background Technology
[0002] In the field of handwriting image identification, traditional handwriting identification methods mainly rely on human experience and limited feature analysis, judging the consistency of handwriting by comparing simple features such as the basic shape and proportion of strokes. With the development of computer technology, some image processing and machine learning-based methods have begun to be applied to handwriting identification. These methods usually extract some features of the handwriting, such as the length and angle of the strokes, and then use classifiers for classification and identification.
[0003] However, the aforementioned traditional methods rely too heavily on human experience, are highly subjective, and make it difficult to guarantee the accuracy and reliability of the identification results. They are also inefficient and cannot meet the needs of large-scale identification. Furthermore, existing computer-aided identification methods extract only single features and lack comprehensive and in-depth analysis of handwriting, making it difficult to accurately distinguish similar handwriting. In particular, the identification results are poor when faced with complex writing styles and environmental interference.
[0004] As a result, existing technologies are unable to fully extract the potential information from handwriting, reducing the accuracy and reliability of identification. Summary of the Invention
[0005] This invention provides a handwriting image identification method and system based on artificial intelligence.
[0006] In a first aspect, embodiments of the present invention provide an artificial intelligence-based handwriting image identification method, applied to a handwriting image identification system, the method comprising: Collect handwriting images to be identified and known sample handwriting images, and generate a target handwriting image set based on the handwriting images to be identified and the known sample handwriting images; The handwriting images in the target handwriting image set are sequentially subjected to stroke region segmentation, stroke feature extraction and texture pattern analysis to generate a hybrid feature dataset containing a traditional handwriting feature set and a deep learning handwriting feature set. The traditional handwriting feature set includes stroke tilt angle, stroke frequency and turning curvature features, while the deep learning handwriting feature set includes high-level semantic features output by the convolutional layer and temporal dependency features extracted by the recurrent layer. The hybrid feature dataset is input into the Triplet network model, and the traditional handwriting feature set and the deep learning handwriting feature set are jointly optimized and trained through the triplet loss function to generate the first discriminative feature vector of the handwriting image to be identified and the second discriminative feature vector of the known sample handwriting image. The triplet loss function realizes the aggregation of similar handwriting features and the separation of dissimilar handwriting features through the feature distance constraints of anchor samples, positive samples and negative samples. Based on the first and second discriminative feature vectors, a handwriting image identification report is generated, which includes feature matching scores and author attribution probabilities.
[0007] Secondly, embodiments of the present invention provide a handwriting image identification system, comprising: processor; Storage device, on which computer programs are stored, When the computer program is executed by the processor, the processor implements any of the described artificial intelligence-based handwriting image identification methods.
[0008] This invention provides a readable storage medium storing a program or instructions, which, when executed by a processor, implement the steps of the artificial intelligence-based handwriting image identification method.
[0009] This invention achieves efficient and accurate handwriting image identification. First, it collects handwriting images to be identified and known sample handwriting images to generate a target handwriting image set, avoiding data confusion and missing information, and ensuring the integrity and representativeness of the samples. Then, it sequentially performs stroke region segmentation, stroke feature extraction, and texture pattern analysis on the target handwriting image set. The resulting hybrid feature dataset integrates traditional handwriting feature sets and deep learning handwriting feature sets. Traditional handwriting feature sets, such as stroke tilt angle, continuous stroke frequency, and turning curvature features, can characterize the morphological features of handwriting at a microscopic level. Meanwhile, the high-level semantic features output by convolutional layers and the temporal dependency features extracted by recurrent layers in the deep learning handwriting feature set mine abstract information and writing order patterns from macroscopic and temporal dimensions. The combination of these two greatly enriches the expression of handwriting features and improves the comprehensiveness and depth of the features. The hybrid feature dataset is input into the Triplet network model, and the traditional and deep learning handwriting feature sets are jointly optimized and trained using the triplet loss function. By constraining the feature distances of anchor samples, positive samples, and negative samples, effective aggregation of similar handwriting features and clear separation of dissimilar handwriting features are achieved, resulting in stronger discriminative first and second discriminative feature vectors. The handwriting image identification report generated based on these two discriminative feature vectors, which includes feature matching scores and author attribution probabilities, provides objective, accurate, and detailed evidence for handwriting identification, improving the accuracy, reliability, and efficiency of handwriting identification. Attached Figure Description
[0010] Figure 1 This is a flowchart of an artificial intelligence-based handwriting image identification method provided in an embodiment of the present invention.
[0011] Figure 2 This is a schematic diagram of the basic structure of a handwriting image identification system provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0013] See Figure 1 As shown, this figure is a flowchart of an artificial intelligence-based handwriting image identification method provided by an embodiment of the present invention. This method can be applied to handwriting image identification systems. Figure 1 As shown, the method includes steps 110-140.
[0014] Step 110: Collect handwriting images to be identified and known sample handwriting images, and generate a target handwriting image set based on the handwriting images to be identified and the known sample handwriting images.
[0015] In this embodiment of the invention, in a handwriting image identification scenario, the handwriting image to be identified is the handwriting image whose authorship needs to be determined, while the known sample handwriting images are handwriting images of known authors, used for comparison with the handwriting image to be identified. After collecting these images, they are processed accordingly to ultimately generate a target handwriting image set. This set contains the sorted and classified handwriting images to be identified and the known sample handwriting images. For example, in a document authenticity identification case, the signature handwriting on a document needs to be identified; this signature handwriting image is the handwriting image to be identified. Simultaneously, signature handwriting images of authors who may be involved in the document at different times and under different writing conditions are collected as known sample handwriting images.
[0016] Optionally, the step of collecting the handwriting image to be identified and the known sample handwriting images, and generating a target handwriting image set based on the handwriting image to be identified and the known sample handwriting images, includes: Step 111: Receive handwriting images to be identified and known sample handwriting images acquired through multiple acquisition devices. The handwriting images to be identified contain handwriting patterns formed by different writing media and writing tools, and the known sample handwriting images contain handwriting sample patterns formed by multiple known authors under different writing conditions.
[0017] In this embodiment of the invention, various acquisition devices can include scanners, high-definition cameras, etc. The handwriting image to be identified may be formed on different writing media such as paper and electronic screens, using different writing tools such as pens, pencils, and electronic pens. Known sample handwriting images are handwriting sample patterns left by multiple known authors under different writing environments, writing times, and other conditions. For example, in the aforementioned document signature identification case, the signature to be identified may be written with a pen on ordinary paper, while known sample handwriting images may include the author's signature in pencil on a notebook, signature in electronic documents using an electronic pen, etc. Acquiring these images through different acquisition devices can more comprehensively cover various possible handwriting situations.
[0018] Step 112: Perform image quality assessment processing on the handwriting image to be identified and the known sample handwriting image. Filter out valid handwriting images that meet the preset quality standards through edge clarity detection and contrast analysis, and remove invalid handwriting images with blurred areas or incomplete strokes.
[0019] In this embodiment of the invention, image quality assessment processing is used to ensure the accuracy of subsequent feature extraction and analysis. Edge sharpness detection determines whether the edges of the handwriting image are clear, while contrast analysis determines the distinction between the handwriting and the background in the image. Preset quality standards are set according to actual needs; only images that meet these standards can be used as valid handwriting images for subsequent processing. For example, in the aforementioned case, if a known sample handwriting image contains blurred areas, making some strokes unclear or containing incomplete strokes, then this image will be determined as an invalid handwriting image and removed, thereby avoiding the impact of low-quality images on the final identification result.
[0020] Step 113: Based on the screened handwriting images to be identified and the known sample handwriting images, generate optimized handwriting images with the same resolution and color mode.
[0021] In this embodiment of the invention, images acquired by different acquisition devices may have different resolutions and color modes, which can complicate subsequent feature extraction and comparison. Therefore, it is necessary to unify the resolution and color mode of the screened handwriting images to be identified and the known sample handwriting images to generate optimized handwriting images. For example, the resolution of all images can be unified to a fixed value, and the color mode can be unified to grayscale. In the aforementioned document signature identification case, converting all the signature images to be identified and the known sample signature images to the same resolution and grayscale color mode can ensure consistency in feature extraction and improve the accuracy of identification.
[0022] Step 114: Perform classification and labeling processing based on the source type and author identification information of the optimized handwriting image, and establish a classification index table containing the category to be identified and the known sample categories.
[0023] In this embodiment of the invention, the source types of optimized handwriting images are divided into two categories: those to be identified and known samples. Author identification information is key to distinguishing different authors. By classifying and labeling the images and establishing a classification index table, subsequent image management and retrieval can be facilitated. For example, in the aforementioned case, the signature image to be identified is marked as the category to be identified, and the known sample signature images of different authors are marked with their corresponding author identifiers. The classification index table is similar to a directory, recording the category to which each image belongs and the corresponding author information, facilitating quick location and use.
[0024] Step 115: Based on the classification index table, arrange and combine the optimized handwriting images according to a preset category order to generate a target handwriting image set containing category labels and image identifiers. Each image unit in the target handwriting image set is associated with corresponding classification index information.
[0025] In this embodiment of the invention, optimized handwriting images are arranged and combined according to a preset category order based on a classification index table. The preset category order may be: first, the category to be identified, then the known sample categories arranged according to the author identifier order. Each image unit in the generated target handwriting image set has a corresponding category label and image identifier, and is associated with classification index information. For example, in the aforementioned document signature identification case, the signature image to be identified is placed first, and then the known sample signature images of different authors are arranged in alphabetical order of the authors. Each image has a unique image identifier, which is also associated with its category and author information. Thus, each image can be clearly identified and used in subsequent processing.
[0026] Step 120: Perform stroke region segmentation, stroke feature extraction, and texture pattern analysis operations sequentially on the handwriting images in the target handwriting image set to generate a hybrid feature dataset containing a traditional handwriting feature set and a deep learning handwriting feature set; the traditional handwriting feature set includes stroke tilt angle, stroke frequency, and turning curvature features, and the deep learning handwriting feature set includes high-level semantic features output by convolutional layers and temporal dependency features extracted by recurrent layers.
[0027] In this embodiment of the invention, a series of feature extraction operations are performed on each handwriting image in the target handwriting image set. First, stroke region segmentation is performed, dividing the handwriting into independent stroke regions to provide a foundation for subsequent feature extraction. Then, stroke features are extracted, which reflect the writer's writing habits and style. Next, texture pattern analysis is performed, which reveals the microscopic features of the handwriting. Through these operations, a traditional handwriting feature set and a deep learning handwriting feature set are generated, and finally, they are merged into a hybrid feature dataset. For example, in the aforementioned document signature authentication case, each signature image in the target handwriting image set is processed. Stroke region segmentation separates each stroke in the signature, facilitating the analysis of each stroke's features. Stroke feature extraction reveals the stroke characteristics at the beginning and end of strokes in the signature, while texture pattern analysis reveals the texture patterns of the strokes in the signature. The stroke tilt angle, connecting stroke frequency, and turning curvature features in the traditional handwriting feature set can describe the form of the signature from different aspects, while the high-level semantic features and temporal dependency features in the deep learning handwriting feature set can uncover deeper features of the signature.
[0028] In step 120, stroke region segmentation is performed on the handwriting images in the target handwriting image set, including: Step 121: Perform grayscale processing on the handwriting images in the target handwriting image set, convert the color handwriting images into single-channel grayscale images, and retain the grayscale gradient information of the handwriting area; perform binarization processing on the single-channel grayscale image through an adaptive threshold segmentation algorithm to separate the handwriting foreground area and the background area, and generate a binarized handwriting image; perform stroke area segmentation processing based on the contour features of the binarized handwriting image, obtain the central skeleton line of the handwriting through a skeleton extraction algorithm, and divide the complete handwriting into multiple independent stroke area units according to the branch points and intersection points of the central skeleton line; perform contour tracking processing on each stroke area unit, extract the boundary pixel sequence of the stroke area unit, and determine the trend feature and length feature of the stroke area unit in combination with the coordinate change trend of the boundary pixel sequence; use the trend feature and the length feature as the stroke area segmentation result, and the stroke area segmentation result is used to provide regional positioning information for pen tip feature extraction.
[0029] In the embodiment of the present invention, first, perform grayscale processing on the color handwriting images in the target handwriting image set and convert them into single-channel grayscale images, which can simplify the image processing process and retain the grayscale gradient information of the handwriting area at the same time. The grayscale gradient information can reflect the edge and contour features of the handwriting. Then, use an adaptive threshold segmentation algorithm to perform binarization processing on the single-channel grayscale image, divide the image into the handwriting foreground area and the background area, and generate a binarized handwriting image. Next, based on the contour features of the binarized handwriting image, obtain the central skeleton line of the handwriting through a skeleton extraction algorithm. The central skeleton line can represent the core structure of the handwriting. According to the branch points and intersection points of the central skeleton line, divide the complete handwriting into multiple independent stroke area units. Perform contour tracking processing on each stroke area unit, extract its boundary pixel sequence, and determine the trend feature and length feature of the stroke area unit by analyzing the coordinate change trend of the boundary pixel sequence.
[0030] For example, in the above document signature authentication case, for the signature images in the target handwriting image set, the grayscale distribution of the signature can be seen more clearly after grayscale processing. The binarization processing separates the signature from the background, facilitating subsequent stroke segmentation. The skeleton extraction algorithm can obtain the central skeleton of the signature, and divide the signature into multiple independent strokes such as "one" and "slash" according to the branches and intersections of the skeleton. Perform contour tracking on each stroke area unit to determine its trend, such as up, down, left, or right, and the length of the stroke. These trend features and length features are used as the stroke area segmentation result, providing accurate regional positioning information for pen tip feature extraction.
[0031] In step 120, perform pen tip feature extraction on the handwriting images in the target handwriting image set, including: Step 122: Based on the stroke region units in the stroke region segmentation result, locate the start and end endpoints of each stroke region unit to determine the extraction region of the stroke tip feature; perform pixel grayscale value analysis on the extraction region of the stroke tip feature, calculate the grayscale change rate in the preset neighborhood around the start and end endpoints, and determine the sharpness feature of the stroke tip; perform statistical analysis on the edge direction distribution of the stroke tip region using the directional gradient histogram algorithm to generate the directional feature vector of the stroke tip region, which is used to describe the tilt direction and diffusion range of the stroke tip; combine the sharpness feature and the directional feature vector to construct a stroke tip feature descriptor, which includes the morphological features and grayscale distribution features of the stroke tip; summarize the stroke tip feature descriptors of all stroke region units to generate a stroke tip feature set, which is used as a component of the traditional handwriting feature set.
[0032] In this embodiment of the invention, based on the stroke region segmentation results, the start and end endpoints of each stroke region unit are located. The vicinity of these endpoints is the extraction area for stroke features. Pixel grayscale value analysis is performed on the extracted area of stroke features, calculating the grayscale change rate within a preset neighborhood around the start and end endpoints. The greater the grayscale change rate, the sharper the stroke, thus determining the sharpness feature of the stroke. The Histogram of Oriented Gradients (HGP) algorithm can be used to statistically analyze the edge direction distribution of the stroke region, generating a directional feature vector for the stroke region. This vector describes the tilt direction and diffusion range of the stroke. The sharpness feature and the directional feature vector are combined to construct a stroke feature descriptor, which includes the morphological features and grayscale distribution features of the stroke. Finally, the stroke feature descriptors of all stroke region units are summarized to generate a stroke feature set. For example, in the aforementioned document signature authentication case, for the segmented signature stroke region units, the start and end endpoints of each stroke are found. The pixel grayscale value changes around these endpoints are analyzed to determine the sharpness of the stroke. The distribution of edge directions in the stroke region is statistically analyzed using the Histogram of Oriented Gradients (HGP) algorithm to obtain the tilt direction and diffusion range of the stroke. This information is combined into stroke feature descriptors, such as a sharp stroke, a tilt direction pointing to the upper left, and a small diffusion range. All stroke feature descriptors are then aggregated into a stroke feature set, which is used as part of the traditional handwriting feature set for handwriting identification.
[0033] In step 120, a texture pattern analysis operation is performed on the handwriting images in the target handwriting image set to generate a hybrid feature dataset containing both traditional handwriting feature sets and deep learning handwriting feature sets, including: Step 123: Based on the stroke region segmentation results and the stroke feature set, determine the region of interest for texture pattern analysis. The region of interest includes the internal region of the stroke and the edge region of the stroke. Calculate the gray-level co-occurrence matrix for the region of interest, and statistically analyze the gray-level co-occurrence probability of pixels under different directions and distances to generate spatial correlation features of the texture. Extract texture features from the region of interest using the local binary mode algorithm, calculate the gray-level comparison result of each pixel with its neighboring pixels, and generate a local texture pattern histogram. Combine the spatial correlation features and the local texture pattern histogram to construct a texture feature descriptor, which includes the thickness and periodicity of the texture. Fuse the texture feature descriptor with the stroke tilt angle, continuous stroke frequency, and turning curvature features in the traditional handwriting feature set to generate a traditional handwriting feature set.
[0034] In this embodiment of the invention, regions of interest (ROIs) for texture pattern analysis are determined based on stroke region segmentation results and a set of stroke feature characteristics. These regions primarily consist of the internal and edge regions of strokes. A gray-level co-occurrence matrix (GLCM) is calculated for each ROI, and the gray-level co-occurrence probabilities of pixels at different directions and distances are statistically analyzed. These probabilities reflect the spatial correlation characteristics of the texture, such as its continuity and directionality. A local binary pattern algorithm can be used to extract texture features from the ROIs, calculating the gray-level comparison between each pixel and its neighboring pixels to generate a local texture pattern histogram. This histogram describes the local features of the texture. The spatial correlation features and the local texture pattern histogram are combined to construct a texture feature descriptor, which includes the texture's thickness and periodicity. Finally, the texture feature descriptor is fused with stroke tilt angles, continuous stroke frequency, and turning curvature features from a traditional handwriting feature set to generate a new traditional handwriting feature set. For example, in the aforementioned document signature authentication case, the regions inside and at the edges of the signature strokes are determined as ROIs based on the stroke region segmentation results and the set of stroke feature characteristics. Gray-level co-occurrence matrix calculations were performed on these regions, revealing strong continuity in the signature stroke texture along a certain direction. A local texture pattern histogram was generated using a local binary pattern algorithm to illustrate local variations in the signature texture. This information was then used to construct a texture feature descriptor, indicating features such as fine signature stroke texture and a certain periodicity. This texture feature descriptor was then fused with stroke tilt angle, continuous stroke frequency, and turning curvature features from a traditional handwriting feature set to obtain a more comprehensive set of traditional handwriting features.
[0035] Step 124: Invoke the pre-trained convolutional neural network model to extract features from the handwriting images in the target handwriting image set. Perform layer-by-layer feature mapping on the handwriting images through convolutional layers to generate high-level semantic features. Input the high-level semantic features into the recurrent neural network model. Model the temporal dependency relationship of the stroke sequence through recurrent layers to extract temporal dependency features. Concatenate the high-level semantic features and the temporal dependency features along the channel dimension to generate a deep learning handwriting feature set.
[0036] In this embodiment of the invention, a pre-trained convolutional neural network (CNN) model is invoked to extract features from handwriting images in the target handwriting image set. The convolutional layers of the CNN perform layer-by-layer feature mapping on the handwriting images, extracting high-level semantic features from the images. These features reflect the abstract characteristics and semantic information of the handwriting. Then, the high-level semantic features are input into a recurrent neural network (RNN) model. The recurrent layers of the RNN can model the temporal dependencies of the stroke sequence. Because handwriting is a process with a temporal order, there are temporal dependencies between strokes, and temporal dependency features can be extracted through modeling. Finally, the high-level semantic features and temporal dependency features are concatenated along the channel dimension to generate a deep learning handwriting feature set. For example, in the aforementioned document signature identification case, a pre-trained CNN is used to process signature images in the target handwriting image set. The convolutional layers can extract high-level semantic features such as the overall structure and style of the signature. These high-level semantic features are then input into the RNN, and the recurrent layers can analyze the writing order and temporal relationships of the strokes in the signature, extracting temporal dependency features. By concatenating these two features along the channel dimension, a deep learning handwriting feature set containing high-level semantics and temporal dependency information is obtained.
[0037] Step 125: Align the traditional handwriting feature set and the deep learning handwriting feature set by feature dimension to generate a hybrid feature dataset with a unified dimension representation.
[0038] In this embodiment of the invention, the traditional handwriting feature set and the deep learning handwriting feature set may have different dimensions. To merge them, feature dimension alignment processing is required. A certain method is used to make the dimensions of the two feature sets consistent, generating a hybrid feature dataset with a unified dimensional representation. For example, in the aforementioned document signature identification case, the number and dimensions of features in the traditional handwriting feature set and the deep learning handwriting feature set may differ. Interpolation, dimensionality reduction, and other methods can be used to process the two feature sets to make their dimensions the same, and then they can be merged into a unified hybrid feature dataset for model training and analysis.
[0039] Step 130: Input the hybrid feature dataset into the Triplet network model, and jointly optimize and train the traditional handwriting feature set and the deep learning handwriting feature set through the triplet loss function to generate the first discriminative feature vector of the handwriting image to be identified and the second discriminative feature vector of the known sample handwriting image. The triplet loss function realizes the aggregation of similar handwriting features and the separation of dissimilar handwriting features through the feature distance constraints of anchor samples, positive samples and negative samples.
[0040] In this embodiment of the invention, the generated hybrid feature dataset is input into the Triplet network model for training. The Triplet network model performs joint optimization training on the traditional handwriting feature set and the deep learning handwriting feature set through a triplet loss function. The triplet loss function uses anchor samples, positive samples, and negative samples, and by constraining the feature distance between them, it aggregates similar handwriting features together and separates dissimilar handwriting features.
[0041] During training, the model's parameters are continuously adjusted to enable it to learn discriminative features of handwriting. Ultimately, a first discriminative feature vector is generated for the handwriting image to be identified, and a second discriminative feature vector is generated for known sample handwriting images. For example, in the aforementioned document signature identification case, a mixed feature dataset is input into the Triplet network model. One signature image is selected as the anchor sample, other signature images from the same author are used as positive samples, and signature images from different authors are used as negative samples.
[0042] By using a triplet loss function, the model learns the feature similarity of signatures from the same author and the feature differences between signatures from different authors. After training, the model obtains the first discriminative feature vector of the signature image to be identified and the second discriminative feature vector of the known sample signature images, which are then used for subsequent identification.
[0043] As an optional embodiment, step 130 can be implemented by any of the embodiments 1-4 below.
[0044] Example 1: Step 1311: Construct a Triplet network model that includes a feature input layer, a shared feature encoding layer, a triplet loss calculation layer, and a feature output layer. The shared feature encoding layer includes a convolutional sub-network and a recurrent sub-network.
[0045] In this embodiment of the invention, the constructed Triplet network model comprises four main layers. The feature input layer receives a mixed feature dataset and inputs the data into the model. The shared feature encoding layer includes a convolutional sub-network and a recurrent sub-network. The convolutional sub-network performs convolution operations on the input features to extract spatial information, while the recurrent sub-network models the temporal information of the features. The triplet loss calculation layer calculates the triplet loss function value based on the features of the input anchor samples, positive samples, and negative samples. The feature output layer outputs the trained feature vector. For example, in the aforementioned document signature authentication case, the feature input layer of the constructed Triplet network model receives a mixed feature dataset of the signature image. The convolutional sub-network of the shared feature encoding layer performs convolution operations on the features to extract spatial features of the signature, while the recurrent sub-network models the temporal features of the signature strokes. The triplet loss calculation layer calculates the loss value based on the features of the selected anchor samples, positive samples, and negative samples, and the feature output layer outputs the discriminative feature vector of the signature image.
[0046] Step 1312: Input the traditional handwriting feature set in the hybrid feature dataset into the recurrent subnetwork of the shared feature encoding layer, and perform temporal modeling processing on the stroke sequence features through the long short-term memory unit to generate the traditional feature encoding vector.
[0047] In this embodiment of the invention, the traditional handwriting feature set from the mixed feature dataset is input into a recurrent subnetwork with a shared feature encoding layer. The recurrent subnetwork uses Long Short-Term Memory (LSTM) units to perform temporal modeling of the stroke sequence features. LSTM units can process sequential data and remember long-term dependencies within the sequence. By modeling the stroke sequence features, a traditional feature encoding vector is generated. For example, in the aforementioned document signature authentication case, the traditional handwriting feature set of the signature image is input into the recurrent subnetwork. The LSTM unit analyzes the writing order and temporal relationship of the strokes in the signature, models these features, and generates a traditional feature encoding vector containing the temporal information of the signature strokes.
[0048] Step 1313: Input the deep learning handwriting feature set in the hybrid feature dataset into the convolutional sub-network of the shared feature encoding layer, and perform deep encoding of high-level semantic features through multi-layer convolution operations and pooling processing to generate deep learning feature encoding vectors.
[0049] In this embodiment of the invention, a set of deep learning handwriting features from a hybrid feature dataset is input into a convolutional sub-network with a shared feature encoding layer. The convolutional sub-network performs deep encoding of high-level semantic features through multiple convolutional operations and pooling. Multiple convolutional operations progressively extract detailed and abstract information from the features, while pooling reduces the dimensionality of the features and decreases computational cost. These operations generate a deep learning feature encoding vector. For example, in the aforementioned document signature authentication case, the set of deep learning handwriting features from the signature image is input into the convolutional sub-network. The multi-layer convolutional operations of the convolutional sub-network further extract and analyze the high-level semantic features of the signature, and pooling reduces the dimensionality of the features, ultimately generating a deep learning feature encoding vector that contains the high-level semantic information of the signature.
[0050] Step 1314: Input the traditional feature encoding vector and the deep learning feature encoding vector into the feature fusion unit, and perform feature fusion processing through element-level addition to generate a fused feature encoding vector.
[0051] In this embodiment of the invention, the generated traditional feature encoding vector and deep learning feature encoding vector are input into the feature fusion unit. In the feature fusion unit, the two vectors are fused through element-wise addition to generate a fused feature encoding vector. Element-wise addition involves adding the elements at corresponding positions of the two vectors to obtain a new vector. For example, in the aforementioned document signature authentication case, the traditional feature encoding vector and deep learning feature encoding vector are input into the feature fusion unit. Adding the elements at corresponding positions of the two vectors yields a fused feature encoding vector containing signature stroke temporal information and high-level semantic information.
[0052] Step 1315: Select anchor sample feature vector, positive sample feature vector and negative sample feature vector from the mixed feature dataset. The positive sample feature vector and the anchor sample feature vector are from the same author, and the negative sample feature vector and the anchor sample feature vector are from different authors.
[0053] In this embodiment of the invention, anchor sample feature vectors, positive sample feature vectors, and negative sample feature vectors are selected from a mixed feature dataset. The anchor sample feature vector serves as a baseline feature vector; the positive sample feature vector and the anchor sample feature vector originate from the same author, while the negative sample feature vector and the anchor sample feature vector originate from different authors. For example, in the aforementioned document signature authentication case, the feature vector of one signature image is selected as the anchor sample feature vector from the mixed feature dataset of signature images. Feature vectors of other signature images by the same author are selected as positive sample feature vectors, and feature vectors of signature images by different authors are selected as negative sample feature vectors.
[0054] Step 1316: Input the anchor sample feature vector, the positive sample feature vector, and the negative sample feature vector into the Triplet network model respectively, and generate the corresponding anchor sample encoding vector, positive sample encoding vector, and negative sample encoding vector through the shared feature encoding layer.
[0055] In this embodiment of the invention, the selected anchor sample feature vector, positive sample feature vector, and negative sample feature vector are respectively input into the Triplet network model. In the shared feature encoding layer, these feature vectors are encoded to generate corresponding anchor sample encoding vectors, positive sample encoding vectors, and negative sample encoding vectors. For example, in the aforementioned document signature authentication case, the selected anchor sample feature vector, positive sample feature vector, and negative sample feature vector are input into the shared feature encoding layer of the Triplet network model. The shared feature encoding layer processes these vectors to generate corresponding encoding vectors, which contain higher-level feature information of the signature.
[0056] Step 1317: Calculate the Euclidean distance between the anchor sample encoding vector and the positive sample encoding vector as the same-class feature distance, and calculate the Euclidean distance between the anchor sample encoding vector and the negative sample encoding vector as the different-class feature distance.
[0057] In this embodiment of the invention, the Euclidean distance between the anchor sample encoding vector and the positive sample encoding vector is calculated. This distance serves as the distance between similar features, because positive samples and anchor samples come from the same author, their features should be similar, and the distance should be small. Simultaneously, the Euclidean distance between the anchor sample encoding vector and the negative sample encoding vector is calculated. This distance serves as the distance between dissimilar features, because negative samples and anchor samples come from different authors, their features should be significantly different, and the distance should be large. For example, in the aforementioned document signature authentication case, calculating the Euclidean distance between the anchor sample encoding vector and the positive sample encoding vector reflects the similarity of signature features of the same author. Calculating the Euclidean distance between the anchor sample encoding vector and the negative sample encoding vector reflects the degree of difference in signature features of different authors.
[0058] Step 1318: Construct a triplet loss function based on the distance between similar features and the distance between dissimilar features, minimize the triplet loss function using an optimization algorithm, and adjust the network parameters of the shared feature encoding layer; when the value of the triplet loss function converges to a preset threshold, stop the model training process and save the current network parameters as the optimal model parameters.
[0059] In this embodiment of the invention, a triplet loss function is constructed based on the distances of similar and dissimilar features. The purpose of this loss function is to minimize the distances of similar features and maximize the distances of dissimilar features. Through optimization algorithms, such as stochastic gradient descent, the network parameters of the shared feature encoding layer are continuously adjusted, causing the value of the triplet loss function to gradually decrease. When the value of the loss function converges to a preset threshold, it indicates that the model has learned a better feature representation, and the model training process is stopped, saving the current network parameters as the optimal model parameters. For example, in the aforementioned document signature authentication case, the constructed triplet loss function reduces the feature distance of signatures from the same author and increases the feature distance of signatures from different authors. The stochastic gradient descent algorithm can be used to adjust the network parameters of the shared feature encoding layer, and when the loss function converges to the preset threshold, training is stopped, and the current network parameters are saved.
[0060] Step 1319: Input the mixed feature dataset of the handwriting image to be identified into the trained Triplet network model, and generate a first discriminative feature vector through the shared feature encoding layer; input the mixed feature data of the known sample handwriting image into the trained Triplet network model, and generate a second discriminative feature vector through the shared feature encoding layer.
[0061] In this embodiment of the invention, a mixed feature dataset of the handwriting images to be identified is input into a trained Triplet network model. In the shared feature encoding layer, this data is encoded to generate a first discriminative feature vector. Similarly, mixed feature data of known sample handwriting images is input into the trained model to generate a second discriminative feature vector. For example, in the aforementioned document signature identification case, a mixed feature dataset of the signature images to be identified is input into the trained Triplet network model to generate a first discriminative feature vector. Mixed feature data of known sample signature images is input into the model to generate a second discriminative feature vector.
[0062] Example 2: Step 1321: Initialize the network parameters of the shared feature encoding layer of the Triplet network model, and set the initial weight values of the convolutional sub-network and the recurrent sub-network.
[0063] In this embodiment of the invention, the network parameters of the shared feature encoding layer of the Triplet network model are initialized, including the initial weight values of the convolutional sub-network and the recurrent sub-network. The setting of the initial weight values affects the model's training process and final performance. Random initialization can be used to set the weight values. For example, in the aforementioned document signature authentication case, the weight values of the convolutional and recurrent sub-networks of the shared feature encoding layer of the Triplet network model are randomly initialized to provide an initial state for the model's training.
[0064] Step 1322: Divide the hybrid feature dataset into a training dataset and a validation dataset according to a preset ratio. The training dataset is used for model parameter optimization, and the validation dataset is used for model performance evaluation.
[0065] In this embodiment of the invention, the hybrid feature dataset is divided into a training dataset and a validation dataset according to a preset ratio. The training dataset is used for model parameter optimization. During training, the model's parameters are continuously adjusted to enable the model to learn the features of the data. The validation dataset is used to evaluate the model's performance. During training, the model is periodically evaluated using the validation dataset to check its generalization ability. For example, in the aforementioned document signature identification case, the hybrid feature dataset of the signature image is divided into a training dataset and a validation dataset in a ratio of 80% and 20%. The training dataset can be used to optimize the parameters of the Triplet network model, and the validation dataset can be used to evaluate the model's performance.
[0066] Step 1323: The training dataset is iteratively trained using the batch gradient descent algorithm. In each iteration, multiple triplet samples are randomly selected from the training dataset. Each triplet sample contains an anchor sample, a positive sample, and a negative sample.
[0067] In this embodiment of the invention, the batch gradient descent algorithm is used to iteratively train the training dataset. During each iteration, multiple triplet samples are randomly selected from the training dataset. Each triplet sample contains an anchor sample, a positive sample, and a negative sample. The batch gradient descent algorithm updates the model parameters by calculating the gradient of the samples. For example, in the aforementioned document signature identification case, the batch gradient descent algorithm is used to iteratively train the training dataset of signature images. In each iteration, multiple triplet samples of signature images are randomly selected from the training dataset. Each triplet sample contains an anchor sample signature image, a positive sample signature image, and a negative sample signature image.
[0068] Step 1324: Input each triplet sample into the Triplet network model, calculate the triplet loss function value, and update the network parameters of the shared feature encoding layer through the backpropagation algorithm.
[0069] In this embodiment of the invention, each selected triplet sample is input into the Triplet network model, and the triplet loss function value is calculated. Using the backpropagation algorithm, the gradient is calculated based on the loss function value, and the network parameters of the shared feature encoding layer are updated. The backpropagation algorithm can propagate the error from the output layer back to the input layer, adjusting the weight values in the network. For example, in the aforementioned document signature authentication case, selected triplet samples of the signature image are input into the Triplet network model, and the triplet loss function value is calculated. The backpropagation algorithm can be used to update the weight values of the convolutional and recurrent sub-networks of the shared feature encoding layer based on the loss function value.
[0070] Step 1325: After each iteration, use the validation dataset to evaluate the performance of the current model parameters, and calculate the triplet loss function value and feature matching accuracy on the validation dataset.
[0071] In this embodiment of the invention, after each iteration, the performance of the current model parameters is evaluated using a validation dataset. The triplet loss function value and feature matching accuracy on the validation dataset are calculated. The triplet loss function value reflects the model's loss on the validation dataset, and the feature matching accuracy reflects the model's classification accuracy on the validation dataset. For example, in the aforementioned document signature identification case, after each iteration, the parameters of the current Triplet network model are evaluated using a validation dataset of signature images, and the triplet loss function value and signature image feature matching accuracy on the validation dataset are calculated.
[0072] Step 1326: Compare the validation loss function value of the current iteration with the validation loss function value of the previous iteration. If the validation loss function value does not decrease for several consecutive iterations, reduce the learning rate parameter and continue training.
[0073] In this embodiment of the invention, the validation loss function value of the current iteration is compared with that of the previous iteration. If the validation loss function value fails to decrease for several consecutive iterations, it indicates that the model may be trapped in a local optimum or the learning rate is too high. In this case, the learning rate parameter is reduced and training continues, allowing the model to escape the local optimum and continue learning. For example, in the aforementioned document signature authentication case, during the training of the Triplet network model, the validation loss function value is compared for each iteration. If the validation loss function value fails to decrease for several consecutive iterations, the learning rate parameter is reduced, and the model training continues.
[0074] Step 1327: When the preset maximum number of iterations is reached or the verification loss function value is less than the preset threshold, stop the model training process and save the current model parameters as the final training result.
[0075] In this embodiment of the invention, the model training process is stopped when the preset maximum number of iterations is reached or the verification loss function value is less than a preset threshold. The preset maximum number of iterations is to prevent the model training time from becoming too long, and the preset threshold is to determine whether the model has converged. After training is stopped, the current model parameters are saved as the final training result. For example, in the aforementioned document signature authentication case, when the training of the Triplet network model reaches the preset maximum number of iterations or the verification loss function value is less than the preset threshold, training is stopped, and the current model parameters are saved.
[0076] Step 1328: Configure the Triplet network model using the model parameters of the final training result, and perform feature encoding processing on the mixed feature data of the handwriting image to be identified and the known sample handwriting image to generate a first discriminative feature vector and a second discriminative feature vector respectively.
[0077] In this embodiment of the invention, the Triplet network model is configured using the model parameters from the final training result. The mixed feature data of the handwriting image to be identified and the handwriting images of known samples are input into the configured model for feature encoding processing, generating a first discriminative feature vector and a second discriminative feature vector, respectively. For example, in the aforementioned document signature identification case, the Triplet network model is configured using the model parameters from the final training result. The mixed feature data of the signature image to be identified and the signature images of known samples are input into the model to generate a first discriminative feature vector and a second discriminative feature vector.
[0078] Example 3: Step 1331: Select triplet samples that meet the hard sample condition from the mixed feature dataset. The hard sample condition is a combination of samples whose distance between features of different classes is less than the sum of the distance between features of the same class and a preset margin value.
[0079] In this embodiment of the invention, triplet samples that meet the hard example sample criteria are selected from the mixed feature dataset. The hard example sample criteria are sample combinations where the distance between dissimilar features is less than the sum of the distances between similar features and a preset margin value. These samples are relatively difficult for the model to distinguish. By training with these hard example samples, the model's ability to distinguish similar handwriting features can be improved. For example, in the aforementioned document signature identification case, triplet samples of signature images where the distance between dissimilar features is less than the sum of the distances between similar features and a preset margin value are selected from the mixed feature dataset of signature images. These samples may represent cases where the signature features of different authors are relatively similar.
[0080] Step 1332: Train the Triplet network model using the difficult examples, and improve the model's ability to distinguish similar handwriting features by increasing the training weights of the difficult examples.
[0081] In this embodiment of the invention, the Triplet network model is trained using selected difficult example samples. During training, the training weights of the difficult example samples are increased, causing the model to pay more attention to these hard-to-distinguish samples, thereby improving the model's ability to distinguish similar handwriting features. For example, in the aforementioned document signature identification case, the selected difficult example signature image triplet samples are input into the Triplet network model for training, increasing the training weights of these difficult example samples, causing the model to work harder to learn the feature differences of these samples.
[0082] Step 1333: Dynamically adjust the selection threshold for difficult examples during model training. Gradually lower the selection criteria for difficult examples as the number of training iterations increases, thereby expanding the selection range of difficult examples.
[0083] In this embodiment of the invention, the selection threshold for difficult examples is dynamically adjusted during model training. As the number of training iterations increases, the selection criteria for difficult examples are gradually lowered, expanding the range of difficult examples to be selected. This allows the model to encounter samples of varying difficulty at different stages, improving the model's generalization ability. For example, in the aforementioned document signature identification case, during the training of the Triplet network model, as the number of iterations increases, the selection criteria for difficult signature image triplet samples are gradually lowered, selecting more difficult examples for training.
[0084] Step 1334: Evaluate the feature importance of the traditional feature encoding vector and the deep learning feature encoding vector during the training process, and dynamically adjust the weight ratio of the two feature encoding vectors in the fusion process through an attention mechanism.
[0085] In this embodiment of the invention, the importance of traditional feature encoding vectors and deep learning feature encoding vectors during the training process is evaluated. An attention mechanism is used to dynamically adjust the weight ratio of the two types of feature encoding vectors during the fusion process based on the importance of the features. The attention mechanism allows the model to focus more on important features. For example, in the aforementioned document signature authentication case, during the training of the Triplet network model, an attention mechanism is used to evaluate the importance of traditional feature encoding vectors and deep learning feature encoding vectors, and their weight ratios during the fusion process are dynamically adjusted based on the evaluation results.
[0086] Step 1335: When the contribution of the traditional feature encoding vector to the loss function is higher than that of the deep learning feature encoding vector, increase the fusion weight of the traditional feature encoding vector; when the contribution of the deep learning feature encoding vector exceeds the set contribution, increase the fusion weight of the deep learning feature encoding vector.
[0087] In this embodiment of the invention, the fusion weights of the traditional feature encoding vector and the deep learning feature encoding vector are adjusted based on their contributions to the loss function. When the contribution of the traditional feature encoding vector to the loss function is higher than that of the deep learning feature encoding vector, the fusion weight of the traditional feature encoding vector is increased. When the contribution of the deep learning feature encoding vector exceeds a set contribution level, the fusion weight of the deep learning feature encoding vector is increased. For example, in the aforementioned document signature authentication case, during the training of the Triplet network model, the contributions of the traditional feature encoding vector and the deep learning feature encoding vector to the loss function are calculated. If the contribution of the traditional feature encoding vector is high, its fusion weight is increased; if the contribution of the deep learning feature encoding vector exceeds a set value, its fusion weight is increased.
[0088] Step 1336: Achieve adaptive fusion of traditional handwriting features and deep learning handwriting features through dynamic weight adjustment to generate a fusion feature encoding vector with stronger discriminative ability.
[0089] In this embodiment of the invention, the adaptive fusion of traditional handwriting features and deep learning handwriting features is achieved by dynamically adjusting the fusion weights of the traditional feature encoding vector and the deep learning feature encoding vector. This fusion method can automatically adjust the fusion ratio of the two features according to the characteristics of the data and the training status of the model, generating a fused feature encoding vector with stronger discriminative power. For example, in the aforementioned document signature identification case, during the training process of the Triplet network model, the traditional handwriting features and deep learning handwriting features are adaptively fused by dynamically adjusting the weights, generating a fused feature encoding vector with stronger discriminative power.
[0090] Step 1337: Based on the optimization results of the fused feature encoding vector and the triplet loss function, extract the high-dimensional feature representations of the handwriting image to be identified and the known sample handwriting images, respectively, as the first discriminative feature vector and the second discriminative feature vector.
[0091] In this embodiment of the invention, based on the optimization results of the fused feature encoding vector and the triplet loss function, high-dimensional feature representations are extracted from the handwriting image to be identified and the known sample handwriting image, respectively. These high-dimensional feature representations serve as the first and second discriminative feature vectors. For example, in the aforementioned document signature identification case, based on the optimization results of the fused feature encoding vector and the triplet loss function, high-dimensional feature representations are extracted from the signature image to be identified and the known sample signature image, respectively, serving as the first and second discriminative feature vectors.
[0092] Example 4: Step 1341: Introduce a feature normalization processing unit in the shared feature encoding layer of the Triplet network model to perform L2 normalization processing on the traditional feature encoding vector and the deep learning feature encoding vector, so that the magnitude of the feature vector is unified to a preset constant.
[0093] In this embodiment of the invention, a feature normalization unit is introduced into the shared feature encoding layer of the Triplet network model. L2 normalization is performed on both traditional and deep learning feature encoding vectors, unifying the magnitude of the feature vectors to a preset constant. This allows feature vectors to be compared on the same scale, improving model stability and training performance. For example, in the aforementioned document signature authentication case, the feature normalization unit introduced into the shared feature encoding layer of the Triplet network model performs L2 normalization on both traditional and deep learning feature encoding vectors, unifying the magnitude of the signature image feature vector to a preset constant.
[0094] Step 1342: Calculate the distance between similar features and the distance between dissimilar features using the normalized feature vectors.
[0095] In this embodiment of the invention, normalized feature vectors are used to calculate the distances between similar and dissimilar features. Normalized feature vectors, operating at the same scale, yield more comparable distances. For example, in the aforementioned document signature authentication case, using normalized signature image feature vectors to calculate the distances between similar and dissimilar features allows for a more accurate reflection of the similarity and differences in signature features.
[0096] Step 1343: Construct a variant of the triplet loss function that includes a temperature coefficient. By adjusting the temperature coefficient, the weight ratio of the distance between similar and dissimilar features is controlled, thereby enhancing the sensitivity of the loss function to differences in feature distances.
[0097] In this embodiment of the invention, a variant of the triplet loss function incorporating a temperature coefficient is constructed. By adjusting the temperature coefficient, the weight ratio of similar feature distances to dissimilar feature distances can be controlled, enhancing the sensitivity of the loss function to differences in feature distances. The larger the temperature coefficient, the more sensitive the loss function is to differences in feature distances. For example, in the aforementioned document signature authentication case, a variant of the triplet loss function incorporating a temperature coefficient is constructed. By adjusting the temperature coefficient, the loss function is made to pay more attention to the differences between similar and dissimilar feature distances.
[0098] Step 1344: When the difference between the distance of similar features and the distance of dissimilar features does not reach the preset difference range, increase the temperature coefficient to increase the loss function value; when the difference reaches the preset difference range, decrease the temperature coefficient to decrease the loss function value.
[0099] In this embodiment of the invention, the temperature coefficient is adjusted based on the difference between the distances of similar and dissimilar features. When the difference does not reach a preset range, it indicates that the model is not distinguishing between similar and dissimilar features well enough. In this case, the temperature coefficient is increased to improve the loss function value, making the model pay more attention to these differences. When the difference reaches the preset range, it indicates that the model can distinguish between similar and dissimilar features well. In this case, the temperature coefficient is decreased to reduce the loss function value and avoid overfitting. For example, in the aforementioned document signature authentication case, during the training of the Triplet network model, the temperature coefficient is adjusted based on the difference between the distances of similar and dissimilar signature features to optimize the loss function.
[0100] Step 1345: During model training, monitor the degree of aggregation of similar features and the degree of separation of dissimilar features in real time, and evaluate the feature learning effect of the model by calculating the intra-class variance and inter-class distance of the feature vectors.
[0101] In this embodiment of the invention, the degree of aggregation of similar features and the degree of separation of dissimilar features are monitored in real time during model training. The feature learning effect of the model is evaluated by calculating the intra-class variance and inter-class distance of the feature vectors. The smaller the intra-class variance, the higher the degree of aggregation of similar features; the larger the inter-class distance, the higher the degree of separation of dissimilar features. For example, in the aforementioned document signature authentication case, during the training of the Triplet network model, the intra-class variance and inter-class distance of the signature image feature vectors are calculated in real time to evaluate the model's learning effect on signature features.
[0102] Step 1346: When the intra-class variance is less than the preset variance threshold and the inter-class distance is greater than the preset distance threshold, the model training is deemed successful and the parameter optimization process is stopped.
[0103] In this embodiment of the invention, when the intra-class variance is less than a preset variance threshold and the inter-class distance is greater than a preset distance threshold, it indicates that the model has been able to effectively aggregate similar features together and separate dissimilar features. The model training is then deemed successful, and the parameter optimization process is stopped. For example, in the aforementioned document signature authentication case, when the intra-class variance of the signature image feature vector is less than a preset variance threshold and the inter-class distance is greater than a preset distance threshold, the Triplet network model training is deemed successful, and the optimization of the model parameters is stopped.
[0104] Step 1347: Apply the trained model parameters to the feature encoding process, and encode the mixed feature data of the handwriting image to be identified and the known sample handwriting image respectively to generate a first discriminative feature vector and a second discriminative feature vector with normalization properties.
[0105] In this embodiment of the invention, the trained model parameters are applied to the feature encoding process. The mixed feature data of the handwriting image to be identified and the known sample handwriting images are encoded separately to generate a first and a second discriminative feature vector with normalization properties. For example, in the aforementioned document signature identification case, the parameters of the trained Triplet network model are applied to the feature encoding process to encode the mixed feature data of the signature image, generating a first and a second discriminative feature vector with normalization properties.
[0106] Step 140: Based on the first discriminative feature vector and the second discriminative feature vector, generate a handwriting image identification report that includes feature matching scores and author attribution probability.
[0107] In this embodiment of the invention, a handwriting image identification report is generated based on the generated first and second discriminative feature vectors. The report includes a feature matching score and an author attribution probability. The feature matching score reflects the degree of feature similarity between the handwriting image to be identified and known sample handwriting images, while the author attribution probability indicates the likelihood that the handwriting image to be identified belongs to a particular author. For example, in the aforementioned document signature identification case, a signature identification report containing a feature matching score and an author attribution probability is generated based on the first discriminative feature vector of the signature image to be identified and the second discriminative feature vector of the known sample signature image.
[0108] In a non-limiting embodiment, step 140 includes: Step 1410: Calculate the cosine similarity between the first discriminative feature vector and the second discriminative feature vector as the basic metric for the feature matching score; convert the cosine similarity into a feature matching score using a preset similarity mapping function; collect the second discriminative feature vectors of the known sample handwriting images to construct a feature database, where each feature vector in the feature database is associated with corresponding author identification information; compare the first discriminative feature vector of the handwriting image to be identified with all the second discriminative feature vectors in the feature database one by one, and calculate the feature matching score for each comparison result; sort the feature matching scores of all comparison results, and select the top N comparison results with the highest scores as candidate matching items, where N is a preset number of candidates; based on the feature matching scores of the candidate matching items and the corresponding author identification information, calculate the attribution probability of each candidate author using a probability normalization algorithm; organize the feature matching scores and author attribution probabilities according to a preset report format to generate a handwriting image identification report containing the handwriting image identifier to be identified, a list of candidate authors, the distribution of feature matching scores, and the ranking of author attribution probabilities.
[0109] In this embodiment of the invention, the cosine similarity between the first and second discriminative feature vectors is first calculated. Cosine similarity measures the degree of similarity between two vectors and serves as the basic metric for feature matching scores. Then, a preset similarity mapping function is used to convert the cosine similarity into a feature matching score. Next, second discriminative feature vectors from known sample handwriting images are collected to construct a feature database, with each feature vector associated with corresponding author identification information. The first discriminative feature vector of the handwriting image to be identified is compared one by one with all the second discriminative feature vectors in the feature database, and the feature matching score for each comparison result is calculated. The feature matching scores of all comparison results are sorted, and the top N comparison results with the highest scores are selected as candidate matches. Based on the feature matching scores of the candidate matches and the corresponding author identification information, a probability normalization algorithm is used to calculate the attribution probability of each candidate author. Finally, the feature matching scores and author attribution probabilities are organized according to a preset report format to generate a handwriting image identification report.
[0110] For example, in the aforementioned document signature authentication case, the cosine similarity between the first discriminative feature vector of the signature image to be authenticated and the second discriminative feature vector of a known sample signature image is calculated and converted into a feature matching score. A feature database of known sample signature images is constructed, and the first discriminative feature vector of the signature image to be authenticated is compared with all feature vectors in the database to calculate the feature matching score. The top N comparison results with the highest scores are selected as candidate matches, and the attribution probability of each candidate author is calculated. A signature authentication report is generated according to a preset report format, containing the identifier of the signature image to be authenticated, a list of candidate authors, the distribution of feature matching scores, and the ranking of author attribution probabilities.
[0111] In a non-limiting embodiment, the method further includes: inputting the high-level semantic features and the temporal dependency features into a bidirectional attention mechanism module; calculating a first attention weight distribution of the high-level semantic features on the temporal dependency features and a second attention weight distribution of the temporal dependency features on the high-level semantic features through the bidirectional attention mechanism module; performing weighted optimization processing on the temporal dependency features based on the first attention weight distribution to obtain optimized temporal dependency features; performing weighted optimization processing on the high-level semantic features based on the second attention weight distribution to obtain optimized high-level semantic features; concatenating the optimized temporal dependency features and the optimized high-level semantic features by feature dimensions to generate an optimized deep learning handwriting feature set; and feeding back the optimized deep learning handwriting feature set to the step of inputting the hybrid feature dataset into the Triplet network model to replace the original deep learning handwriting feature set for joint optimization training.
[0112] In this embodiment of the invention, high-level semantic features and temporal dependent features are input into a bidirectional attention mechanism module. Within this module, a first attention weight distribution of the high-level semantic features on the temporal dependent features and a second attention weight distribution of the temporal dependent features on the high-level semantic features are calculated. The temporal dependent features are then weighted and optimized according to the first attention weight distribution to obtain optimized temporal dependent features. Similarly, the high-level semantic features are weighted and optimized according to the second attention weight distribution to obtain optimized high-level semantic features. The optimized temporal dependent features and optimized high-level semantic features are then concatenated along their feature dimensions to generate an optimized deep learning handwriting feature set. This optimized deep learning handwriting feature set is then fed back into the step of inputting the hybrid feature dataset into the Triplet network model, replacing the original deep learning handwriting feature set for joint optimization training.
[0113] For example, in the aforementioned document signature authentication case, the high-level semantic features and temporal dependency features of the signature image are input into the bidirectional attention mechanism module. The attention weight distribution is calculated, and the two types of features are weighted and optimized to generate an optimized deep learning handwriting feature set. This set is then fed back into the training process of the Triplet network model, replacing the original deep learning handwriting feature set, allowing the model to be trained using the optimized features and improving the model's performance.
[0114] In a non-limiting embodiment, the method further includes: performing feature decoupling processing on the first discriminative feature vector and the second discriminative feature vector; decomposing the first discriminative feature vector into a first writing style sub-feature vector, a first stroke shape sub-feature vector, and a first temporal dynamic sub-feature vector; and decomposing the second discriminative feature vector into a second writing style sub-feature vector, a second stroke shape sub-feature vector, and a second temporal dynamic sub-feature vector, wherein the number and corresponding dimensions of each corresponding sub-feature vector remain consistent; calculating the style similarity between the first writing style sub-feature vector and the second writing style sub-feature vector, and the morphological similarity between the first stroke shape sub-feature vector and the second stroke shape sub-feature vector. The process involves: determining the dynamic similarity between the first and second temporal dynamic sub-feature vectors; generating dynamic weighting coefficients for each sub-feature vector based on style similarity, morphological similarity, and dynamic similarity using a feature importance evaluation algorithm; performing weighted fusion processing on each sub-feature vector using the dynamic weighting coefficients to obtain a weighted first discriminative feature vector and a weighted second discriminative feature vector; and re-inputting the weighted first and second discriminative feature vectors into the step of generating a handwriting image identification report containing feature matching scores and author attribution probabilities based on the first and second discriminative feature vectors, thereby obtaining an optimized handwriting image identification report.
[0115] In this embodiment of the invention, the first and second discriminative feature vectors are decoupled, decomposing them into first handwriting style sub-feature vectors, first stroke shape sub-feature vectors, first temporal dynamic sub-feature vectors, and second handwriting style sub-feature vectors, second stroke shape sub-feature vectors, and second temporal dynamic sub-feature vectors, respectively, with the number and corresponding dimensions of each sub-feature vector remaining consistent. The style similarity between the first and second handwriting style sub-feature vectors, the morphological similarity between the first and second stroke shape sub-feature vectors, and the dynamic similarity between the first and second temporal dynamic sub-feature vectors are calculated. Based on these similarities, dynamic weighting coefficients are generated for each sub-feature vector using a feature importance evaluation algorithm. The sub-feature vectors are then weighted and fused using these dynamic weighting coefficients to obtain a weighted first discriminative feature vector and a weighted second discriminative feature vector. These weighted first and second discriminative feature vectors are then re-inputted into the step of generating a handwriting image identification report to obtain an optimized handwriting image identification report.
[0116] For example, in the aforementioned document signature authentication case, the first and second discriminative feature vectors of the signature image are decoupled into handwriting style, stroke shape, and temporal dynamic sub-feature vectors. The similarity of these sub-feature vectors is calculated, dynamic weighting coefficients are generated, and the sub-feature vectors are weighted and fused to obtain a weighted discriminative feature vector. This weighted discriminative feature vector is then re-inputted into the signature authentication report generation process, resulting in an optimized signature authentication report with higher accuracy and reliability.
[0117] This invention achieves efficient and accurate handwriting image identification. First, it collects handwriting images to be identified and known sample handwriting images to generate a target handwriting image set, avoiding data confusion and missing information, and ensuring the integrity and representativeness of the samples. Then, it sequentially performs stroke region segmentation, stroke feature extraction, and texture pattern analysis on the target handwriting image set. The resulting hybrid feature dataset integrates traditional handwriting feature sets and deep learning handwriting feature sets. Traditional handwriting feature sets, such as stroke tilt angle, continuous stroke frequency, and turning curvature features, can characterize the morphological features of handwriting at a microscopic level. Meanwhile, the high-level semantic features output by convolutional layers and the temporal dependency features extracted by recurrent layers in the deep learning handwriting feature set mine abstract information and writing order patterns from macroscopic and temporal dimensions. The combination of these two greatly enriches the expression of handwriting features and improves the comprehensiveness and depth of the features. The hybrid feature dataset is input into the Triplet network model, and the traditional and deep learning handwriting feature sets are jointly optimized and trained using the triplet loss function. By constraining the feature distances of anchor samples, positive samples, and negative samples, effective aggregation of similar handwriting features and clear separation of dissimilar handwriting features are achieved, resulting in stronger discriminative first and second discriminative feature vectors. The handwriting image identification report generated based on these two discriminative feature vectors, which includes feature matching scores and author attribution probabilities, provides objective, accurate, and detailed evidence for handwriting identification, improving the accuracy, reliability, and efficiency of handwriting identification.
[0118] See Figure 2 As shown in the figure, this is a schematic diagram of the basic structure of a handwriting image identification system 200 provided in an embodiment of the present invention. The handwriting image identification system 200 includes: Processor 201; Storage device 202, on which computer program 2020 is stored; When the computer program 2020 is executed by the processor 201, the processor 201 implements any of the aforementioned artificial intelligence-based handwriting image identification methods.
[0119] Based on the above, a readable storage medium is provided, on which a program or instructions are stored, which, when executed by a processor, implement the steps of the above-described method. It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
Claims
1. A handwriting image identification method based on artificial intelligence, characterized in that, The method includes: Collect handwriting images to be identified and known sample handwriting images, and generate a target handwriting image set based on the handwriting images to be identified and the known sample handwriting images; The handwriting images in the target handwriting image set are sequentially subjected to stroke region segmentation, stroke feature extraction and texture pattern analysis to generate a hybrid feature dataset containing a traditional handwriting feature set and a deep learning handwriting feature set. The traditional handwriting feature set includes stroke tilt angle, stroke frequency and turning curvature features, while the deep learning handwriting feature set includes high-level semantic features output by the convolutional layer and temporal dependency features extracted by the recurrent layer. The hybrid feature dataset is input into the Triplet network model, and the traditional handwriting feature set and the deep learning handwriting feature set are jointly optimized and trained through the triplet loss function to generate the first discriminative feature vector of the handwriting image to be identified and the second discriminative feature vector of the known sample handwriting image. The triplet loss function realizes the aggregation of similar handwriting features and the separation of dissimilar handwriting features through the feature distance constraints of anchor samples, positive samples and negative samples. Based on the first and second discriminative feature vectors, a handwriting image identification report is generated, which includes feature matching scores and author attribution probabilities.
2. The method according to claim 1, characterized in that, The process of collecting handwriting images to be identified and known sample handwriting images, and generating a target handwriting image set based on the handwriting images to be identified and the known sample handwriting images, includes: The system receives handwriting images to be identified and known sample handwriting images acquired through various acquisition devices. The handwriting images to be identified contain handwriting patterns formed by different writing media and writing tools, and the known sample handwriting images contain handwriting sample patterns formed by multiple known authors under different writing conditions. Image quality assessment processing is performed on the handwriting image to be identified and the known sample handwriting image. Valid handwriting images that meet the preset quality standards are screened out by edge clarity detection and contrast analysis, and invalid handwriting images with blurred areas or incomplete strokes are removed. Based on the screened handwriting images to be identified and the known sample handwriting images, an optimized handwriting image with the same resolution and color mode is generated. Based on the source type and author identification information of the optimized handwriting image, classification and labeling processing is performed to establish a classification index table containing the category to be identified and the known sample categories; Based on the classification index table, the optimized handwriting images are arranged and combined according to a preset category order to generate a target handwriting image set containing category labels and image identifiers. Each image unit in the target handwriting image set is associated with corresponding classification index information.
3. The method according to claim 1, characterized in that, Perform stroke region segmentation on the handwriting images in the target handwriting image set, including: The handwriting images in the target handwriting image set are processed to grayscale, converting the color handwriting images into single-channel grayscale images while preserving the grayscale gradient information of the handwriting region; The single-channel grayscale image is binarized using an adaptive threshold segmentation algorithm to separate the foreground and background regions of the handwriting, thereby generating a binarized handwriting image. The stroke region segmentation process is performed based on the contour features of the binarized handwriting image. The central skeleton line of the handwriting is obtained through the skeleton extraction algorithm. The complete handwriting is divided into multiple independent stroke region units according to the branch points and intersection points of the central skeleton line. Contour tracking is performed on each stroke region unit to extract the boundary pixel sequence of the stroke region unit. The orientation and length features of the stroke region unit are determined by combining the coordinate change trend of the boundary pixel sequence. The direction feature and the length feature are used as the stroke region segmentation result, which is used to provide regional positioning information for the extraction of stroke features.
4. The method according to claim 3, characterized in that, Perform stroke feature extraction on the handwriting images in the target handwriting image set, including: Based on the stroke region units in the stroke region segmentation result, locate the start and end endpoints of each stroke region unit to determine the extraction area of the stroke features. Pixel grayscale value analysis is performed on the extracted area of the pen stroke feature to calculate the grayscale change rate in the preset neighborhood around the start and end points, and to determine the sharpness feature of the pen stroke. The edge direction distribution of the brush tip region is statistically analyzed by the directional gradient histogram algorithm to generate the directional feature vector of the brush tip region. The directional feature vector is used to describe the tilt direction and diffusion range of the brush tip. By combining the sharpness feature and the direction feature vector, a pen stroke feature descriptor is constructed, which includes the morphological features and grayscale distribution features of the pen stroke. The stroke feature descriptors of all stroke region units are summarized to generate a stroke feature set, which is used as a component of the traditional handwriting feature set.
5. The method according to claim 4, characterized in that, Perform texture pattern analysis on the handwriting images in the target handwriting image set to generate a hybrid feature dataset containing both traditional handwriting feature sets and deep learning handwriting feature sets, including: Based on the stroke region segmentation results and the stroke feature set, the region of interest for texture pattern analysis is determined, which includes the internal region of the stroke and the edge region of the stroke. The gray-level co-occurrence matrix is calculated for the region of interest, and the gray-level co-occurrence probability of pixels under different directions and distances is statistically analyzed to generate spatial correlation features of the texture. The texture features of the region of interest are extracted by the local binary mode algorithm, and the grayscale comparison results of each pixel with its neighboring pixels are calculated to generate a local texture mode histogram. By combining the spatial correlation features and the local texture pattern histogram, a texture feature descriptor is constructed, which includes the texture coarseness and periodicity features; The texture feature descriptor is fused with the stroke tilt angle, stroke frequency and turning curvature features in the traditional handwriting feature set to generate the traditional handwriting feature set. The pre-trained convolutional neural network model is invoked to extract features from the handwriting images in the target handwriting image set. The handwriting images are then mapped layer by layer through convolutional layers to generate high-level semantic features. The high-level semantic features are input into a recurrent neural network model, and the temporal dependency relationship of the stroke sequence is modeled through the recurrent layer to extract the temporal dependency features. The high-level semantic features and the temporal dependency features are concatenated along the channel dimension to generate a deep learning handwriting feature set. The traditional handwriting feature set and the deep learning handwriting feature set are aligned in terms of feature dimensions to generate a hybrid feature dataset with a unified dimensional representation.
6. The method according to any one of claims 1-5, characterized in that, The process involves inputting the hybrid feature dataset into a Triplet network model, and jointly optimizing and training the traditional handwriting feature set and the deep learning handwriting feature set using a triplet loss function to generate a first discriminative feature vector for the handwriting image to be identified and a second discriminative feature vector for the known sample handwriting images, including: A Triplet network model is constructed, comprising a feature input layer, a shared feature encoding layer, a triplet loss calculation layer, and a feature output layer. The shared feature encoding layer includes a convolutional sub-network and a recurrent sub-network. The traditional handwriting feature set in the hybrid feature dataset is input into the recurrent subnetwork of the shared feature encoding layer. The stroke sequence features are processed by temporal modeling through the long short-term memory unit to generate the traditional feature encoding vector. The deep learning handwriting feature set in the hybrid feature dataset is input into the convolutional sub-network of the shared feature encoding layer. The high-level semantic features are deeply encoded through multi-layer convolution operations and pooling to generate a deep learning feature encoding vector. The traditional feature encoding vector and the deep learning feature encoding vector are input into the feature fusion unit, and feature fusion processing is performed through element-level addition to generate a fused feature encoding vector. Anchor sample feature vectors, positive sample feature vectors, and negative sample feature vectors are selected from the hybrid feature dataset, wherein the positive sample feature vectors and the anchor sample feature vectors are from the same author, and the negative sample feature vectors and the anchor sample feature vectors are from different authors; The anchor sample feature vector, the positive sample feature vector, and the negative sample feature vector are respectively input into the Triplet network model, and the corresponding anchor sample encoding vector, positive sample encoding vector, and negative sample encoding vector are generated through the shared feature encoding layer. The Euclidean distance between the anchor sample encoding vector and the positive sample encoding vector is calculated as the same-class feature distance, and the Euclidean distance between the anchor sample encoding vector and the negative sample encoding vector is calculated as the different-class feature distance; A triplet loss function is constructed based on the distance between similar features and the distance between dissimilar features. The network parameters of the shared feature encoding layer are adjusted by minimizing the triplet loss function through an optimization algorithm. When the value of the triplet loss function converges to a preset threshold, the model training process is stopped, and the current network parameters are saved as the optimal model parameters. The mixed feature dataset of the handwriting image to be identified is input into the trained Triplet network model, and a first discriminative feature vector is generated through the shared feature encoding layer; The mixed feature data of the known sample handwriting images are input into the trained Triplet network model, and a second discriminative feature vector is generated through the shared feature encoding layer.
7. The method according to any one of claims 1-5, characterized in that, The process involves inputting the hybrid feature dataset into a Triplet network model, and jointly optimizing and training the traditional handwriting feature set and the deep learning handwriting feature set using a triplet loss function to generate a first discriminative feature vector for the handwriting image to be identified and a second discriminative feature vector for the known sample handwriting images, including: Initialize the network parameters of the shared feature encoding layer of the Triplet network model, and set the initial weight values of the convolutional sub-network and the recurrent sub-network; The hybrid feature dataset is divided into a training dataset and a validation dataset according to a preset ratio. The training dataset is used for model parameter optimization, and the validation dataset is used for model performance evaluation. The training dataset is iteratively trained using the batch gradient descent algorithm. In each iteration, multiple triplet samples are randomly selected from the training dataset. Each triplet sample contains an anchor sample, a positive sample, and a negative sample. Each triplet sample is input into the Triplet network model, the triplet loss function value is calculated, and the network parameters of the shared feature encoding layer are updated through the backpropagation algorithm. After each iteration, the performance of the current model parameters is evaluated using the validation dataset, and the triplet loss function value and feature matching accuracy on the validation dataset are calculated. Compare the validation loss function value of the current iteration with the validation loss function value of the previous iteration. If the validation loss function value does not decrease for several consecutive iterations, reduce the learning rate parameter and continue training. When the preset maximum number of iterations is reached or the value of the validation loss function is less than the preset threshold, the model training process is stopped and the current model parameters are saved as the final training result. The Triplet network model is configured using the model parameters from the final training results. Feature encoding is performed on the mixed feature data of the handwriting image to be identified and the known sample handwriting images to generate a first discriminative feature vector and a second discriminative feature vector, respectively.
8. The method according to any one of claims 1-5, characterized in that, The process involves inputting the hybrid feature dataset into a Triplet network model, and jointly optimizing and training the traditional handwriting feature set and the deep learning handwriting feature set using a triplet loss function to generate a first discriminative feature vector for the handwriting image to be identified and a second discriminative feature vector for the known sample handwriting images, including: Select triplet samples that meet the hard sample criteria from the mixed feature dataset. The hard sample criteria are sample combinations whose distance between different features is less than the sum of the distance between similar features and a preset margin value. The Triplet network model is trained using the difficult examples, and the model's ability to distinguish similar handwriting features is improved by increasing the training weights of the difficult examples. During model training, the selection threshold for difficult examples is dynamically adjusted. As the number of training iterations increases, the selection criteria for difficult examples are gradually lowered to expand the range of difficult examples. The importance of the traditional feature encoding vector and the deep learning feature encoding vector during the training process is evaluated, and the weight ratio of the two feature encoding vectors in the fusion process is dynamically adjusted through an attention mechanism. When the contribution of the traditional feature encoding vector to the loss function is higher than that of the deep learning feature encoding vector, the fusion weight of the traditional feature encoding vector is increased; when the contribution of the deep learning feature encoding vector exceeds a set contribution level, the fusion weight of the deep learning feature encoding vector is increased. By dynamically adjusting the weights, the traditional handwriting features and deep learning handwriting features are adaptively fused to generate a fused feature encoding vector with stronger discriminative power. Based on the optimization results of the fused feature encoding vector and the triplet loss function, high-dimensional feature representations of the handwriting image to be identified and the known sample handwriting images are extracted respectively, serving as the first discriminative feature vector and the second discriminative feature vector.
9. The method according to any one of claims 1-5, characterized in that, The process involves inputting the hybrid feature dataset into a Triplet network model, and jointly optimizing and training the traditional handwriting feature set and the deep learning handwriting feature set using a triplet loss function to generate a first discriminative feature vector for the handwriting image to be identified and a second discriminative feature vector for the known sample handwriting images, including: In the shared feature encoding layer of the Triplet network model, a feature normalization processing unit is introduced to perform L2 normalization processing on the traditional feature encoding vector and the deep learning feature encoding vector, so that the magnitude of the feature vector is unified to a preset constant. The distances between similar and dissimilar features are calculated using the normalized feature vectors. A variant of the triplet loss function with a temperature coefficient is constructed. By adjusting the temperature coefficient, the weight ratio of the distance between similar and dissimilar features is controlled, thereby enhancing the sensitivity of the loss function to differences in feature distances. When the difference between the distance of similar features and the distance of dissimilar features does not reach the preset difference range, the temperature coefficient is increased to improve the loss function value; when the difference reaches the preset difference range, the temperature coefficient is decreased to reduce the loss function value. During model training, the degree of aggregation of similar features and the degree of separation of dissimilar features are monitored in real time, and the feature learning effect of the model is evaluated by calculating the intra-class variance and inter-class distance of the feature vectors. When the intra-class variance is less than the preset variance threshold and the inter-class distance is greater than the preset distance threshold, the model training is deemed to have met the target, and the parameter optimization process is stopped. The trained model parameters are applied to the feature encoding process to encode the mixed feature data of the handwriting image to be identified and the known sample handwriting images, respectively, to generate a first discriminative feature vector and a second discriminative feature vector with normalization properties.
10. A handwriting image identification system, characterized in that, include: processor; A storage device having a computer program stored thereon, which, when executed by the processor, causes the processor to implement the handwriting image identification method based on artificial intelligence as described in any one of claims 1-9.
Citation Information
Cited By
OCR (Optical Character Recognition) image recognition method and system based on large model self-learning
CN121459369A