Signature handwriting identification method and system based on prompt learning, and storage medium
Patent Information
- Application Number
- CN202310635178.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-31
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-05-31
AI Technical Summary
[0006]有鉴于此,本发明针对现有技术跨膜态笔迹识别中存在的精度与泛化性不足,容易误判,准确率较低等问题,提出一种基于提示学习的签名笔迹鉴别方法,可以支持解决多模态泛场景下的签名笔迹鉴别问题
[0017]根据本申请另一方面,提供一种存储有计算机指令的非瞬时计算机可读存储介质,其中,所述计算机指令用于使所述计算机执行根据权利要求1-6中任一项所述的基于提示学习的签名笔迹鉴别方法。
Smart Images

Figure CN116645683B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer information processing technology and information security technology, specifically a signature handwriting identification method based on prompting learning. Background Technology
[0002] In the digital age, a large number of official documents, contracts, transactions, and authorizations rely on electronic signatures to verify identity, making handwriting authentication a crucial technology. Currently, electronic signature handwriting authentication is a very challenging task because handwriting can be collected under different devices, angles, lighting conditions, and writing styles. Traditional handwriting authentication methods primarily rely on manual feature extraction and measurement for comparison, which suffers from low accuracy, vulnerability to attacks, and high data volume requirements. Deep learning algorithms can automatically extract features and obtain good comparison results through end-to-end learning. However, conventional deep learning methods require a large amount of genuine and counterfeit signature data for training. Due to the privacy of this data and the difficulty of collecting handwriting imitation data, the overall richness of the training data is not high. Furthermore, conventional signature handwriting identification methods generally only focus on the feature information of the signature itself, without further considering the inherent attribute information of the signature or other prior knowledge. This results in limited algorithm accuracy and relatively slow training convergence speed. Although many algorithms now introduce prior knowledge to assist the model in learning quickly and accurately, most are based on manually designed prompt templates, which are relatively time-consuming, labor-intensive, and limited. In addition, most signature handwriting identification algorithms have relatively limited application scenarios, focusing more on handwriting comparison under the same modality and device type, and cannot support the identification of multimodal signature handwriting data in general scenarios. Therefore, there is an urgent need for a new, high-accuracy, stable signature handwriting identification method that can be used in complex multimodal scenarios to meet practical application needs.
[0003] Publication No. CN115601843A, entitled "Multimodal Signature Handwriting Identification System and Method Based on Two-Stream Network," discloses a multimodal signature handwriting identification system based on a two-stream network. This system collects paper and electronic signature data from different signatories to obtain offline and online signature image data, constructing corresponding orthographic and imitation samples of paper and electronic signature sequences. A recurrent generative adversarial network is used to migrate the signature images from one domain to another, and the same-modal data is combined and spliced to convert them into images of the same scale. The spliced two types of same-modal signature image data are input into the two-stream network for signature feature vectors obtained from a feature extraction network and a loss function for signature authenticity binary classification training. The system outputs probability scores for the similarity of signature image pairs. By training and optimizing the two-stream network with cross-modal orthographic and imitation sample signature handwriting pairs and the loss function, a multimodal signature handwriting identification model is obtained. This method uses modal transformation to convert cross-modal alignment into single-modal alignment, which can indeed improve the poor accuracy caused by modal differences. However, this method does not take into account the strong prior attribute information inherent in the actual data acquisition scenario, such as the acquisition device and writing type. Such features have a significant impact on handwriting style.
[0004] Publication number CN106803082A, titled "An Online Handwriting Recognition Method Based on Conditional Generative Adversarial Networks," describes a method that trains an adversarial network (AAN) on a handwriting signature dataset using category labels as conditions. This network generates corresponding directional numerical features based on the label conditions. The method mines personalized handwriting data from users through the conditional AAN, and authentication is achieved through an adversarial network signature discrimination model. This model captures the distribution of received text sample data and then adds category labels to the text samples, forming handwriting data samples of the specified categories. While this method first generates pseudo-signature data with data acquisition attributes such as writing style using the conditional AAN, and then uses a signature discrimination model to determine the authenticity of the signature, it utilizes strong prior attribute information to some extent. However, the prior attributes of a single user's handwriting input are too fixed. When encountering pseudo-signature attacks that exceed the defined range in real-world scenarios, misjudgments may occur.
[0005] Existing cross-membrane handwriting recognition methods do not fully consider the strong prior attribute information inherent in the data acquisition scenario that greatly affects the writing style, nor do they fully consider the diversity and variability of handwriting prior attributes. As a result, they are prone to misjudgment in cross-membrane handwriting recognition and have low accuracy. Summary of the Invention
[0006] In view of this, this invention addresses the problems of insufficient accuracy and generalization, easy misjudgment, and low accuracy in existing cross-membrane handwriting recognition technologies. It proposes a signature handwriting identification method based on cue learning, which can support the solution of signature handwriting identification problems in multimodal and general scenarios.
[0007] This approach combines signature image data with prior knowledge of handwriting characteristics, such as the writing content, style, writing method, and writing device, to serve as input for a signature handwriting comparison learning pre-training architecture. This enables the signature image feature extractor to quickly and accurately learn the personalized signature features of the signer, thus improving subsequent signature authentication. Unlike traditional cue-based comparison learning methods, this approach eliminates the need for manually designing cue vectors as input to the attribute feature extractor. The automatically extracted and adaptively adjustable signature attribute cue vector learning method not only saves time-consuming manual design steps but also more accurately and stably describes the personalized features of the signature handwriting, further improving comparison accuracy. Simultaneously, relevant signature attribute information obtained during the signature handwriting collection process is introduced as a cue into the pre-trained model for comparative learning. This approach addresses two main concerns: firstly, the significant impact of cross-device and cross-writing style attributes on signature handwriting style changes, mitigating the model's insufficient generalization ability; and secondly, the inherent features of signatures, such as writing type and content, which promote handwriting comparison feature learning, accelerating the model's learning speed and serving as prior knowledge to guide the pre-trained model towards better learning outcomes. Furthermore, thanks to the pre-trained model, this invention eliminates the need to collect large amounts of one-to-one signature imitation data during the signature handwriting comparison model training phase, further reducing data collection cycle and cost.
[0008] According to one aspect of this application, a signature handwriting identification method based on cue learning is proposed. This method involves collecting signature images and / or data sequences containing signatures from different signers using different writing devices, writing methods, and calligraphic styles, and recording the attribute information of the signatures. Different types of signatures are converted into signature images with fixed formats and sizes. The labels corresponding to the attributes of the signature images are confirmed, and classification training data, pre-training data, and fine-tuning data for different attributes are obtained. Multiple attribute feature extractors are trained, setting the signer's identity as the target value, and each feature extractor extracts cue feature vectors corresponding to each attribute of the signature image. Based on the signature image data and its corresponding attribute cue vector sequence, a multimodal signature comparison learning pre-trained model is trained, adaptively adjusting the attribute cue vector sequence. The multimodal signature comparison learning pre-trained model is adapted to the signature handwriting authenticity identification task. Fine-tuning data is input to the image feature extractor to extract signature feature vectors, and then training is performed based on the signature feature vector comparison learning to obtain a signature handwriting identifier.
[0009] Further preferred, the training of multiple attribute feature extractors includes: training a multi-classification model for writing devices, a binary classification model for writing methods, a multi-classification model for writing content, and a multi-classification model for writing styles based on different signature attribute information and their corresponding different attribute category labels, to obtain a feature extractor for writing devices, a feature extractor for writing styles, a feature extractor for writing methods, and a feature extractor for writing content.
[0010] Further optimization involves extracting feature vectors that include: feature representations of the writing device, writing method, script style, and writing content of the signature handwriting image. Based on contrastive learning pre-trained signature image data and each attribute feature extractor, feature vectors corresponding to the pre-trained data are extracted respectively. Each feature vector is concatenated and combined in the order of writing device vector, writing method vector, writing style vector, writing content vector, and {class} to obtain the attribute prompt vector sequence <writing device><writing method><writing style><writing content>{class}. Here, the attributes in <> are filled with corresponding descriptions according to the signature attribute labels. class represents the target signer identity label category to be predicted. The output vector dimension is 1*1.
[0011] Further optimization confirms that the tags corresponding to the signature image attributes include: setting tags for data types; each signature data tag includes writing device, writing method, writing style, writing content, signature authenticity, and signer identity; the writing device tag is a four-category tag; the writing method tag is a three-category tag; the writing style tag is a three-category tag; the writing content corresponds to the signer's name and is a multi-category tag; the signature authenticity tag is a two-category tag; and the signer identity tag is a multi-category tag, with each signer having a specific type of representation.
[0012] Further optimization involves obtaining four-class classification training data for writing devices, three-class classification training data for writing methods, three-class classification training data for writing styles, and N-class classification training data for writing content based on signature image data and their corresponding attribute labels. Based on this training data, four-class classification models for writing devices, three-class classification models for writing methods, three-class classification models for writing styles, and N-class classification models for writing content are trained respectively. The backbone network of all classification models uses ResNet50, and the dimensions of the classification layers are uniformly set. Corresponding feature extractors for writing devices, writing methods, writing styles, and writing content are obtained. Based on these feature extractors, feature vectors for writing devices, writing methods, writing styles, and writing content from the comparative learning pre-trained signature data are extracted. The feature vector output by each feature extractor represents a fixed dimension.
[0013] Further optimization involves pre-training multimodal signature comparison learning based on a dual-tower network structure. The image feature extractor branch extracts signature image features, with its backbone network using a residual network structure (ResNet101). The attribute feature extractor branch extracts attribute hint feature sequences corresponding to the signature image, with its backbone network using an attention-based transformer architecture. After feature extraction based on the attribute hint feature sequences, it outputs attribute feature vectors. Combining the implicit signature attribute information contained in the attribute hint feature sequences, the feature alignment stage of the comparison learning training process ignores the differences in signatures across different modalities, devices, and writing styles, focusing instead on learning the signature style features unique to the signer. The genuine and counterfeit signature data of some forged signers are used as fine-tuning data for multimodal comparison learning, while the remaining data serves as pre-training data. The similarity between the output feature vectors of the image feature extractor is evaluated using cosine similarity loss calculation, and this cosine similarity loss is further optimized until convergence.
[0014] According to another aspect of this application, a signature handwriting identification system based on cue learning is also proposed, comprising: a data acquisition part, a data preprocessing part, a writing attribute feature extractor, an attribute cue vector sequence acquisition part, and a multimodal pre-training part. The data acquisition part collects paper images and electronic signature sequence data containing signatures from different signers using different acquisition devices, with different writing methods and styles, and simultaneously records relevant attribute information of the signature data, and collects imitation signature data corresponding to some signatures. The data preprocessing part converts different types of signatures into signature images with fixed formats and sizes, confirms the corresponding attributes and labels of the signature images, acquires and determines training signature image data with different attributes, and acquires a number of pre-trained signature images. The training data for the signature authenticity detector is fine-tuned. The handwriting attribute feature extractor trains multiple attribute feature extractors based on training signature image data with different attributes and corresponding attribute category labels, setting the signer's identity as the target value. Each feature extractor extracts the hint feature vector corresponding to each attribute of the signature image. For the attribute hint vector sequence acquisition part, a multimodal signature comparison learning pre-trained model is trained based on the signature image data and its corresponding attribute hint vector sequence, adaptively adjusting the attribute hint vector sequence. In the multimodal pre-training part, adapted to the signature handwriting authenticity identification task, the signature feature vector extracted from the input image feature extractor is fine-tuned, and then the signature handwriting detector is obtained through comparison learning training based on the extracted signature feature vectors.
[0015] Further optimization involves using a dual-tower network structure for the multimodal signature comparison learning pre-training. Specifically, the image feature extractor branch performs comparison learning pre-training to extract signature image features, with the backbone network using a ResNet101 network structure. The attribute feature extractor branch extracts attribute hint feature sequences corresponding to the signature image, with the backbone network using a Transformer architecture. After feature extraction based on the attribute hint feature sequences, it outputs attribute feature vectors. Combining the implicit signature attribute information contained in the attribute hint feature sequences, the feature alignment stage during comparison learning ignores the differences in signatures across different modalities, devices, and writing styles, focusing instead on learning the signature style features unique to the signer. The genuine and counterfeit signature data of some forged signatures are used as fine-tuning data for multimodal comparison learning, while the remaining data serves as pre-training data. The similarity between the output feature vectors of the image feature extractor is evaluated using cosine similarity loss calculation, and this cosine similarity loss is further optimized until convergence.
[0016] Further preferred, the training of multiple attribute feature extractors includes: training a multi-classification model for writing devices, a binary classification model for writing methods, a multi-classification model for writing content, and a multi-classification model for writing styles based on different signature attribute information and their corresponding different attribute category labels, to obtain a feature extractor for writing devices, a feature extractor for writing styles, a feature extractor for writing methods, and a feature extractor for writing content.
[0017] According to another aspect of this application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the signature handwriting identification method based on prompting learning according to any one of claims 1-6.
[0018] This invention employs a cue-based learning method to further utilize signature-related attribute information combined with the original signature image for comparative learning, thereby more accurately describing the characteristics of the signature handwriting and improving the accuracy of handwriting identification. Compared with traditional signature handwriting comparison algorithms, it does not require a large amount of one-to-one corresponding positive imitation signature data for signature comparison model training and easily avoids getting trapped in local optima. Furthermore, unlike traditional cue-based comparative learning training methods, this algorithm no longer requires manually designing cue vectors as input to the attribute feature extractor. Instead, it proposes an automatically extractable and adaptively adjustable signature attribute cue vector learning method. This method not only saves time-consuming and laborious manual design steps but also more accurately and stably describes the personalized characteristics of the signature handwriting, further improving the comparison accuracy. Therefore, this method has high practical value. The application scenarios of this invention are wide-ranging, and it can be used for signature handwriting comparison and recognition in fields such as finance, law, and government, as well as for identity verification in the field of network security.
[0019] This invention supports the solution of signature handwriting identification problems in multimodal and generalized scenarios. Addressing the insufficient accuracy and generalization of current signature handwriting identification algorithms, this patent proposes a cue-based signature handwriting identification algorithm. This algorithm combines signature image data with prior knowledge of handwriting content, style, writing method, and writing device—all containing strong identity information—as input to a signature handwriting comparison learning pre-training architecture. This facilitates the signature image feature extractor to quickly and accurately learn the personalized signature features of the signer, thereby improving subsequent signature handwriting authenticity verification. Furthermore, unlike traditional cue-based comparison learning training methods, this algorithm eliminates the need for manually designing cue vectors as input to the attribute feature extractor. Instead, it proposes an automatically extractable and adaptively adjustable signature attribute cue vector learning method. This method not only saves time-consuming and laborious manual design steps but also more accurately and stably describes the personalized features of the signature handwriting, further improving comparison accuracy. By incorporating relevant signature attribute information obtained during the signature handwriting collection process into the pre-trained model for comparative learning, this approach addresses two key issues. First, it considers the significant impact of cross-device and cross-writing style attributes on signature handwriting style changes, mitigating the model's insufficient generalization ability. Second, it recognizes that attributes such as writing type and content, as inherent signature features, promote handwriting comparison feature learning, accelerating the model's learning speed and serving as prior knowledge to guide the pre-trained model towards better learning outcomes. Furthermore, thanks to the pre-trained model, this invention eliminates the need to collect large amounts of one-to-one signature imitation data during the signature handwriting comparison model training phase, further reducing data collection time and costs. Attached Figure Description
[0020] Figure 1 A schematic diagram of the signature handwriting identification method based on prompting learning in an exemplary embodiment of this application; Figure 2 A schematic diagram of the multimodal pre-trained network design structure in an exemplary embodiment of this application; Figure 3 A schematic diagram of signature authenticity verification device training in an exemplary embodiment of this application; Figure 4 The diagram shown is a structural block diagram of an exemplary electronic device that can be used to implement embodiments of this application. Detailed Implementation
[0021] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.
[0022] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.
[0023] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this application are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0024] The terms “a” and “a plurality” used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as “one or more”.
[0025] The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0026] This invention proposes a signature handwriting identification system based on prompting learning. It collects paper images and electronic signature sequence data containing signatures from different signers using different acquisition devices, writing methods, and calligraphic styles. Simultaneously, it records and saves attribute information such as the writing device, writing method, calligraphic style, and writing content of the relevant signature data. It also collects imitation signature data corresponding to some signatures. The paper signature data and electronic signature sequence data are processed, mainly including: extracting paper signature images; data type conversion, uniformly converting different types of electronic signatures into signature images with fixed formats and sizes; cleaning and filtering out illegible, poor-quality signature data that does not conform to actual signing scenarios; confirming the labels of the corresponding attributes of the signature images; obtaining data attribute-related labels; obtaining and determining the training signature image data for different attributes; and comparing the pre-trained signature image data with the signature authenticity detector for fine-tuning training. The process involves: 1) Signature image data; 2) Training of feature extractors for each writing attribute: Based on training signature image data for different attributes and their corresponding attribute category labels, training multi-classification models for writing devices, writing methods, writing content, and writing styles, respectively, to obtain feature extractors for writing devices, writing methods, writing content, and writing styles; 3) Obtaining attribute hint vector sequences: Based on pre-trained signature image data using contrastive learning and each attribute feature extractor, extracting feature vectors corresponding to the pre-trained data, and concatenating these feature vectors in the order of writing device vector, writing method vector, writing style vector, writing content vector, and {class} from left to right to obtain the attribute hint vector sequence; 4) Multimodal pre-training: Based on pre-trained signature image data using contrastive learning and its corresponding attribute hint vector sequence, training a multimodal signature contrastive learning pre-trained model. One branch of the pre-trained model takes the signature image as input, and the other branch takes the attribute hint feature vector sequence as input. Upon completion, the image feature extractor and attribute feature extractor are obtained. The image feature extractor backbone network can use a convolutional neural network (CNN), and the attribute sequence feature extraction backbone network can use a transformer architecture.
[0027] The signature authenticity detector training is based on the pre-trained model for subsequent signature authenticity identification and classification tasks. The signature image data obtained above for fine-tuning the signature authenticity detector training is input into the image feature extractor for feature extraction. Then, the extracted features are compared for feature similarity, and the signature authenticity detector is obtained based on contrastive learning.
[0028] To further elaborate on the technical solution of this method, the present invention will be further described in conjunction with specific embodiments and accompanying drawings.
[0029] like Figure 1The diagram illustrates the overall process of the signature handwriting identification method based on cue learning in an exemplary embodiment of this application. It includes: a data acquisition module that collects multimodal signature data across devices, writing styles, and genres; a data preprocessing module that standardizes the multimodal signature data, including paper signature image extraction, data type conversion, label acquisition, and dataset segmentation; an attribute cue feature sequence acquisition module that trains different attribute feature extractors based on the signature attributes of the multimodal data and obtains feature vectors corresponding to the signature images, combining the {class} flag to construct an attribute cue feature vector sequence; a multimodal pre-training module that pre-trains the multimodal signature data model based on a contrastive learning method; a signature authenticity detector training module that trains the authenticity classification model of the signature authenticity detector; and a testing and evaluation module that tests and evaluates the model.
[0030] The data acquisition module collects and gathers paper documents and corresponding electronic signature sequence data related to the signatures of different signers using different writing devices, media, writing methods, and calligraphic styles.
[0031] Different media include paper-based documents (not limited to contracts, forms, and other documents) and data collection from various electronic devices (not limited to mobile phones, tablets, and signature pads). Different writing methods include touchscreen, pen, and electronic stylus writing; different handwriting styles include regular script, running script, and cursive script, etc. The signatures of the signatory on different media, devices, and with different writing styles are all their own names. During the data collection process, the writing device, writing method, writing style, and related attributes of the written content of the signature data must be recorded synchronously and saved in the corresponding data list format.
[0032] The above data are all signatures written in the correct form. A portion of the signatories (about 10%) were selected and their signatures were copied by observation and imitation. The corresponding copy data was collected for subsequent signature identification and classification training.
[0033] The data preprocessing module preprocesses the collected data, mainly including: extraction of paper signatures, data type conversion, acquisition of data attribute-related tags, and data segmentation. Specifically, it includes: Paper signature extraction. This involves detecting and matting paper signature data. Generally, paper data such as contracts, forms, and documents have relatively complex and interfering backgrounds, and their appearance is completely different from electronic signature sequences. Therefore, signature detection algorithms and signature matting algorithms are needed to detect and matte the regions of paper signatures. Open-source models such as DBNet and SegNet can be used for detection and matting preprocessing of paper signature data to obtain corresponding paper signature mask image representations.
[0034] Data type conversion. Signature data from different sources and writing devices are uniformly converted into a single-modal image data representation. Based on the statistical distribution of normal signature pixel size and aspect ratio, a uniform fixed size is set (e.g., the optimal standard width and height scale for conventional signatures is set to 256*128 based on statistical analysis). Signature mask images extracted from different media, such as paper signature data, can be binarized, denoised, and scaled to obtain uniformly sized paper-based binarized signature image data. Electronic signature sequences acquired from different acquisition devices, writing styles, and writing methods can be upsampled or downsampled at a fixed sampling rate (e.g., a sampling rate of 50%) before being displayed as a signature image representation of the aforementioned fixed size.
[0035] Obtain data attribute-related tags. Set tags for the data types of the preprocessed data. For example, each signature data tag should include six tags: writing device, writing method, writing style, writing content, signature authenticity, and signatory identity. For instance, the writing device can include types such as paper, mobile phone, tablet, and signature board, with corresponding four-category tags {0, 1, 2, 3}; the writing method includes three types: touchscreen, pen, and electronic touchscreen pen, with corresponding three-category tags {0, 1, 2}; the writing style includes three types: regular script, running script, and cursive script, with corresponding three-category tags {0, 1, 2}; and the writing content is the signatory's surname. A name can generally be represented by the content of a signature. However, due to the possibility of duplicate names, signatories with the same name need to be represented by discrete numbers. This is because there may be signatories with the same name, and the signature content does not fully represent their true identity. For example, if there are 1000 signatories, their corresponding written content can be categorized into labels {0, ..., N}, where N represents the different types of signature content among the 1000 people, and N < 999. The authenticity of the signature is set to two categories: true or false, with corresponding labels {0, 1}. The identity of a signator must be unique, so each signator's identity needs to have a unique type representation. For example, if there are 1000 signatories, their corresponding signator identity labels can be {0, ..., 999}.
[0036] Data segmentation involves classifying training data by different attributes, obtaining pre-training data for contrastive learning, and fine-tuning data for handwriting authenticity identification classification. Based on the preprocessed signature image data, we can first obtain classification training data regarding writing device, writing method, writing style, and writing content, as well as the authenticity of the data, according to the label information of the signature data. The authentic and forged signature data of some signatories who were forged are used as fine-tuning data, while the remaining data is used as pre-training data for multimodal contrastive learning. If only 100 people participated in the collection of forged signature data, the signature data corresponding to these 100 people with multi-class labels are retained as fine-tuning data.
[0037] Attribute hint feature sequence acquisition module. Since the main purpose of this invention is to identify signature handwriting, ultimately recognizing the signer's identity information, the target value is set as the signer's identity (signer's name). Based on the pre-processed signature image data and its corresponding attribute label information, four-classification training data for writing devices, three-classification training data for writing methods, three-classification training data for writing styles, and N-classification training data for writing content can be obtained. Based on these four types of training data, four-classification models for writing devices, three-classification models for writing methods, three-classification models for writing styles, and N-classification models for writing content are trained respectively. The backbone network of all classification models can use ResNet50, and the dimension of the final classification layer is uniformly set (e.g., uniformly set to 1*64), obtaining the corresponding writing device feature extractor, writing method feature extractor, writing style feature extractor, and writing content feature extractor. Based on the above four feature extractors, feature vectors for writing device, writing method, writing style, and writing content are extracted from the comparative learning pre-trained signature data. The feature vector output by each feature extractor represents a fixed dimension (all 1*64). Therefore, the arrangement of the attribute hint feature sequence can be referenced in the following example: <Writing Device><Writing Method><Writing Style><Writing Content>{class} In the above expression, the contents within <> are the feature vector representations of the corresponding attributes to be filled. The attributes within <> can be filled in according to this template based on the obtained signature attribute labels. The contents within {} do not need to be filled. Here, class represents the target signer identity label category that the model needs to learn to predict. The final output vector dimension is 1*1.
[0038] Following the left-to-right order in the example above, the attribute feature vectors are merged with the target category to be predicted. This yields the final attribute hint feature vector sequence corresponding to the signature image data, with an output dimension of 1*257. This allows the generation of the attribute hint feature sequence corresponding to the pre-trained signature image data for subsequent multimodal contrastive learning.
[0039] Figure 2 The diagram shown is a schematic diagram of the multimodal pre-trained network design structure in an exemplary embodiment of this application.
[0040] Multimodal pre-training module. Based on the contrastive learning pre-trained signature image data and its corresponding attribute hint feature sequences obtained above, this embodiment of the invention adopts a CLIP-based dual-tower network structure for multimodal signature contrastive learning pre-training. The dual-tower network structure has two network branches, among which the image feature extractor branch mainly performs feature extraction of the contrastive learning pre-trained signature images. The backbone network of this branch can use the classic ResNet101 network. For example, if the input signature image data is 256*128, after passing through this network, it will output a 1*512 dimensional feature vector representation.
[0041] The attribute feature extractor branch primarily extracts features from the attribute cue feature sequence corresponding to the signature image. The backbone network of this branch can use a transformer architecture, outputting an attribute feature vector representation after feature extraction from the attribute cue feature sequence. For example, if a 1*257 dimensional attribute cue feature sequence is input, after feature extraction through a temporal transformer structure, a 1*512 dimensional attribute feature vector representation is output.
[0042] The purpose of using this dual-tower structure for learning different modal data is to combine the implicit signature attribute information contained in the attribute hint feature sequence. This can further promote the signature image branch to ignore the differences in signatures under different modalities, different devices and different writing styles during the feature alignment stage of the contrastive learning training process, so that the model can focus more on learning the signature style features unique to the signer.
[0043] Multimodal signature comparison learning is trained using pre-trained data. Since the multimodal input consists of one-to-one correspondence between signature image data and attribute hint feature sequences, a significant feature similarity can be assumed between them. Therefore, it can also be approximated that the similarity of their respective 1*512-dimensional feature vectors obtained through different feature extractors should also be significant. An adaptive adjustment method for the attribute hint feature sequences is employed during training. Although the attribute hint feature sequences corresponding to the signature images are fixed at input, it is difficult to guarantee completely accurate feature extraction, nor can it be guaranteed that the four attribute types can fully represent the signature style and habitual characteristics of the signer. Therefore, by continuously updating and correcting the attribute hint feature sequences through adaptive adjustment during training, the image feature extractor can be further improved. The similarity between the network output feature vectors can be evaluated using cosine similarity loss calculation, and this loss can be further optimized until convergence. This leads to the acquisition of the image feature extractor and the attribute feature extractor.
[0044] Figure 3 The diagram shown is a schematic diagram of the signature authenticity detector training in an exemplary embodiment of this application.
[0045] The signature authentication discriminator training module adapts the image feature extractor from the aforementioned multimodal pre-trained model for subsequent signature authentication classifier tasks. It trains the feature extractor using a positively copied signature image as input, optimizing the cosine similarity loss function until convergence.
[0046] First, based on the data (e.g., 100) obtained above, which contains both genuine and counterfeit signatures for fine-tuning the signature authenticity detector, a training list of 1vs1 comparison signature image pairs is constructed. Signature image features are extracted from the two sets of data in the image pairs respectively. Then, the extracted feature vectors are subjected to comparative learning training. The training label is the signer identity category {0, 1} previously labeled in the data preprocessing. Specifically, if the image pairs in the 1vs1 comparison are signed by the same person, the label is 1, and otherwise it is 0. The loss function used for training can also be the classic cosine similarity loss. The training objective is to optimize this loss until it converges. Finally, an image feature extraction model with good image feature extraction capabilities for both genuine and counterfeit data can be obtained. This model can be used as the final signature authenticity detector.
[0047] Testing and Evaluation Module. Further, after training the aforementioned signature authentication and comparison learning model, the signature authentication effector is tested and evaluated to determine its performance on the signature authentication task. A portion of the pre-processed fine-tuned dataset is used as the test set, which is then input into the trained signature authentication detector for testing. Specifically, the following method can be employed: First, construct 1vs1 signature image pairs and input them into the signature authentication detector for signature feature extraction. Finally, calculate the cosine similarity between the two feature vectors. If the similarity score exceeds a threshold (e.g., 0.5), the input signature image pair is considered to have been signed by the same person; otherwise, it is considered a forgery. The model's recall, precision, and F1 score are calculated to evaluate its performance in different scenarios.
[0048] In actual testing, the signature authenticity detector using this implementation scheme performed well.
[0049] This invention first requires collecting signature handwriting images and text descriptions. Then, a deep learning model is used to extract feature representations. An attribute hint learning module reduces the differences between different modalities. Pre-trained feature extractors for different attributes further accelerate model learning. Furthermore, a training strategy based on adaptively adjusted attribute hint vectors significantly improves the model's accuracy. Finally, a matching module is used to compare and recognize the signature handwriting to verify its authenticity.
[0050] Compared with traditional signature handwriting comparison methods, (1) it can effectively handle the problem of signature handwriting identification and improve the accuracy of signature handwriting comparison; (2) the automatic extraction and adaptive adjustment prompt learning method can combine signature attributes to further enhance the model's feature learning for different scenarios and conditions, improve the model's generalization ability, expand the application scenarios and comparison accuracy; (3) it does not require a large amount of one-to-one corresponding positive imitation signature data for signature comparison model training, effectively saving related data collection costs and training costs.
[0051] This invention can support the solution of signature handwriting identification problems in multimodal and generalized scenarios. Addressing the insufficient accuracy and generalization of current signature handwriting identification algorithms, this application combines signature image data with prior knowledge such as the writing content, style, writing method, and writing device that contain strong identity information, as input to a signature handwriting comparison learning pre-training architecture. This facilitates the signature image feature extractor to quickly and accurately learn the personalized signature features of the signer, thereby better enabling subsequent signature handwriting authentication tasks. Furthermore, unlike traditional cue-based contrastive learning training methods, this invention eliminates the need for manually designed cue vectors as input to the attribute feature extractor. Instead, it proposes an automatically extractable and adaptively adjustable signature attribute cue vector learning method. This method can more accurately and stably describe the personalized features of signature handwriting, further improving the comparison accuracy. Simultaneously, relevant signature attribute information obtained during signature handwriting collection is introduced as cue into the contrastive learning pre-training model. This addresses two issues: firstly, the significant impact of cross-device and cross-writing style attributes on signature handwriting style changes, thus mitigating insufficient model generalization; and secondly, the inherent features of handwriting type and content as signatures promote handwriting comparison feature learning, accelerating the model's learning speed and serving as prior knowledge to guide the pre-trained model to achieve better learning results. Moreover, due to the assistance of the pre-trained model, this invention eliminates the need to collect large amounts of one-to-one correspondence signature imitation data during the signature handwriting comparison model training phase, further reducing the data collection cycle and cost.
[0052] In summary, this invention proposes a signature handwriting identification method based on cue learning, which has the advantages of wide application and significant effect. It is applicable to the comparison and identification of signature handwriting in the fields of finance, law, and government, and can also be applied to identity verification in the field of network security.
[0053] like Figure 4The diagram shows a structural block diagram of an exemplary electronic device that can be used to implement embodiments of this application. The electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of the device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0054] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, output unit 307, storage unit 308, and communication unit 309. Input unit 306 can be any type of device capable of inputting information to electronic device 300. Input unit 306 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 307 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 308 may include, but is not limited to, disk and optical disk. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0055] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above. For example, in some embodiments, the reconstruction and decomposition of the muscle movement trajectory based on the original trajectory of the signature stroke, and the decomposition of its logarithmic velocity curve, can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. In some embodiments, the computing unit 301 can be configured in any other suitable manner to perform signature handwriting comparison and verification methods.
[0056] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0057] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0058] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0059] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0060] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0061] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
Claims
1. A signature handwriting identification method based on cue learning, characterized in that, The process involves collecting signature images and / or data sequences from different signers using different writing devices, writing methods, and writing styles, and recording the signature attribute information; converting different types of signatures into signature images with fixed formats and sizes; confirming the labels corresponding to the attributes of the signature images, and obtaining classification training data, pre-training data, and fine-tuning data for different attributes; training multiple feature extractors, setting the signer's identity as the target value, and having each feature extractor extract the cue feature vectors corresponding to each attribute of the signature image; and training a multimodal signature contrastive learning pre-trained model based on the signature image data and its corresponding attribute cue feature vector sequences, adaptively adjusting the attribute cue feature vector sequences. A multimodal signature comparison learning pre-trained model is adapted to the signature handwriting authenticity identification task. The signature feature vector is extracted from the data input image feature extractor and then trained based on the signature feature vector comparison learning to obtain a signature authenticity detector. The confirmed signature image attribute corresponding labels include: setting labels for data types. The labels for each signature data include writing device, writing method, writing style, writing content, signature authenticity, and signer identity. The writing device label is a four-category label, the writing method is a three-category label, the writing style is a three-category label, the writing content corresponds to the signer's name and is a multi-category label, the signature authenticity is a two-category label, and the signer identity is a multi-category label corresponding to each signer having a specific type representation. Based on signature image data and its corresponding label information of different attributes, training data for four categories of writing devices, three categories of writing methods, three categories of writing styles, and N categories of writing content are obtained. Based on the above training data, four-category models for writing devices, three-category models for writing methods, three-category models for writing styles, and N-category models for writing content are trained respectively. The backbone network of the classification models all adopts ResNet50, and the dimension of the classification layer is uniformly set. Corresponding feature extractors for writing devices, writing methods, writing styles, and writing content are obtained. Based on the above feature extractors, prompt feature vectors for writing devices, writing methods, writing styles, and writing content are extracted from the pre-trained signature data of the comparative learning. The prompt feature vector output by each feature extractor represents a fixed dimension. The data of the forged signatures, some of which are genuine and fake, are used as fine-tuning data to train the signature authenticity detector, while the rest of the data is used as pre-training data. Multimodal signature comparison learning pre-training is performed based on a dual-tower network structure. The image feature extractor branch extracts signature image features, with its backbone network using a ResNet101 network structure. The attribute feature extractor branch extracts features from the attribute hint feature vector sequence corresponding to the signature image, with its backbone network using a Transformer architecture. After feature extraction based on the attribute hint feature vector sequence, it outputs attribute feature vector representations. Combining the implicit signature attribute information contained in the attribute hint feature vector sequence, the feature alignment stage of the comparison learning training process ignores the differences in signatures under different writing styles, writing devices, and writing methods, focusing instead on learning the signature style features unique to the signer. Multimodal signature comparison learning training is performed using the pre-trained data. This strategy of adaptively adjusting the attribute hint feature vector sequence during training continuously updates and corrects the data, promoting better learning of the image feature extractor. The similarity between the output feature vectors of the image feature extractor is evaluated using cosine similarity loss calculation, and this cosine similarity loss is further optimized until convergence, resulting in the image feature extractor and the attribute feature extractor. Adaptation and transfer of image feature extractors from multimodal pre-trained models to subsequent signature identification classifier tasks.
2. The method according to claim 1, characterized in that, Based on the pre-trained signature image data from contrastive learning and each feature extractor, the prompt feature vectors corresponding to the pre-trained data are extracted respectively. The prompt feature vectors are concatenated and combined in the order of writing device vector, writing method vector, writing style vector, writing content vector, and {class} to obtain the attribute prompt feature vector sequence <writing device><writing method><writing style><writing content>{class}. Among them, the attributes in <> are filled with corresponding descriptions according to the signature attribute labels. class represents the target signer identity label category to be predicted. The output vector dimension is 1*1.
3. A signature handwriting identification system based on cue learning, characterized in that, include: The system includes a data acquisition section, a data preprocessing section, a feature extractor for writing attributes, a feature vector sequence acquisition section for attribute hints, and a multimodal pre-training section. The data acquisition section collects signature images and / or data sequences from different signers using different writing devices, writing methods, and writing styles, and records the attribute information of the signatures. In the data preprocessing section, different types of signatures are converted into signature images with fixed formats and sizes. The labels corresponding to the attributes of the signature images are confirmed, and classification training data, pre-training data, and fine-tuning data for different attributes are obtained. For the writing attribute feature extractor, multiple feature extractors are trained, with the signer's identity set as the target value. Each feature extractor extracts the cue feature vector corresponding to each attribute of the signature image. In the feature vector sequence acquisition part, based on the signature image data and its corresponding attribute hint feature vector sequence, a multimodal signature comparison learning pre-trained model is trained, and the attribute hint feature vector sequence is adaptively adjusted; in the multimodal pre-training part, the multimodal signature comparison learning pre-trained model is adapted to the signature handwriting authenticity identification task, the signature feature vector extracted from the data input image feature extractor is fine-tuned, and then the signature authenticity detector is obtained based on the signature feature vector comparison learning training. The confirmed signature image attribute corresponding labels include: setting labels for data types. The labels for each signature data include writing device, writing method, writing style, writing content, signature authenticity, and signer identity. The writing device label is a four-category label, the writing method is a three-category label, the writing style is a three-category label, the writing content corresponds to the signer's name and is a multi-category label, the signature authenticity is a two-category label, and the signer identity is a multi-category label corresponding to each signer having a specific type representation. Based on signature image data and its corresponding label information of different attributes, training data for four categories of writing devices, three categories of writing methods, three categories of writing styles, and N categories of writing content are obtained. Based on the above training data, four-category models for writing devices, three-category models for writing methods, three-category models for writing styles, and N-category models for writing content are trained respectively. The backbone network of the classification models all adopts ResNet50, and the dimension of the classification layer is uniformly set. Corresponding feature extractors for writing devices, writing methods, writing styles, and writing content are obtained. Based on the above feature extractors, prompt feature vectors for writing devices, writing methods, writing styles, and writing content are extracted from the pre-trained signature data of the comparative learning. The prompt feature vector output by each feature extractor represents a fixed dimension. The data of the forged signatures, some of which are genuine and fake, are used as fine-tuning data to train the signature authenticity detector, while the rest of the data is used as pre-training data. Multimodal signature comparison learning pre-training is performed based on a dual-tower network structure. The image feature extractor branch extracts signature image features, with its backbone network using a ResNet101 network structure. The attribute feature extractor branch extracts features from the attribute hint feature vector sequence corresponding to the signature image, with its backbone network using a Transformer architecture. After feature extraction based on the attribute hint feature vector sequence, it outputs attribute feature vector representations. Combining the implicit signature attribute information contained in the attribute hint feature vector sequence, the feature alignment stage of the comparison learning training process ignores the differences in signatures under different writing styles, writing devices, and writing methods, focusing instead on learning the signature style features unique to the signer. Multimodal signature comparison learning training is performed using the pre-trained data. This strategy of adaptively adjusting the attribute hint feature vector sequence during training continuously updates and corrects the data, promoting better learning of the image feature extractor. The similarity between the output feature vectors of the image feature extractor is evaluated using cosine similarity loss calculation, and this cosine similarity loss is further optimized until convergence, resulting in the image feature extractor and the attribute feature extractor. Adaptation and transfer of image feature extractors from multimodal pre-trained models to subsequent signature identification classifier tasks.
4. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, in, The computer instructions are used to cause the computer to execute the signature handwriting identification method based on prompting learning according to any one of claims 1-2.
Citation Information
Patent Citations
Conditional generative adversarial network-based online handwriting identification method
CN106803082A
Intelligent training method and device fusing semantics and image features, equipment and medium
CN113449821A
Multi-modal signature handwriting identification system and method based on double-flow network
CN115601843A