Handwriting comparison method and device based on multi-modal characteristics, equipment and medium
By constructing a multimodal feature handwriting comparison method, using stroke texture, curve and neural network features, combined with the training of classification loss function and triple metric loss function, the problems of low accuracy and insufficient interpretability of handwriting ratio are solved, and higher discrimination and accuracy are achieved.
Patent Information
- Application Number
- CN202510146817.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-08
AI Technical Summary
In the prior art, the accuracy of the word mark is low and the lack of interpretability, resulting in insufficient distinction of feature when distinguishing word marks.
Using a multimodal feature-based handwriting comparison method, by constructing multiple triplets, the stroke texture features, curve features and neural network features of each image in each triplet are extracted, and these features are fused to generate multimodal features. Then, the initial model is constructed using the encoder and trained through the classification loss function and the triple metric loss function to obtain the word comparison model. Finally, the division threshold is determined based on the image feature expression output by the model, and it is determined whether there is a risk of copying the handwriting to be compared.
It improves the accuracy and distinction ability of handwriting comparison, solves the problem of insufficient feature distinction, and improves the interpretability of the model by integrating multiple features, especially when the data volume is insufficient.
Smart Images

Figure CN119992567A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and medical health technology, and in particular to a handwriting comparison method, device, equipment and medium based on multimodal features. Background Art
[0002] Handwriting comparison is a common authentication method, generally used to verify identity, such as contract signing and document identification in financial scenarios, procurement contract signing and family signatures in medical scenarios, etc. Accurate handwriting comparison can effectively reduce the economic losses and security risks caused to users by handwriting copying or forgery. In addition, handwriting, as an effective supplement to biometrics, is also of great value in verifying whether it is the real person.
[0003] However, in the existing technology, handwriting comparison still has problems such as difficulty in identification. Because the characteristics of handwriting are unstable, the distance within a class is often greater than the distance between classes. In addition, due to the limitations of manual experience and the improvement of copying methods, traditional features are often not enough to cover the cases in the scene in terms of discrimination.
[0004] Therefore, how to provide a handwriting comparison method that takes into account both explainability and strong distinguishing ability has become an urgent problem to be solved. Summary of the invention
[0005] In view of the above, it is necessary to provide a handwriting comparison method, device, equipment and medium based on multimodal features, aiming to solve the problems of low handwriting comparison accuracy and lack of interpretability.
[0006] A handwriting comparison method based on multimodal features, the handwriting comparison method based on multimodal features comprising:
[0007] In response to a handwriting comparison instruction based on the handwriting of a target user, collecting a historical handwriting image of the target user and a copy image of the historical handwriting image;
[0008] Constructing a plurality of triplets according to the historical handwriting image and the copied image;
[0009] For each triplet, extracting the stroke texture features, curve features and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet;
[0010] Constructing an initial model according to the encoder, and constructing a classification loss function and a triplet metric loss function of the initial model;
[0011] Taking each multimodal feature as a training sample, and training the initial model according to the classification loss function and the triplet metric loss function, to obtain a handwriting comparison model;
[0012] Determine a segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model, and determine the benchmark features of the historical handwriting image;
[0013] Acquire a handwriting image to be compared, and input the handwriting image to be compared into the handwriting comparison model to obtain the features to be compared of the handwriting image to be compared;
[0014] Calculating the distance between the feature to be compared and the reference feature as the target distance;
[0015] It is determined whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold.
[0016] A handwriting comparison device based on multimodal features, the handwriting comparison device based on multimodal features comprising:
[0017] A collection unit, configured to collect a historical handwriting image of the target user and a copy image of the historical handwriting image in response to a handwriting comparison instruction based on the handwriting of the target user;
[0018] A construction unit, used for constructing a plurality of triplets according to the historical handwriting image and the copied image;
[0019] An extraction unit, for extracting, for each triplet, stroke texture features, curve features, and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet;
[0020] The construction unit is further used to construct an initial model according to the encoder, and to construct a classification loss function and a triplet metric loss function of the initial model;
[0021] A training unit, used for taking each multimodal feature as a training sample and training the initial model according to the classification loss function and the triplet metric loss function to obtain a handwriting comparison model;
[0022] A determination unit, used to determine a segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model, and to determine the reference feature of the historical handwriting image;
[0023] An input unit, used for acquiring a handwriting image to be compared, and inputting the handwriting image to be compared into the handwriting comparison model to obtain a feature to be compared of the handwriting image to be compared;
[0024] A calculation unit, used for calculating the distance between the feature to be compared and the reference feature as a target distance;
[0025] The determination unit is further used to determine whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold.
[0026] A computer device, comprising:
[0027] a memory storing at least one instruction; and
[0028] A processor executes the instructions stored in the memory to implement the handwriting comparison method based on multimodal features.
[0029] A computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the handwriting comparison method based on multimodal features.
[0030] It can be seen from the above technical scheme that the present invention can construct multiple triplets based on historical handwriting images and copied images, and extract the stroke texture features, curve features and neural network features of each image in each triplet to form the multimodal features corresponding to each triplet. Due to the fusion of stroke texture features, curve features and neural network features, the problem of insufficient feature discrimination is effectively solved, and better results can be obtained when the amount of data is insufficient. In addition, the use of stroke texture features and curve features also solves the problem of lack of interpretability of using a single neural network feature; constructing a classification loss function and a triplet metric loss function can effectively shorten the distance between the same samples and increase the distance between different samples during the model training process, so as to better distinguish between positive and negative samples and improve the accuracy of the model; according to the image feature expression output by the last fully connected layer of the handwriting comparison model, the division threshold is reasonably determined, so that it can more accurately detect whether the handwriting image to be compared has the risk of copying the handwriting of the target user. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 It is a flow chart of a preferred embodiment of the handwriting comparison method based on multimodal features of the present invention.
[0032] Figure 2 It is a functional module diagram of a preferred embodiment of the handwriting comparison device based on multimodal features of the present invention.
[0033] Figure 3 It is a structural schematic diagram of a computer device of a preferred embodiment of the present invention for implementing a handwriting comparison method based on multimodal features. DETAILED DESCRIPTION
[0034] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0035] like Figure 1 FIG. 1 is a flowchart of a preferred embodiment of the handwriting comparison method based on multimodal features of the present invention. According to different requirements, the order of the steps in the flowchart can be changed, and some steps can be omitted.
[0036] The handwriting comparison method based on multimodal features is applied to one or more computer devices, wherein the computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASIC), programmable gate arrays (FPGA), digital signal processors (DSP), embedded devices, etc.
[0037] The computer device may be any electronic product that can perform human-computer interaction with a user, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive network television (IPTV), a smart wearable device, etc.
[0038] The computer device may also include a network device and / or a user device, wherein the network device includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud consisting of a large number of hosts or network servers based on cloud computing.
[0039] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms.
[0040] Among them, Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0041] AI basic technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include computer vision technology, robotics technology, biometrics technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0042] The network where the computer device is located includes but is not limited to the Internet, a wide area network, a metropolitan area network, a local area network, a virtual private network (VPN), etc.
[0043] S10, in response to a handwriting comparison instruction based on the handwriting of a target user, collecting a historical handwriting image of the target user and a copy image of the historical handwriting image.
[0044] In this embodiment, the handwriting comparison instruction can be triggered by relevant personnel according to actual needs. For example: in the financial field, when it is necessary to authenticate the signature of a contract, the handwriting comparison instruction can be triggered to verify whether the signature on the contract is signed by the contracting party himself, so as to avoid economic losses caused by malicious forgery of signatures. In the medical and health field, when it is necessary to authenticate the signature of a family member, the handwriting comparison instruction can also be triggered to verify whether the surgical consent form is signed by the family member himself, so as to avoid safety risks to the patient.
[0045] In this embodiment, the historical handwriting image may be a real handwriting image of the target user.
[0046] In order to expand the data, existing samples can be rotated, scaled, and noise added to increase sample diversity.
[0047] At the same time, the historical handwriting image may also include handwriting samples under different conditions such as different writing speeds, strengths, writing tools, paper types, etc., so that subsequent models can learn more generalized features and improve the ability to recognize handwriting in various actual scenarios.
[0048] In this embodiment, the copied image and the historical handwriting image have the same handwriting content and are used as negative samples.
[0049] For example, when the content of the handwriting in the historical handwriting image is "peace", the copying image is "peace" written by imitating the handwriting characteristics, handwriting form and other writing features of the historical handwriting image.
[0050] S11, constructing a plurality of triplets according to the historical handwriting image and the copied image.
[0051] In this embodiment, constructing a plurality of triplets according to the historical handwriting image and the copied image includes:
[0052] For each triple in the triples, obtaining a first image from the historical handwriting image as an anchor point sample;
[0053] Acquire a second image having the same handwriting content as the first image from the historical handwriting image as a positive sample similar to the anchor point sample;
[0054] Acquire a copy image of the first image from the copy image as a negative sample corresponding to the anchor point sample;
[0055] The anchor sample, the positive sample and the negative sample are combined to obtain the triplet.
[0056] Through the above embodiments, each triple feature can be established to correspond to the subsequent triple metric loss function.
[0057] S12: For each triplet, extract the stroke texture features, curve features, and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet.
[0058] In this embodiment, extracting the stroke texture features, curve features and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet includes:
[0059] For each triplet, extracting the gradient and the rotation of each image in the triplet to obtain the stroke texture feature of each image in the triplet;
[0060] Extracting the curvature of each image in the triplet to obtain a curve feature of each image in the triplet;
[0061] Extracting neural network features of each image in the triplet using a residual network;
[0062] The stroke texture features, curve features and neural network features of each image in the triplet are combined to obtain the multimodal features corresponding to the triplet.
[0063] Through the above embodiments, the stroke texture features, curve features and neural network features of each image are integrated to form a multimodal feature corresponding to each triplet, which can effectively solve the problem of insufficient feature discrimination.
[0064] In addition, the integration of stroke texture features and curve features can effectively solve the problem of lack of interpretability when using a single neural network feature.
[0065] In addition, the integration of stroke texture features and curve features that are consistent with manual authentication can also effectively improve the model's performance.
[0066] S13, constructing an initial model according to the encoder, and constructing a classification loss function and a triplet metric loss function of the initial model.
[0067] In this embodiment, the initial model includes two encoders connected in sequence.
[0068] S14, taking each multimodal feature as a training sample, and training the initial model according to the classification loss function and the triplet metric loss function to obtain a handwriting comparison model.
[0069] In this embodiment, the method of using each multimodal feature as a training sample and training the initial model according to the classification loss function and the triplet metric loss function to obtain the handwriting comparison model includes:
[0070] Inputting each multimodal feature into the initial model in sequence for training; wherein the initial model includes a first encoder and a second encoder, each multimodal feature is first input into the first encoder to obtain a first output feature of the first encoder, and then each first output feature is input into the second encoder;
[0071] During the training process, the network parameters of the initial model are optimized using the classification loss function and the triplet metric loss function; wherein the classification loss function is used to distinguish samples of different categories; the triplet metric loss function is used to shorten the distance between samples of the same category and shorten the distance between samples of different categories;
[0072] When the loss value of the classification loss function and the loss value of the triplet metric loss function reach a balance, the training is stopped to obtain the handwriting comparison model.
[0073] In the above embodiment, by continuously optimizing the network parameters of the model through training, a model with smaller values of the classification loss function and the triplet metric loss function can be obtained, thereby effectively distinguishing positive and negative samples and improving the accuracy of sample output.
[0074] S15, determining a segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model, and determining the benchmark features of the historical handwriting image.
[0075] In this embodiment, determining the segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model includes:
[0076] For each triplet, the feature expression corresponding to the positive sample and the feature expression corresponding to the negative sample are obtained from the image feature expression output by the last fully connected layer of the handwriting comparison model, and the distance between the feature expression corresponding to the positive sample and the feature expression corresponding to the negative sample is calculated as the distance corresponding to each triplet;
[0077] Select the minimum distance and the maximum distance from the distances corresponding to each triple to construct a distance interval, and traverse all possible distance values in the distance interval in sequence according to a preset step size as each candidate threshold;
[0078] For each candidate threshold, a verification set is obtained, and samples in the verification set that are greater than or equal to the candidate threshold are determined as candidate positive samples, and samples in the verification set that are less than the candidate threshold are determined as candidate negative samples;
[0079] Calculate the sample segmentation accuracy corresponding to each candidate threshold value based on the true positive samples, true negative samples marked in the verification set, and the candidate positive samples and candidate negative samples corresponding to each candidate threshold value;
[0080] The candidate threshold corresponding to the sample segmentation accuracy with the highest value is obtained as the segmentation threshold.
[0081] For example, when the minimum distance is 0 and the maximum distance is 10, the step size is 0.1, starting from 0 and increasing by 0.1 each time, to obtain each candidate threshold of 0.1, 0.2, 0.3...10. Each candidate threshold is used to divide the candidate positive sample and the candidate negative sample, and the comparison and calculation are performed with the real positive sample and the real negative sample to obtain the sample division accuracy corresponding to each candidate threshold, and the candidate threshold corresponding to the highest sample division accuracy is selected as the division threshold, so that the final division threshold is more accurate and reasonable, thereby improving the accuracy of subsequent handwriting comparison.
[0082] In this embodiment, determining the reference features of the historical handwriting image includes:
[0083] The feature expression corresponding to each anchor point sample is obtained from the image feature expression output by the last fully connected layer of the handwriting comparison model as the reference feature.
[0084] Through the above embodiments, accurate handwriting features can be obtained to serve as a reference for subsequent comparison.
[0085] S16, obtaining a handwriting image to be compared, and inputting the handwriting image to be compared into the handwriting comparison model to obtain features to be compared of the handwriting image to be compared.
[0086] In this embodiment, after the handwriting image to be compared is input into the handwriting comparison model, the image feature expression output by the last fully connected layer of the handwriting comparison model is obtained as the feature to be compared of the handwriting image to be compared.
[0087] S17, calculating the distance between the feature to be compared and the reference feature as a target distance.
[0088] In this embodiment, the cosine distance or the Euclidean distance between the feature to be compared and the reference feature may be calculated and used as the target distance.
[0089] S18, determining whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold.
[0090] In this embodiment, determining whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold comprises:
[0091] comparing the target distance with the division threshold;
[0092] When the target distance is greater than or equal to the division threshold, determining that the handwriting image to be compared has a risk of copying the handwriting of the target user; or
[0093] When the target distance is smaller than the division threshold, it is determined that the handwriting image to be compared does not have the risk of copying the handwriting of the target user.
[0094] Specifically, if the distance between the new handwriting (i.e., the handwriting image to be compared) and the positive sample (i.e., the benchmark feature) is less than the division threshold, it means that the feature of the new handwriting is similar to the feature of the positive sample of the signature of the person himself, then the new handwriting can be judged to be authentic, that is, it is considered to be written by the person himself.
[0095] If the distance between the new handwriting and the positive sample is greater than or equal to the division threshold, it means that the characteristics of the new handwriting are greatly different from the characteristics of the positive sample of the signature of the author, and the new handwriting can be judged as fake, that is, it is considered that it is not written by the author, but copied or forged by others.
[0096] Through the above embodiments, it is possible to accurately determine whether the handwriting has the risk of copying.
[0097] By conducting experiments on 6,000 pairs of test samples, we can see that even when the amount of data is insufficient, good results can be achieved, and the accuracy can even reach more than 95%.
[0098] Of course, by improving the training samples, even if the handwriting content is inconsistent, it is possible to identify whether the handwriting is from the same person by learning the writing style, stroke details, pause habits, etc., without limiting the handwriting to be identified to be exactly the same as the training sample.
[0099] It can be seen from the above technical scheme that the present invention can construct multiple triplets based on historical handwriting images and copied images, and extract the stroke texture features, curve features and neural network features of each image in each triplet to form the multimodal features corresponding to each triplet. Due to the fusion of stroke texture features, curve features and neural network features, the problem of insufficient feature discrimination is effectively solved, and better results can be obtained when the amount of data is insufficient. In addition, the use of stroke texture features and curve features also solves the problem of lack of interpretability of using a single neural network feature; constructing a classification loss function and a triplet metric loss function can effectively shorten the distance between the same samples and increase the distance between different samples during the model training process, so as to better distinguish between positive and negative samples and improve the accuracy of the model; according to the image feature expression output by the last fully connected layer of the handwriting comparison model, the division threshold is reasonably determined, so that it can more accurately detect whether the handwriting image to be compared has the risk of copying the handwriting of the target user.
[0100] like Figure 2 As shown, it is a functional module diagram of a preferred embodiment of the handwriting comparison device based on multimodal features of the present invention. The handwriting comparison device 11 based on multimodal features includes a collection unit 110, a construction unit 111, an extraction unit 112, a training unit 113, a determination unit 114, an input unit 115, and a calculation unit 116. The module / unit referred to in the present invention refers to a series of computer program segments that can be executed by a processor and can perform fixed functions, which are stored in a memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0101] The acquisition unit 110 is used to acquire a historical handwriting image of the target user and a copy image of the historical handwriting image in response to a handwriting comparison instruction based on the handwriting of the target user.
[0102] In this embodiment, the handwriting comparison instruction can be triggered by relevant personnel according to actual needs. For example: in the financial field, when it is necessary to authenticate the signature of a contract, the handwriting comparison instruction can be triggered to verify whether the signature on the contract is signed by the contracting party himself, so as to avoid economic losses caused by malicious forgery of signatures. In the medical and health field, when it is necessary to authenticate the signature of a family member, the handwriting comparison instruction can also be triggered to verify whether the surgical consent form is signed by the family member himself, so as to avoid safety risks to the patient.
[0103] In this embodiment, the historical handwriting image may be a real handwriting image of the target user.
[0104] In order to expand the data, existing samples can be rotated, scaled, and noise added to increase sample diversity.
[0105] At the same time, the historical handwriting image may also include handwriting samples under different conditions such as different writing speeds, strengths, writing tools, paper types, etc., so that subsequent models can learn more generalized features and improve the ability to recognize handwriting in various actual scenarios.
[0106] In this embodiment, the copied image and the historical handwriting image have the same handwriting content and are used as negative samples.
[0107] For example, when the content of the handwriting in the historical handwriting image is "peace", the copying image is "peace" written by imitating the handwriting characteristics, handwriting form and other writing features of the historical handwriting image.
[0108] The construction unit 111 is used to construct a plurality of triplets according to the historical handwriting image and the copied image.
[0109] In this embodiment, the construction unit 111 constructs a plurality of triplets according to the historical handwriting image and the copied image, including:
[0110] For each triple in the triples, obtaining a first image from the historical handwriting image as an anchor point sample;
[0111] Acquire a second image having the same handwriting content as the first image from the historical handwriting image as a positive sample similar to the anchor point sample;
[0112] Acquire a copy image of the first image from the copy image as a negative sample corresponding to the anchor point sample;
[0113] The anchor sample, the positive sample and the negative sample are combined to obtain the triplet.
[0114] Through the above embodiments, each triple feature can be established to correspond to the subsequent triple metric loss function.
[0115] The extraction unit 112 is used to extract, for each triplet, the stroke texture features, curve features and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet.
[0116] In this embodiment, the extraction unit 112 extracts the stroke texture features, curve features and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet, including:
[0117] For each triplet, extracting the gradient and the rotation of each image in the triplet to obtain the stroke texture feature of each image in the triplet;
[0118] Extracting the curvature of each image in the triplet to obtain a curve feature of each image in the triplet;
[0119] Extracting neural network features of each image in the triplet using a residual network;
[0120] The stroke texture features, curve features and neural network features of each image in the triplet are combined to obtain the multimodal features corresponding to the triplet.
[0121] Through the above embodiments, the stroke texture features, curve features and neural network features of each image are integrated to form a multimodal feature corresponding to each triplet, which can effectively solve the problem of insufficient feature discrimination.
[0122] In addition, the integration of stroke texture features and curve features can effectively solve the problem of lack of interpretability when using a single neural network feature.
[0123] In addition, the integration of stroke texture features and curve features that are consistent with manual authentication can also effectively improve the model's performance.
[0124] The construction unit 111 is further used to construct an initial model according to the encoder, and to construct a classification loss function and a triplet metric loss function of the initial model.
[0125] In this embodiment, the initial model includes two encoders connected in sequence.
[0126] The training unit 113 is used to use each multimodal feature as a training sample, and train the initial model according to the classification loss function and the triplet metric loss function to obtain a handwriting comparison model.
[0127] In this embodiment, the training unit 113 uses each multimodal feature as a training sample, and trains the initial model according to the classification loss function and the triplet metric loss function to obtain the handwriting comparison model including:
[0128] Inputting each multimodal feature into the initial model in sequence for training; wherein the initial model includes a first encoder and a second encoder, each multimodal feature is first input into the first encoder to obtain a first output feature of the first encoder, and then each first output feature is input into the second encoder;
[0129] During the training process, the network parameters of the initial model are optimized using the classification loss function and the triplet metric loss function; wherein the classification loss function is used to distinguish samples of different categories; the triplet metric loss function is used to shorten the distance between samples of the same category and shorten the distance between samples of different categories;
[0130] When the loss value of the classification loss function and the loss value of the triplet metric loss function reach a balance, the training is stopped to obtain the handwriting comparison model.
[0131] In the above embodiment, by continuously optimizing the network parameters of the model through training, a model with smaller values of the classification loss function and the triplet metric loss function can be obtained, thereby effectively distinguishing positive and negative samples and improving the accuracy of sample output.
[0132] The determining unit 114 is used to determine the segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model, and to determine the reference feature of the historical handwriting image.
[0133] In this embodiment, the determining unit 114 determines the segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model, including:
[0134] For each triplet, the feature expression corresponding to the positive sample and the feature expression corresponding to the negative sample are obtained from the image feature expression output by the last fully connected layer of the handwriting comparison model, and the distance between the feature expression corresponding to the positive sample and the feature expression corresponding to the negative sample is calculated as the distance corresponding to each triplet;
[0135] Select the minimum distance and the maximum distance from the distances corresponding to each triple to construct a distance interval, and traverse all possible distance values in the distance interval in sequence according to a preset step size as each candidate threshold;
[0136] For each candidate threshold, a verification set is obtained, and samples in the verification set that are greater than or equal to the candidate threshold are determined as candidate positive samples, and samples in the verification set that are less than the candidate threshold are determined as candidate negative samples;
[0137] Calculate the sample segmentation accuracy corresponding to each candidate threshold value based on the true positive samples, true negative samples marked in the verification set, and the candidate positive samples and candidate negative samples corresponding to each candidate threshold value;
[0138] The candidate threshold corresponding to the sample segmentation accuracy with the highest value is obtained as the segmentation threshold.
[0139] For example, when the minimum distance is 0 and the maximum distance is 10, the step size is 0.1, starting from 0 and increasing by 0.1 each time, to obtain each candidate threshold of 0.1, 0.2, 0.3...10. Each candidate threshold is used to divide the candidate positive sample and the candidate negative sample, and the comparison and calculation are performed with the real positive sample and the real negative sample to obtain the sample division accuracy corresponding to each candidate threshold, and the candidate threshold corresponding to the highest sample division accuracy is selected as the division threshold, so that the final division threshold is more accurate and reasonable, thereby improving the accuracy of subsequent handwriting comparison.
[0140] In this embodiment, the determination unit 114 determines the reference features of the historical handwriting image including:
[0141] The feature expression corresponding to each anchor point sample is obtained from the image feature expression output by the last fully connected layer of the handwriting comparison model as the reference feature.
[0142] Through the above embodiments, accurate handwriting features can be obtained to serve as a reference for subsequent comparison.
[0143] The input unit 115 is used to obtain a handwriting image to be compared, and input the handwriting image to be compared into the handwriting comparison model to obtain the features to be compared of the handwriting image to be compared.
[0144] In this embodiment, after the handwriting image to be compared is input into the handwriting comparison model, the image feature expression output by the last fully connected layer of the handwriting comparison model is obtained as the feature to be compared of the handwriting image to be compared.
[0145] The calculation unit 116 is used to calculate the distance between the feature to be compared and the reference feature as a target distance.
[0146] In this embodiment, the cosine distance or the Euclidean distance between the feature to be compared and the reference feature may be calculated and used as the target distance.
[0147] The determination unit 114 is further configured to determine whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold.
[0148] In this embodiment, the determining unit 114 determines whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold, including:
[0149] comparing the target distance with the division threshold;
[0150] When the target distance is greater than or equal to the division threshold, determining that the handwriting image to be compared has a risk of copying the handwriting of the target user; or
[0151] When the target distance is smaller than the division threshold, it is determined that the handwriting image to be compared does not have the risk of copying the handwriting of the target user.
[0152] Specifically, if the distance between the new handwriting (i.e., the handwriting image to be compared) and the positive sample (i.e., the benchmark feature) is less than the division threshold, it means that the feature of the new handwriting is similar to the feature of the positive sample of the signature of the person himself, then the new handwriting can be judged to be authentic, that is, it is considered to be written by the person himself.
[0153] If the distance between the new handwriting and the positive sample is greater than or equal to the division threshold, it means that the characteristics of the new handwriting are greatly different from the characteristics of the positive sample of the signature of the author, and the new handwriting can be judged as fake, that is, it is considered that it is not written by the author, but copied or forged by others.
[0154] Through the above embodiments, it is possible to accurately determine whether the handwriting has the risk of copying.
[0155] By conducting experiments on 6,000 pairs of test samples, we can see that even when the amount of data is insufficient, good results can be achieved, and the accuracy can even reach more than 95%.
[0156] Of course, by improving the training samples, even if the handwriting content is inconsistent, it is possible to identify whether the handwriting is from the same person by learning the writing style, stroke details, pause habits, etc., without limiting the handwriting to be identified to be exactly the same as the training sample.
[0157] It can be seen from the above technical scheme that the present invention can construct multiple triplets based on historical handwriting images and copied images, and extract the stroke texture features, curve features and neural network features of each image in each triplet to form the multimodal features corresponding to each triplet. Due to the fusion of stroke texture features, curve features and neural network features, the problem of insufficient feature discrimination is effectively solved, and better results can be obtained when the amount of data is insufficient. In addition, the use of stroke texture features and curve features also solves the problem of lack of interpretability of using a single neural network feature; constructing a classification loss function and a triplet metric loss function can effectively shorten the distance between the same samples and increase the distance between different samples during the model training process, so as to better distinguish between positive and negative samples and improve the accuracy of the model; according to the image feature expression output by the last fully connected layer of the handwriting comparison model, the division threshold is reasonably determined, so that it can more accurately detect whether the handwriting image to be compared has the risk of copying the handwriting of the target user.
[0158] like Figure 3The figure is a schematic diagram of the structure of a computer device of a preferred embodiment of the handwriting comparison method based on multimodal features of the present invention.
[0159] The computer device 1 may include a memory 12, a processor 13 and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a handwriting comparison program based on multimodal features.
[0160] Those skilled in the art will appreciate that the schematic diagram is merely an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 may be a bus-type structure or a star-type structure. The computer device 1 may also include more or less other hardware or software than shown in the diagram, or a different arrangement of components. For example, the computer device 1 may also include input and output devices, network access devices, etc.
[0161] It should be noted that the computer device 1 is only an example, and other existing or future electronic products that are suitable for the present invention should also be included in the protection scope of the present invention and included here by reference.
[0162] Among them, the memory 12 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (for example: SD or DX memory, etc.), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a mobile hard disk of the computer device 1. In other embodiments, the memory 12 can also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the computer device 1. Further, the memory 12 can also include both an internal storage unit of the computer device 1 and an external storage device. The memory 12 can not only be used to store application software and various types of data installed in the computer device 1, such as the code of a handwriting comparison program based on multimodal features, but also can be used to temporarily store data that has been output or is to be output.
[0163] In some embodiments, the processor 13 may be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 13 is the control core (Control Unit) of the computer device 1, and uses various interfaces and lines to connect the various components of the entire computer device 1, and executes or executes programs or modules stored in the memory 12 (for example, executing a handwriting comparison program based on multimodal features, etc.), and calls the data stored in the memory 12 to execute various functions of the computer device 1 and process data.
[0164] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the above-mentioned various multimodal feature-based handwriting comparison method embodiments, for example Figure 1 Steps shown.
[0165] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into a collection unit 110, a construction unit 111, an extraction unit 112, a training unit 113, a determination unit 114, an input unit 115, and a calculation unit 116.
[0166] The above-mentioned integrated unit implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, and includes a number of instructions for enabling a computer device (which can be a personal computer, a computer device, or a network device, etc.) or a processor to execute the part of the handwriting comparison method based on multimodal features described in various embodiments of the present invention.
[0167] If the module / unit integrated in the computer device 1 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware devices through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of each of the above-mentioned method embodiments can be implemented.
[0168] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form, etc. The computer readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.
[0169] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0170] The blockchain referred to in this invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, encryption algorithm, etc. Blockchain is essentially a decentralized database, a string of data blocks generated by cryptographic methods. Each data block contains a batch of network transaction information, which is used to verify the validity of its information (anti-counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, platform product service layer, and application service layer.
[0171] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The bus is represented by only one straight line, but it does not mean that there is only one bus or one type of bus. The bus is configured to realize the connection and communication between the memory 12 and at least one processor 13, etc.
[0172] Although not shown, the computer device 1 may also include a power source (such as a battery) for supplying power to each component. Preferably, the power source may be logically connected to the at least one processor 13 through a power management device, so that the power management device can realize functions such as charging management, discharging management, and power consumption management. The power source may also include any components such as one or more DC or AC power sources, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, etc. The computer device 1 may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0173] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.
[0174] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), or a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touch device. The display may also be appropriately referred to as a display screen or a display unit, which is used to display information processed in the computer device 1 and to display a visual user interface.
[0175] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0176] It can be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0177] Combination Figure 1 , the memory 12 in the computer device 1 stores a plurality of instructions to implement a handwriting comparison method based on multimodal features, and the processor 13 can execute the plurality of instructions to implement:
[0178] In response to a handwriting comparison instruction based on the handwriting of a target user, collecting a historical handwriting image of the target user and a copy image of the historical handwriting image;
[0179] Constructing a plurality of triplets according to the historical handwriting image and the copied image;
[0180] For each triplet, extracting the stroke texture features, curve features and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet;
[0181] Constructing an initial model according to the encoder, and constructing a classification loss function and a triplet metric loss function of the initial model;
[0182] Taking each multimodal feature as a training sample, and training the initial model according to the classification loss function and the triplet metric loss function, to obtain a handwriting comparison model;
[0183] Determine a segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model, and determine the benchmark features of the historical handwriting image;
[0184] Acquire a handwriting image to be compared, and input the handwriting image to be compared into the handwriting comparison model to obtain the features to be compared of the handwriting image to be compared;
[0185] Calculating the distance between the feature to be compared and the reference feature as the target distance;
[0186] It is determined whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold.
[0187] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here.
[0188] It should be noted that the data involved in this case were all obtained legally.
[0189] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0190] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0191] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0192] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0193] It is obvious to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0194] Therefore, no matter from which point of view, the embodiments should be regarded as illustrative and non-restrictive, and the scope of the present invention is limited by the appended claims rather than the above description, so it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention. Any attached figure mark in the claims should not be regarded as limiting the claims involved.
[0195] In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the present invention can also be implemented by one unit or device through software or hardware. The words first, second, etc. are used to indicate names, and do not indicate any particular order.
[0196] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.
Claims
1. A handwriting comparison method based on multimodal features, characterized in that: The handwriting comparison method based on multimodal features includes: In response to a handwriting comparison instruction based on the handwriting of a target user, collecting a historical handwriting image of the target user and a copy image of the historical handwriting image; Constructing a plurality of triplets according to the historical handwriting image and the copied image; For each triplet, extracting the stroke texture features, curve features and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet; Constructing an initial model according to the encoder, and constructing a classification loss function and a triplet metric loss function of the initial model; Taking each multimodal feature as a training sample, and training the initial model according to the classification loss function and the triplet metric loss function, to obtain a handwriting comparison model; Determine a segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model, and determine the benchmark features of the historical handwriting image; Acquire a handwriting image to be compared, and input the handwriting image to be compared into the handwriting comparison model to obtain the features to be compared of the handwriting image to be compared; Calculating the distance between the feature to be compared and the reference feature as the target distance; It is determined whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold.
2. The handwriting comparison method based on multimodal features as claimed in claim 1, characterized in that: The constructing of a plurality of triplets according to the historical handwriting image and the copied image comprises: For each triple in the triples, obtaining a first image from the historical handwriting image as an anchor point sample; Acquire a second image having the same handwriting content as the first image from the historical handwriting image as a positive sample similar to the anchor point sample; Acquire a copy image of the first image from the copy image as a negative sample corresponding to the anchor point sample; The anchor sample, the positive sample and the negative sample are combined to obtain the triplet.
3. The handwriting comparison method based on multimodal features as claimed in claim 1, characterized in that: The step of extracting the stroke texture features, curve features and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet includes: For each triplet, extracting the gradient and the rotation of each image in the triplet to obtain the stroke texture feature of each image in the triplet; Extracting the curvature of each image in the triplet to obtain a curve feature of each image in the triplet; Extracting neural network features of each image in the triplet using a residual network; The stroke texture features, curve features and neural network features of each image in the triplet are combined to obtain the multimodal features corresponding to the triplet.
4. The handwriting comparison method based on multimodal features as claimed in claim 1, characterized in that: The method of taking each multimodal feature as a training sample and training the initial model according to the classification loss function and the triplet metric loss function to obtain the handwriting comparison model includes: Inputting each multimodal feature into the initial model in sequence for training; wherein the initial model includes a first encoder and a second encoder, each multimodal feature is first input into the first encoder to obtain a first output feature of the first encoder, and then each first output feature is input into the second encoder; During the training process, the network parameters of the initial model are optimized using the classification loss function and the triplet metric loss function; wherein the classification loss function is used to distinguish samples of different categories; the triplet metric loss function is used to shorten the distance between samples of the same category and shorten the distance between samples of different categories; When the loss value of the classification loss function and the loss value of the triplet metric loss function reach a balance, the training is stopped to obtain the handwriting comparison model.
5. The handwriting comparison method based on multimodal features as claimed in claim 1, characterized in that: The determining of the segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model comprises: For each triplet, the feature expression corresponding to the positive sample and the feature expression corresponding to the negative sample are obtained from the image feature expression output by the last fully connected layer of the handwriting comparison model, and the distance between the feature expression corresponding to the positive sample and the feature expression corresponding to the negative sample is calculated as the distance corresponding to each triplet; Select the minimum distance and the maximum distance from the distances corresponding to each triple to construct a distance interval, and traverse all possible distance values in the distance interval in sequence according to a preset step size as each candidate threshold; For each candidate threshold, a verification set is obtained, and samples in the verification set that are greater than or equal to the candidate threshold are determined as candidate positive samples, and samples in the verification set that are less than the candidate threshold are determined as candidate negative samples; Calculate the sample segmentation accuracy corresponding to each candidate threshold value based on the true positive samples, true negative samples marked in the verification set, and the candidate positive samples and candidate negative samples corresponding to each candidate threshold value; The candidate threshold corresponding to the sample segmentation accuracy with the highest value is obtained as the segmentation threshold.
6. The handwriting comparison method based on multimodal features as claimed in claim 1, characterized in that: The step of determining the reference features of the historical handwriting image comprises: The feature expression corresponding to each anchor point sample is obtained from the image feature expression output by the last fully connected layer of the handwriting comparison model as the reference feature.
7. The handwriting comparison method based on multimodal features as claimed in claim 1, characterized in that: The step of determining whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold comprises: comparing the target distance with the division threshold; When the target distance is greater than or equal to the division threshold, determining that the handwriting image to be compared has a risk of copying the handwriting of the target user; or When the target distance is smaller than the division threshold, it is determined that the handwriting image to be compared does not have the risk of copying the handwriting of the target user.
8. A handwriting comparison device based on multimodal features, characterized in that: The handwriting comparison device based on multimodal features comprises: A collection unit, configured to collect a historical handwriting image of the target user and a copy image of the historical handwriting image in response to a handwriting comparison instruction based on the handwriting of the target user; A construction unit, used for constructing a plurality of triplets according to the historical handwriting image and the copied image; An extraction unit, for extracting, for each triplet, stroke texture features, curve features, and neural network features of each image in the triplet to form a multimodal feature corresponding to each triplet; The construction unit is further used to construct an initial model according to the encoder, and to construct a classification loss function and a triplet metric loss function of the initial model; A training unit, used for taking each multimodal feature as a training sample and training the initial model according to the classification loss function and the triplet metric loss function to obtain a handwriting comparison model; A determination unit, used to determine a segmentation threshold according to the image feature expression output by the last fully connected layer of the handwriting comparison model, and to determine the reference feature of the historical handwriting image; An input unit, used for acquiring a handwriting image to be compared, and inputting the handwriting image to be compared into the handwriting comparison model to obtain a feature to be compared of the handwriting image to be compared; A calculation unit, used for calculating the distance between the feature to be compared and the reference feature as a target distance; The determination unit is further used to determine whether the handwriting image to be compared has a risk of copying the handwriting of the target user according to the target distance and the division threshold.
9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the handwriting comparison method based on multimodal features as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the handwriting comparison method based on multimodal features as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Cross-modal character handwriting verification method, system and equipment and storage medium
CN115620312A
Data vectorization offline signature authentication method and system
CN117877127A
Systems and methods for adaptive preprocessor selection for efficient multi-modal classification
US20240386897A1