Handwritten signature recognition method and device, equipment and storage medium

Through the YOLOv8 and PP-OCRv4 model combined with BERT's multi-network joint judgment method, the problem of high data set preparation requirements and low recognition accuracy in handwritten signature recognition is solved, and a more efficient recognition effect is achieved.

CN120236289APending Publication Date: 2025-07-01JIANGSU HONGXIN SYST INTEGRATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510177398.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art has problems in handwritten signature recognition with high data set preparation requirements and low recognition accuracy, especially the difficulty in dealing with misdetection caused by differences in structurally similar fonts and personal signature habits.

Method used

The YOLOv8 model is used for signature positioning, the PP-OCRv4 model is used for text recognition, and the relationship between surnames and names is learned through the BERT model, combined with the joint judgment of multiple networks, the recognition results of top5 are retained and processed, and the final results are output according to different weights.

Benefits of technology

It improves the accuracy of handwritten signature recognition, reduces the requirements for data sets, reduces labor costs, has high concurrency and better recognition results than multimodal large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236289A_ABST
    Figure CN120236289A_ABST
Patent Text Reader

Abstract

The invention provides a handwritten signature recognition method and device, equipment and a storage medium, and relates to the technical field of computers. The identification method comprises the following steps: collecting a handwritten name data set; respectively training a YOLOv8 model and a PP-OCRv4 model by utilizing the handwritten name data set; taking the trained YOLOv8 model as a signature position positioning network, and outputting a signature positioning result; taking the trained PP-OCRv4 model as an OCR (Optical Character Recognition) text detection network, and outputting a signature character recognition result; and performing joint judgment on the signature character recognition result and the signature positioning result according to different weights, and taking the result with the highest probability as a final handwritten signature recognition result and outputting the final handwritten signature recognition result. Compared with the existing multi-mode large model detection, the algorithm has an identification result which is not inferior to that of the large model, but is lower in deployment requirement and higher in concurrency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a handwritten signature recognition method, apparatus, device, and storage medium. Background Art

[0002] Due to the different signing habits of customers and the varying lengths of names, it is difficult to effectively digitize the recognition of handwritten signatures. Currently, the recognition of handwritten names in contracts still mainly relies on manual labor, which greatly hinders the digital development of contract management.

[0003] In recent years, with the continuous development of computer vision technology, using OCR technology for handwritten name recognition has become the current mainstream method. However, simply using OCR technology cannot well solve the above characteristics of handwritten signatures, resulting in low recognition accuracy and difficulty in meeting the requirements of the digital transformation of contract management.

[0004] Currently, solutions using AI technology to solve the above problems have emerged, which are roughly divided into the following two solutions:

[0005] Solution 1: As disclosed in the Chinese invention patent with the publication number CN109086653B, a handwritten model training method, a handwritten character recognition method, apparatus, device, and medium are disclosed. A handwritten character recognition model is trained in stages using Chinese character samples of standard characters, non-standard Chinese character samples, and Chinese character samples with recognition errors in the validation set to continuously optimize the weights and biases of the handwritten character model, thereby obtaining a handwritten character recognition model with relatively high accuracy.

[0006] Solution 2: As disclosed in the Chinese invention patent with the publication number CN 117423119A, a method for recognizing handwritten Chinese characters in a scene based on a transformer is disclosed. First, a control point image correction method is used to correct the handwritten image. Secondly, a text line detection and segmentation algorithm based on a text kernel is used for handwritten detection and segmentation to obtain a handwritten text line image to be recognized. Then, a handwritten character recognition model based on a Transformer is constructed and trained using the above text, thereby obtaining a handwritten character recognition model with high generalization and accuracy.

[0007] However, the above solutions have the following problems:

[0008] Question 1: In order to train the OCR text recognition model described in Solution 1 and Solution 2, a very demanding handwritten character dataset is required. For example: in the handwritten characters "博" and "搏", due to their similar structures and the same fonts, if the OCR model wants to recognize the characters accurately, it needs to prepare a large amount of such relevant data. However, there are many fonts with similar structures in Chinese, and it is difficult to enumerate all similar fonts in production. Therefore, the use of OCR model detection alone cannot meet production requirements. In order to meet production needs and reduce the requirements for data set preparation, the present invention proposes to retain the top 5 test results each time, that is, whether it is the handwritten character "博" or "搏", as long as the Chinese character is among the top 5 results with the highest probability of model recognition, the model requirements are met, the data set preparation requirements are reduced, and it is easier to obtain an OCR recognition model with better generalization.

[0009] Question 2: Both Solution 1 and Solution 2 use the one-hot recognition method, that is, the result with the highest probability of handwriting model recognition is taken as the final recognition result. However, due to the different handwriting habits and character connection, relying solely on the recognition model to identify a unique result is likely to cause false detection. Summary of the invention

[0010] Purpose of the invention: To propose a handwritten signature recognition method, device, equipment and storage medium, which no longer relies solely on the final recognition result of the OCR model, but retains the top 5 recognition results of probability detected by the model each time, and performs joint processing based on the output Chinese character information and recognition strategy to select the most accurate effect from the top 5 results, greatly improving the recognition accuracy of the model, thereby effectively solving the above-mentioned problems existing in the prior art.

[0011] A first aspect of the present invention provides a handwritten signature recognition method, comprising the following steps:

[0012] Collect a dataset of handwritten names;

[0013] Using the handwritten name dataset, a YOLOv8 model and a PP-OCRv4 model are trained respectively;

[0014] The trained YOLOv8 model is used as the signature location network to output the signature location result; the trained PP-OCRv4 model is used as the OCR text detection network to output the signature text recognition result;

[0015] The signature text recognition result and the signature positioning result are jointly judged according to different weights, and the result with the highest probability is taken as the final handwritten signature recognition result and output.

[0016] In a further embodiment of the first aspect, the handwritten name dataset is derived from contract pictures containing handwritten signatures and the text information in the contracts; the YOLOv8 model is trained using the contract pictures containing handwritten signatures; and the PP-OCRv4 model is trained using the text information in the contracts.

[0017] In a further embodiment of the first aspect, the training of the YOLOv8 model using the contract pictures containing handwritten signatures specifically includes:

[0018] If the input is a single contract picture, it is directly input into the YOLOv8 model for detection;

[0019] If the input is a PDF file, the PDF file is first converted into pictures and then fed into the YOLOv8 model for detection.

[0020] In a further embodiment of the first aspect, the trained YOLOv8 model outputs the original position information as a four-dimensional vector

[0021] Based on the original position information, the current page number information and the supplementary position information coordinates and are added to form a nine-dimensional vector:

[0022]

[0023] where: and are the position information coordinates of the upper left corner and the lower right corner of the i-th detection target on the j-th page respectively; indicates that the detection target exists on the j-th page; and are the two supplementary information coordinates of the i-th detection target on the j-th page;

[0024] where:

[0025]

[0026] where: is the picture width of the j-th page; min(*) and max(*) are the maximum value function and the minimum value function respectively.

[0027] In a further embodiment of the first aspect, before using the trained PP-OCRv4 model to perform text recognition on the valid information it further includes:

[0028] Using the BERT model to learn the relationship between existing surnames and names, specifically including:

[0029] Construct the name division sample data, which is sourced from dictionaries, ancient documents, existing personal names, solar terms, allusions, and festivals;

[0030] Feed the name division sample data into the BERT model for fine-tuning.

[0031] In a further embodiment of the first aspect, use the trained PP-OCRv4 model to perform text recognition on the valid information Specifically, it includes:

[0032] Based on the supplementary information coordinates and the hint information set S in the valid information p Perform valid information recognition and invalid information filtering to obtain the name recognition information set Among them, the hint information set S p = {"signature", "sign", "seal"};

[0033] Perform surname recognition on the name recognition information set to obtain the surname result t surn ;

[0034] According to the surname result t surn divide the final name set to be inspected t name ;

[0035] Send the surname result t surn and the final name set to be inspected t name into the trained Bert model to judge the relationship between the surname and the name.

[0036] In a further embodiment of the first aspect, the expression of the name recognition information set is as follows:

[0037]

[0038] In the formula, f(*) represents the text recognition process; g(*) v represents the process of obtaining the top five results of the model output; represents the probability of the v-th result text recognition; η name is the result retention threshold.

[0039] In a further embodiment of the first aspect, the step of sending the surname result t surn and the final name set to be inspected t name into the trained Bert model to judge the relationship between the surname and the name specifically includes:

[0040] For t name and t surn construct the input sequence format as follows:

[0041]

[0042] wherein represents the information set to be detected in the k-th row;

[0043] Input the input sequence into the fine-tuned Bert model for name connection probability calculation, and record the calculation result as

[0044]

[0045] t name A total of a rows of name sets in are aggregated to obtain the final probability set

[0046] In a further embodiment of the first aspect, the signature text recognition result and the signature positioning result are jointly judged according to different weights, and the result with the highest probability is used as the final handwritten signature recognition result and output. Specifically, it includes:

[0047] The name connection probability result and the text recognition probability are used for joint probability calculation:

[0048]

[0049] where α is the regulation factor of the joint probability; is the combined probability value of the current name combination;

[0050] Perform joint probability calculation on a rows of name sets and normalize them to obtain the relative joint probability values of the current surname and different name combinations:

[0051]

[0052] where exp(*) represents the natural exponential function;

[0053] The name combination corresponding to the maximum result of the relative joint probability values of the current surname and different name combinations is used as the final recognition result P of the handwritten name final .

[0054] In the second aspect of the present invention, a handwritten signature recognition device is proposed, and the handwritten signature recognition method disclosed in the first aspect can be automatically executed by using this handwritten signature recognition device.

[0055] The handwritten signature recognition device includes:

[0056] A data acquisition module, which is used to collect a handwritten name data set; use the handwritten name data set to train the YOLOv8 model and the PP-OCRv4 model respectively;

[0057] The first output module is used to drive the trained YOLOv8 model to output the signature location result;

[0058] The second output module is used to drive the trained PP-OCRv4 model to output the signature text recognition result;

[0059] The joint judgment module is used to jointly judge the signature text recognition result and the signature location result according to different weights, and take the result with the highest probability as the final handwritten signature recognition result and output it.

[0060] In the third aspect of the present invention, an electronic device is proposed. The electronic device includes a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the handwritten signature recognition method disclosed in the first aspect is implemented.

[0061] In the fourth aspect of the present invention, a computer-readable storage medium is proposed. At least one executable instruction is stored in the storage medium. When the executable instruction runs on an electronic device, the electronic device is caused to execute the handwritten signature recognition method disclosed in the first aspect.

[0062] Compared with the prior art, the present invention has the following beneficial effects:

[0063] (1) By jointly judging the handwritten name through multiple networks, compared with the existing method of solely relying on OCR for text recognition in the market, the present invention can increase semantic connections and the recognition result is more accurate.

[0064] (2) By combining multiple networks, the present invention can reduce the requirements for the accuracy and universality of the handwritten name dataset, greatly releasing labor costs.

[0065] (3) Compared with the existing multi-modal large model detection, the present invention has a recognition result no less than that of the large model, but has lower deployment requirements and higher concurrency. Description of the Drawings

[0066] Figure 1 It is the overall flowchart of the handwritten signature recognition method in the embodiment.

[0067] Figure 2 It is the detailed flowchart of the handwritten signature recognition in the embodiment.

[0068] Figure 3 It is the specific flowchart of the surname and name splitting in the embodiment.

[0069] Figure 4 It is the picture of the contract handwritten text sample to be detected in the embodiment. Detailed Embodiments

[0070] In the following description, numerous specific details are given to provide a more thorough understanding of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced without one or more of these details. In other instances, some well-known technical features are not described in order to avoid obscuring the present invention.

[0071] Before elaborating on the embodiments, the following terms that will appear hereinafter are explained first:

[0072] AI: Artificial Intelligence, artificial intelligence

[0073] OCR: Optical Character Recognition, optical character recognition

[0074] BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, pre-trained language representation model

[0075] To solve the problem that it is difficult to prepare the dataset required for training existing high-quality OCR recognition models and the recognition accuracy of handwritten characters with similar structures is not high, the present invention proposes a method for accurately recognizing handwritten names based on joint judgment of multiple networks. First, modify the output structure of the traditional text detection model and retain the top 5 output results of the model each time. Secondly, according to the number of detected Chinese characters, split the surname and given name. For the detected Chinese characters with a length of 4 characters or less, couple all the detection results of the first character (for a 3-character name) or the first two characters (for a 4-character name) with the 1000 surnames included in the collection, and send the subsequent given name to the BERT model for joint judgment. For ethnic minority names with a length greater than 4 characters, directly input them into the BERT model for joint judgment. Finally, based on multiple results judged by multiple networks, sort them according to probability and take the maximum probability as the final result for output.

[0076] The technical solution involved in the present invention is mainly divided into three steps. One is the preparation of the handwritten name dataset and the training of multiple models; the second is to use the trained models to process the handwritten signature step by step according to the surname and given name; the third is to jointly judge the processing results of multiple networks according to different weights and output the result with the highest probability.

[0077] Regarding the dataset preparation and multi-model training described in the first part, the specific steps for data preparation and model training of multiple models are as follows:

[0078] 1.1. Select the YOLOv8 model for training as the signature position positioning network, and the training method of this model is as follows:

[0079] (1.1.1) Use PyMuPDF to convert the PDF contract into images. Since the handwritten signature is usually on the first and last pages of the contract, only the first 2 pages and the last 2 pages of each contract are saved as the training images for this model.

[0080] (1.1.2) Annotate the images retained in step 1.1.1 and convert the annotation format into a txt file.

[0081] (1.1.3) Use the optical flow method and the multi-sample data augmentation method to augment the data in step 1.2.

[0082] (1.1.4) Feed the augmented data into the YOLOv8 model for training and obtain the final signature position localization network.

[0083] 1.2. Use the PP-OCRv4 model for training as the OCR text detection network. The training method of this model is as follows:

[0084] (1.2.1) Expand the PP-OCRv4 font library from the original 6622 Chinese characters to 11254 Chinese characters to meet the requirements for rare Chinese characters in ethnic minority names.

[0085] (1.2.2) Use the annotation tool to annotate the handwritten text in the file, export all the annotation structures into an overall txt format, and divide the training set and the validation set.

[0086] (1.2.3) Use the PP-OCR serve Chinese and English general recognition model as the pre-trained model (retaining the original general recognition ability), and perform transfer learning training on this model using the training set data in step 1.2.2 to obtain the model training weights.

[0087] (1.2.4) Modify the output structure of the PP-OCRv4 model to output the top 5 recognition results each time. Use the validation set to verify the results. If all the names included in the validation set are among the top 5 recognition results of the model, the model training task ends. Otherwise, retrain the misdetected data in the validation set using the weights in step 1.2.3 through transfer learning.

[0088] (1.2.5) Lightweight export the model trained in step 1.2.4 to obtain the final handwritten name text recognition model.

[0089] 1.3. This invention is applied to the field of handwritten signature recognition in contracts. Through data analysis, it is found that the signer information mainly exhibits two characteristics in this field:

[0090] (1) Due to the large population in China, it often happens that there are duplicate signed names. And because of the existence of clans, some people's names are very similar, such as "Zhenglong", "Zhengyong", "Shouchun", "Shoutian".

[0091] (2) Based on data analysis, it is found that most Chinese names come from allusions, solar terms, idioms, poems, etc. For example: "Bowén", "Xiangyue", "Lixia", "Guoqing", etc.

[0092] In order to utilize these characteristics, the present invention uses the BERT model to learn the relationship between existing surnames and given names. Through a crawler program, 520,000 idioms included in the online Xinhua Dictionary, 50,000 historical names recorded in ancient books, 200,000 existing real names, and 1,500 items such as solar terms, allusions, and major festivals are collected. After filtering out useless punctuation marks, a dataset is formed and fed into the Bert model for fine-tuning. Different from the scenario text generation task, this model only needs to effectively judge the input text as a valid name, which is essentially a binary classification problem. The specific training process is as follows:

[0093] (1.3.1) A total of 200,000 real names, 50,000 historical names recorded in ancient books, and 100,000 virtual names generated based on idioms and solar terms are jointly used to form positive sample data;

[0094] (1.3.2) The original 520,000 idiom data included in the online Xinhua Dictionary and 1,500 data of solar terms, allusions, and major festivals are used as negative sample data;

[0095] (1.3.3) Each group of names in the dataset is labeled. Positive samples are labeled as 1, negative samples are labeled as 0, and the dataset is divided according to the ratio of 8:2;

[0096] (1.3.4) A fully connected layer is added to the top layer of the BERT model to predict whether the input sequence is a Chinese name.

[0097] (1.3.5) Each group of names in the dataset is used as a single input sequence, processed using the BERT tokenizer, and fed into the model for fine-tuning.

[0098] (1.3.6) The weights of the fine-tuned model are retained and used as the final name validity verification model.

[0099] The specific operation steps for processing the handwritten signature step by step according to the surname and given name using the trained model described in the second part are as follows:

[0100] Step 1, use the trained YOLOv8 model to locate the position of the handwritten name:

[0101] Since in the actual handwritten signature scenario, the signature position is usually on the last page of the contract, and customers generally not only sign their names by hand, but also sign contact information and time information, which are all interference information for the scenario of the present invention. However, relying solely on the YOLOv8 model cannot well distinguish between interference handwritten information and handwritten signatures. In order to better distinguish interference information and highlight the importance of effective information in the follow-up, the present invention adds the current page number information d lt ,y lt ,x rd ,y rd ) on the basis of the original model output position information (x page and supplementary position information (x elt ,y elt ,x erd ,y erd ), so that each result detected by the model is a 9-dimensional vector (x lt ,y lt ,x rd ,y rd ,d page ,x elt ,y elt, x erd ,y erd ). At the same time, through actual business analysis, it is known that the position of the handwritten signature is mostly in the lower part of the last page of the contract. The present invention makes full use of this feature and arranges all the information recognized by the model in reverse order according to the page number and the position where the information is located. The specific operation steps are as follows:

[0102] (1) If the input is a single image, it is directly input into the model for detection. If the input is a PDF file, the PDF file needs to be converted into a picture first and then sent to the model for detection.

[0103] (2) Compare the recognition probabilities of all the results detected by the model with the set threshold η, and filter out all the results with probabilities less than η. Define the remaining effective information as where n is the total number of information retained in the current picture, represents the 9-dimensional vector of the i-th detected information in the j-th page, and the specific composition is as follows:

[0104]

[0105] Among them, is the position information of the upper left corner and the lower right corner of the i-th detected target in the j-th page, indicates that the detected target exists in the j-th page, is the supplementary information coordinate, and the relevant calculation method is:

[0106]

[0107] wherein is the width of the picture on the j-th page, and min(*), max(*) are the maximum and minimum value functions respectively.

[0108] (3) For the retained valid information, based on the page number and the where the information is located, perform a reverse sorting from back to front and from bottom to top to highlight the importance of the signature information at the end of the last page.

[0109] Step 2: Use the trained PP-OCRv4 model to perform text recognition on what is obtained in Step 1 :

[0110] Under normal circumstances, there are prompt messages such as "Signature", "Signature with Seal", and "Seal" at the handwritten signature position of the customer name. These prompt messages are important bases for judging which ones among them are the customer's signature. For the convenience of subsequent description, the set of prompt messages is defined as S p = {"Signature with Seal", "Signature", "Seal"}.

[0111] (1) Based on the supplementary information coordinates in and S p perform valid information recognition and invalid information filtering. The specific calculation process is as follows:

[0112]

[0113] In the formula, f(*) represents the text recognition process, and g(*) v represents the process of obtaining the top 5 results of the model output. represents the probability of the v-th result text recognition (representing the index of the top 5 results), and η name is the result retention threshold. The specific meaning of the formula is that in for each information coordinate if there exists a v and an s (representing a certain keyword in the set of prompt messages), such that s appears in then the information recognized by this coordinate is included in the final name recognition information set .

[0114] (2) Perform surname recognition on the name recognition information set . The specific detection steps are as follows:

[0115] (2.1) Construct a surname character library set where c t is a single surname, and n sur is the length of the character library set. A total of 1000 surnames are stored in the present invention, so 1000 is taken here.

[0116] (2.2) Based on the length of each detected name in to confirm the surname, where

[0117]

[0118] Among them, dimensional matrix, is the set of recognition information retained after step (1), p u represents the recognition probability corresponding to the set of recognition information, a is the total number of the top 5 categories retained at the i-th target position on the j-th page. Since the 5 recognized text results may not all meet the probability η set in step (1) name , so here a ∈ [1, 5]. s dt represents the specific text entity detected, n s is the total number of characters detected at the i-th target position on the j-th page. represents the probability value that the recognition information at the i-th target position on the j-th page is u d , d ∈ [1, a].

[0119] According to the construction method of Chinese names, the length of n s and the specific name content to obtain the surname information, the division logic is as follows:

[0120] (2.2.1) If n s ≤4, it is considered that the target information is a Han surname, and the surname is obtained as follows:

[0121] (2.2.1.1) Take the first set of Chinese characters of [s 11 ,…s d1 ,…,s a1 ij T as the set of surnames to be detected for effective surname indexing in C surn .

[0122]

[0123] In the formula, t surn1 represents the process of single-character detection to filter all surname results that meet the C 11 ,…s d1 ,…,s a1 T set and the recognition result with the largest OCR recognition probability. k is the index value of [1, a]. surn

[0124] (2.2.1.2) Take ​​The first two Chinese character sets [(s 11 , s 12 ), … (s d1 , s d2 ), …, (s a1 , s a2 )] ij T are used as the surname set to be detected for effective surname indexing in C surn .

[0125]

[0126] In the formula, t surn2 represents filtering all surname results that satisfy the C 11 , s 12 ), … (s d1 , s d2 ), …, (s a1 , s a2 )] ij T set and the recognition result with the highest OCR recognition probability. surn (2.2.1.3) Based on the recognition situations of t

[0127] (2.2.1.3) Based on the recognition situations of t surn1 , t surn2 , and the detection probability, the specific surname confirmation is as follows:

[0128] When both t surn1 and t surn2 exist, select and retain the one with the highest probability corresponding to t surn1 and t surn2 as the final surname result t surn .

[0129]

[0130] When only one of t surn1 and t surn2 exists, directly take the existing result as the final surname result.

[0131]

[0132] When neither t surn1 nor t surn2 exists, in order to make the most of the subsequent Bert model's surname and name connection logic and the principle of mostly single surnames, directly take the Chinese character with the highest probability in the to-be-detected character set [s 11 , … s d1 , …, s a1 T as the final surname. ​

[0133]

[0134] (2.2.2) If n s > 4, it is considered that the target information is a minority name, and the surname is obtained as follows:

[0135] (2.2.2.1) Based on the minority name interval symbol ".", filter the invalid recognition results. The filtering calculation is as follows: Perform the filtering calculation as follows:

[0136]

[0137] In the formula, C ethinc is the information set containing the interval symbol "." in

[0138] (2.2.2.2) If C ethinc is, it means that the surname writing is not standardized. To facilitate the subsequent inference of the relationship between the surname and the name by the Bert model, directly use the first Chinese character of the character set to be detected [s 11 ,…s d1 ,…,s a1 ij T The Chinese character with the highest probability in the set is tentatively used as the surname.

[0139]

[0140] (2.2.2.3) When C ethinc is not empty, locate the first occurrence position of the interval symbol ".", and record it as κ. Then, use the first κ - 1 Chinese character set of[(s ,…,s 11 ,…,s 1κ-1 ),…(s d1 ,…,s dκ-1 ),…,(s a1 ,…,s aκ-1 )] ij T as the set of surnames to be detected for effective surname indexing in C surn The calculation is as follows:

[0141]

[0142] In the formula, t bsurn represents the surname result obtained by filtering all surnames that satisfy the C surn set during the detection of the first κ - 1 Chinese characters, and it is the recognition result with the highest OCR recognition probability.

[0143] Then, based on whether t bsurn is empty, the final surname result t​surn Confirmation. The specific calculation process is as follows:

[0144]

[0145] In the formula, if t bsurn is not empty, the reserved result of t bsurn is used as the output name and surname result. If t bsurn is empty, then the result with the highest recognition probability before the separator "." is used as the final result.

[0146] (2.3) Based on the surname result t surn confirmed in (2.2), the final set of names to be inspected is partitioned. The specific partitioning process is as follows:

[0147]

[0148] In the formula, when the surname t surn is a single-character surname, means that all information after the first character in the detection information is used as the set of names to be inspected t name When the surname t surn is a compound surname, means that all information after the second character is used as the set of names to be inspected t name When the surname t surn is a surname of ethnic minorities, all information including the separator "." is used as the set of names to be inspected t name .

[0149] Step 3: Send t name and t surn into the trained Bert model to judge the relationship between the surname and the name.

[0150] (3.1) Construct the input format for BERT model classification.

[0151] In the Bert model, [CLS] indicates the start marker of the classification task, and [SEP] is the segmentation marker used to separate different tokens. When there is no new token connected after [SEP], it means the end of the sequence construction. Therefore, for the input t name and t surn , the constructed sequence format is where represents the k-th row of the set of information to be inspected.

[0152] (3.2) Input the input sequence constructed in step 3.1 into the fine-tuned Bert model to calculate the name connection probability, and record the calculated result as

[0153] (3.3) Perform t name For the name sets in row a of t, respectively perform (3.1) and (3.2) to obtain the final probability set

[0154] As described in the third part, the specific calculation process of jointly judging the processing results of multiple networks according to different weights and outputting the result with the highest probability is as follows:

[0155] 3.1 Calculate the joint probability of the relationship probability between the surname and the given name output by the Bert model and the character recognition probability of the OCR model for joint probability calculation:

[0156]

[0157] In the formula, α is the adjustment factor of the joint probability, which is dynamically adjusted based on the recognition scenario, and is the combined probability value of the current name combination.

[0158] 3.2 Calculate the joint probability for each name set in row a according to step 3.1 respectively, and perform normalization to obtain the relative joint probability values of the current surname and different given name combinations.

[0159]

[0160] 3.3 Sort the a names obtained in step 3.2, and use the name combination corresponding to the maximum probability as the final recognition result of the handwritten name.

[0161] The present invention can be used for the recognition of handwritten names in contracts. Only the recognition steps are described in detail below. For model training, please refer to the "Technical Solution of the Invention and Creation" section.

[0162] Take a sample picture such as Figure 4As shown in the figure, in the first step of the second part of the technical solution, two targets are generated during the YOLOv8 detection, namely the position box of the real name "Ma Bowen" and the handwritten "2024" information box of the interference information. The 9-dimensional vectors constructed by the two pieces of information in step 2.1 are [790, 1553, 997, 1940, 5, 997, 1553, 1204, 1940] and [685, 2123, 867, 2456, 5, 776, 2123, 867, 2456] respectively. Among them, (997, 1553, 1204, 1940) and (776, 2123, 867, 2456) are the pictures detected by YOLOv8. (790, 1553, 997, 1940) and (685, 2123, 867, 2456) are the coordinates of the supplementary information, and 5 indicates that this is the handwritten content on the 5th page of the contract.

[0163] In the second step of the second part of the technical solution, the trained PP-OCRv4 model is used to perform text recognition on the coordinate information obtained in the first step, and the detection results with a filtering threshold lower than 0.6 in the top 5 results are output:

[0164] (Fang / signature, 0.753)

[0165] {(Ma Bojiu, 0.729), (Ma Bowen, 0.687), (Ma Bowen, 0.602)}

[0166] (Date:, 0.932)

[0167] {(2024, 0.892), (20 Dou, 0.613)}

[0168] After the above results pass through the information filtering formula in step (1), since the date does not meet the range of S p = {"signature", "sign", "seal"}, the date information is filtered out, and {(Ma Bojiu, 0.729), (Ma Bowen, 0.687), (Ma Bowen, 0.602)} is used as the output result of the handwritten signature on this page.

[0169] Regarding the recognition and retention of the surname t in step (2.2), specifically, since the character length n surn ≤ 4, and [s s ,…s 11 ,…,s d1 ,…,s a1 T constitutes the set of surnames to be detected {Ma, Ma, Ma}, and since Ma is in the surname character library with 1000 entries, t surn1 is Ma. And [(s 11 , s 12 ),…(s d1 , s​d2 ),…,(s a1 ,s a2 )] ij T The set of surnames to be tested {(Ma Bo), (Ma Bo), (Ma Po)} is not included in the surname character library, so So the final surname t surn is Ma, and the set of names to be tested is {(Bo Jiu), (Bo Wen), (Po Wen)}.

[0170] In the third step of the second part of the technical solution, t name and t surn are sent into the trained Bert model to judge the relationship between the surname and the name.

[0171] The input format constructed in step (3.1) is: {([CLS] Ma [SEP] Bo Jiu [SEP]), ([CLS] Ma [SEP] Bo Wen [SEP]), ([CLS] Ma [SEP] Po Wen [SEP])};

[0172] The results recognized by the Bret model in steps (3.2) to (3.3) are {0.70, 0.89, 0.42};

[0173] In the third part of the technical solution, the processing results of multiple networks are jointly judged according to different weights. Specifically, and are jointly calculated according to the technical solution three. The specific calculation is not done here. Finally, {(Ma Bowen, 0.489), (Ma Bo Jiu, 0.395), (Ma Po Wen, 0.107)} is obtained. Therefore, according to the principle of the highest probability, Ma Bowen is selected as the final recognition result.

[0174] The technical idea of the above embodiment can be constructed into a handwritten signature recognition device. Using this handwritten signature recognition device, the publicly disclosed handwritten signature recognition method can be automatically executed. The handwritten signature recognition device includes a data acquisition module, a first output module, a second output module, and a joint judgment module. The data acquisition module is used to collect the handwritten name data set; the YOLOv8 model and the PP-OCRv4 model are respectively trained using the handwritten name data set. The first output module is used to drive the trained YOLOv8 model to output the signature localization result. The second output module is used to drive the trained PP-OCRv4 model to output the signature text recognition result. The joint judgment module is used to jointly judge the signature text recognition result and the signature localization result according to different weights, and output the result with the highest probability as the final handwritten signature recognition result.

[0175] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0176] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0177] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as a limitation of the present invention itself. Various changes can be made to it in form and detail without departing from the spirit and scope of the present invention defined by the appended claims.

Claims

1. A handwritten signature recognition method, characterized in that: The steps include: Collect a dataset of handwritten names; Using the handwritten name dataset, a YOLOv8 model and a PP-OCRv4 model are trained respectively; The trained YOLOv8 model is used as the signature location network to output the signature location result; the trained PP-OCRv4 model is used as the OCR text detection network to output the signature text recognition result; The signature text recognition result and the signature positioning result are jointly judged according to different weights, and the result with the highest probability is taken as the final handwritten signature recognition result and output.

2. The handwritten signature recognition method according to claim 1, characterized in that: The handwritten name dataset comes from contract images containing handwritten signatures and the text information in the contract images; the YOLOv8 model is trained using the contract images containing handwritten signatures; and the PP-OCRv4 model is trained using the text information in the contract images.

3. The handwritten signature recognition method according to claim 2, characterized in that: The YOLOv8 model is trained using a contract image containing a handwritten signature, specifically including: If the input is a single contract image, it is directly input into the YOLOv8 model for detection; If the input is a PDF file, the PDF file is first converted into an image and then sent to the YOLOv8 model for detection.

4. The handwritten signature recognition method according to claim 3, characterized in that: The trained YOLOv8 model outputs the original position information as a four-dimensional vector Based on the original location information, add the current page number information and supplementary location information coordinates and This forms a nine-dimensional vector: In the formula, and are the position information coordinates of the upper left corner and lower right corner of the i-th detection target in the j-th page; Indicates that the detection target exists in page j; and are the two supplementary information coordinates of the i-th detection target in the j-th page; in: In the formula, is the width of the image on page j; min(*) and max(*) are the maximum value function and the minimum value function respectively.

5. The handwritten signature recognition method according to claim 4, characterized in that: When using the trained PP-OCRv4 model to Before text recognition, it also includes: Use the BERT model to learn the relationship between existing surnames and given names, including: Constructing name classification sample data, wherein the name classification sample data is derived from dictionaries, ancient books, existing names, solar terms, allusions, and festivals; The name segmentation sample data is sent to the BERT model for fine-tuning.

6. The handwritten signature recognition method according to claim 5, characterized in that: Use the trained PP-OCRv4 model to Perform text recognition, including: Based on valid information The supplementary information coordinates and prompt information set S in p Perform valid information identification and invalid information filtering to obtain the name recognition information set The prompt information set S p ={"Signature","Signature","Seal"}; Name Identification Information Set Perform surname recognition and obtain surname result t surn ; Results by last name surn , divide the final name set to be checked into t name ; The last name result is t surn , Final name set to be checked name The data is sent to the trained Bert model to determine the relationship between the last name and the first name.

7. The handwritten signature recognition method according to claim 6, characterized in that: The last name result t surn , Final name set to be checked name The trained Bert model is sent to determine the relationship between the surname and the first name, including: For t name and t surn , construct the input sequence format as follows: in represents the k-th row of information set to be checked; The input sequence is input into the fine-tuned BERT model to calculate the name association probability, and the calculated result is recorded as t name There are a total of a rows of name sets in the final probability set.

8. The handwritten signature recognition method according to claim 7, characterized in that: The signature text recognition result and the signature positioning result are jointly judged according to different weights, and the result with the highest probability is taken as the final handwritten signature recognition result and output, specifically including: Associating the name with the probability result and text recognition probability Perform joint probability calculation: In the formula, α is the control factor of the joint probability; is the combined probability value of the current name combination; Calculate the joint probability of the name set in row a and normalize it to get the relative joint probability value of the current surname and different first names: The name combination corresponding to the maximum result of the relative joint probability value of the current surname and different first names is used as the final recognition result of the handwritten name.

9. A handwritten signature recognition device, characterized in that: include: Data acquisition module, used to collect handwritten name dataset; Using the handwritten name dataset, a YOLOv8 model and a PP-OCRv4 model are trained respectively; The first output module is used to drive the trained YOLOv8 model to output signature positioning results; The second output module is used to drive the trained PP-OCRv4 model to output signature text recognition results; A joint judgment module is used to jointly judge the signature text recognition result and the signature positioning result according to different weights, and output the result with the highest probability as the final handwritten signature recognition result; When the handwritten signature recognition device is operated, it can automatically execute the handwritten signature recognition method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the handwritten signature recognition method according to any one of claims 1 to 8 is implemented.

11. A computer-readable storage medium, characterized in that: The storage medium stores at least one executable instruction. When the executable instruction is executed on the electronic device, the electronic device executes the handwritten signature recognition method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Handwriting model training methods, handwritten character recognition methods, devices, equipment and media

    CN109086653B

  • Scene handwritten Chinese character recognition method based on Transform

    CN117423119A