Multi-stage data processing method, system and equipment for handwriting recognition and medium
Through multi-stage data processing methods, including data preprocessing, in-data processing and data postprocessing, the problem of inaccurate handwriting recognition caused by poor quality of electronic signature data is solved, and the robustness and accuracy of handwriting recognition are improved.
Patent Information
- Application Number
- CN202311626563.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, poor quality of electronic signature data leads to inaccurate handwriting recognition results, and there are problems of non-standardization, noise and signature style transformation, which affects the recognition accuracy.
Multi-stage data processing methods are adopted, including data pre-processing, in-data processing and data post-processing. Data preprocessing uses operation specification signature data such as scale normalization, direction normalization, breakpoint smoothing and noise augmentation; data processing uses dynamic time planning to align and fusion features to generate more diverse paired features; data postprocessing adaptively adjusts scores based on the quality of signature sample retention and dynamic corrections.
It effectively improves the robustness and accuracy of the handwriting recognition model, can better deal with the non-standardization, noise and style transformation of signature data, and improves the accuracy of electronic signature recognition.
Smart Images

Figure CN120299095A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer information processing, and in particular to a multi-stage data processing method for improving handwriting recognition effect. Background Art
[0002] As a kind of serial data, electronic signature contains the unique individual characteristics of the signatory, such as writing style, speed, pressure, etc. This information is crucial to ensure the authenticity and accuracy of the signature. In the task of handwriting recognition of electronic signatures, the quality of electronic signature data is crucial. However, due to the different devices for collecting and retaining signature samples, the changes in the writing habits of the signatory, and the influence of various real-life scenarios when signing, electronic signature data often has a series of quality problems, such as:
[0003] Non-standardization problem: electronic signature data may be collected in an irregular form due to the state of the signature, such as the position, tilt angle, size or rotation of the signature;
[0004] Noise problem: Electronic signature data may be interfered by the noise of the device itself, such as hardware failure, unstable input device or sensor, resulting in instability and inaccuracy of the acquired sequence data;
[0005] Style change problem: When different signers sign at different times and on different devices, their writing styles may change, including differences in speed and pressure, which may lead to large differences between different signature data of the same signer.
[0006] Including traditional machine learning methods based on feature engineering and data-based deep learning methods, there are problems with the collected electronic signature data, which makes it difficult to recognize electronic signature handwriting, resulting in inaccurate handwriting recognition results, and brings huge challenges to the accuracy of handwriting recognition. Improving the quality of electronic signature data and improving the accuracy of recognition results are technical problems that need to be solved urgently, which is crucial to improving the accuracy of handwriting recognition.
[0007] The Chinese invention patent application with publication number CN115862153A and name “Signature handwriting authentication method, storage medium, and electronic device based on comprehensive point features, local features, and global features” discloses performing time dynamic programming based on the point features of the sample signature and the signature to be tested to obtain the point feature similarity of the sample signature and the signature to be tested, calculating the local feature similarity of the sample signature and the signature to be tested based on the LNPS feature, calculating the global feature similarity of the sample signature and the signature to be tested based on the Euclidean distance, calculating the comprehensive similarity of the sample signature and the signature to be tested based on the point feature similarity, local feature similarity, and global feature similarity, and authenticating the signature based on the comprehensive similarity.
[0008] This method focuses on the design of features to improve handwriting recognition, but does not involve the consideration of signature data quality and the corresponding processing methods. For signature data of poor quality, there will be differences between the global similarity and local feature similarity calculated based on point features and distance. Summary of the invention
[0009] In view of the problems existing in the prior art, the present invention proposes a multi-stage processing method for handwriting sampling data, which can be effectively applied to machine learning related handwriting recognition methods and deep learning related handwriting recognition methods to solve the problem of inaccurate handwriting recognition results caused by poor signature quality in handwriting recognition methods.
[0010] Based on the first aspect of the present application, a multi-stage data processing method for handwriting recognition is proposed, including: data pre-processing: collecting electronic signature sample data, obtaining electronic signature verification data, and standardizing the electronic signature sample data and signature verification data; data processing: extracting the fusion features of the electronic signature sample data pair and merging them with the signature verification data respectively to obtain the fusion features of the sample-verification data pair; data post-processing: correcting and verifying the score of the electronic signature data pair output by the handwriting recognition model, and outputting the judgment result.
[0011] Further preferably, the data standardization includes: performing a series of data preprocessing to standardize the signature data for data non-standardization problems and noise problems; the data fusion includes: fusing data pair features based on dynamic time planning to generate more diverse paired features for the signature style transformation problem; the electronic signature data pair score correction includes: adaptively adjusting the correction value of the sample-verification signature score according to the signature sample quality during handwriting verification.
[0012] Further preferably, the series of data preprocessing includes: scale normalization, direction normalization, breakpoint smoothing, Gaussian noise data augmentation, mask data augmentation; fusing data pair features further includes: calculating signature writing features, DTW alignment, probabilistic fusion of paired features and labels, and inputting an electronic handwriting recognition model to obtain a verification signature score and a retained sample quality score; adaptively adjusting the retained sample-verification signature score correction includes: calculating the maximum, minimum and average values of the verification score, calculating the maximum, minimum and average values of the quality score, fitting the quality score to a T distribution to calculate the retained sample difference, and correcting the verification score.
[0013] Further preferably, the normalization includes: according to the minimum and maximum coordinates (x min ,y min )、(x max ,y max ), calling the formula: Get the coordinate position of the trajectory point after scale normalization According to the angle θ of the original signature data trace points and the set standard rotation angle θ std , according to the formula: R = θ std - θ to calculate the angle difference of the original signature data trace points, according to the formula:
[0014] Calculate the coordinate position of the trace points after direction normalization Translate the coordinate points (x i , y i , y i ) of the current electronic signature s to the standard direction position to achieve the overall rotation of the electronic signature.
[0015] Further preferably, the breakpoint smoothing process can be achieved by smoothing the high-frequency noise components (corresponding to the breakpoints in the electronic signature data) in the frequency-domain representation of the original electronic signature s i data. Specifically, call the formula:
[0016] Modify the frequency-domain representation (X(ω), Y(ω)) of the original signature data to the smoothed coordinate frequency domain where X(ω) and Y(ω) respectively represent the frequency-domain representations of the abscissa and ordinate, ω represents the frequency, ω c represents the cut-off frequency, which determines the cut-off position of the filter in the frequency domain, and n represents the order of the filter, which determines the steepness of the filter.
[0017] Further preferably, the Gaussian noise data augmentation and mask data augmentation include: according to the noise amplitude A, noise signal frequency f i , y i ) of the original electronic signature data trace points (x i , noise signal phase φ, call the formula:
[0018] G(i) = A·cos(2πf i + φ),
[0019] Calculate the Gaussian noise signal G(i); determine the coordinates after Gaussian noise data augmentation as:
[0020]
[0021] According to the mask state M(i), determine the coordinates after mask data augmentation as:
[0022]
[0023] Further preferably, the correction of adaptively adjusting the signature retention-verification score further includes: calculating the quality score of each signature retention and adaptively giving a larger weight to the high-quality retention according to the quality score, and dynamically correcting the most similar signature retention-verification signature pair and the signature retention-verification signature pair with the largest difference at the same time; the probability fusion of paired features and labels further includes: randomly sampling a random factor λ from the Beta distribution of the electronic signature stroke trajectory, and fusing the aligned electronic signature writing features and the corresponding labels (label, label') with the random factor λ to obtain the fused features and the mixed labels
[0024] Based on the second aspect of the present invention, a multi-stage data processing system for handwriting recognition is proposed, including: a data preprocessing module: used for collecting electronic signature retention data, obtaining electronic signature verification data, and standardizing the electronic signature retention data and signature verification data; a data middle processing module: used for extracting the fused features of the electronic signature retention data and fusing them with the signature verification data respectively to obtain the fused features of the signature retention-verification data pair; a data postprocessing module: used for correcting and verifying the scores of the electronic signature data pairs output by the handwriting recognition model, and outputting a judgment result.
[0025] Further preferably, the standardization of the data includes: performing one or several preprocessings on the non-standardized data and the noise in the data, including but not limited to: scale normalization, direction normalization, breakpoint smoothing, Gaussian noise data augmentation, mask data augmentation, etc., to standardize the signature data; the data fusion includes: fusing the paired features of the data based on dynamic time warping to complete the signature style transformation and generate more diverse paired features; the correction of the scores of the electronic signature data pairs includes: adaptively adjusting the correction value of the signature retention-verification score according to the signature retention quality during handwriting verification.
[0026] Further preferably, the correction value of adaptively adjusting the signature retention-verification score further includes: calculating the quality score of each retention and giving a larger weight to the high-quality retention than the low-quality retention according to the quality score, and dynamically correcting the signature retention-verification signature pair with the smallest difference and the signature retention-verification signature pair with the largest difference according to the weight; the fusion of the paired features of the data further includes: calculating the signature writing features, aligning the paired features of the data based on dynamic time warping, probabilistically fusing the paired features and labels, and inputting into the electronic handwriting recognition model to obtain the signature retention-verification score and the retention quality score
[0027] Further preferably, the correction value of the sample-verification signature score is adaptively adjusted, including: calculating the maximum, minimum and average values of the sample-verification signature score, calculating the maximum, minimum and average values of the quality score, fitting the quality score to the T distribution to calculate the sample difference and correct the sample-verification signature score; the probability fusion of paired features and labels further includes: randomly sampling from the Beta distribution of the electronic signature stroke trajectory to obtain a random factor λ, and aligning the electronic signature handwriting features after alignment And the corresponding label (label, label') is fused with a random factor λ to obtain the fusion feature With mixed tags
[0028] Based on the third aspect of the present application, an electronic device is proposed, comprising: a processor; and a memory for storing a program, wherein the program comprises instructions which, when executed by the processor, cause the processor to execute the multi-stage data processing method for handwriting recognition described above.
[0029] Based on the fourth aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is proposed, wherein the computer instructions are used to enable the computer to execute the multi-stage data processing method for handwriting recognition described above.
[0030] The present invention performs a series of data preprocessing in the process of data preprocessing to address data non-standardization problems and noise problems, so as to standardize signature data and improve the robustness of signature recognition models / methods to noise; in the process of data processing, in order to address the problem of signature style change, paired data features are fused to generate more diverse paired features, making it easier for signature recognition models / methods to focus on the differences between paired data (such as style change), which can effectively improve the discrimination ability and robustness of handwriting recognition models / methods; in the process of data post-processing, a correction method for adaptively adjusting the sample-verification signature score according to the signature sample quality during handwriting verification is used to improve the final accuracy of handwriting recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A flowchart of an electronic handwriting recognition method based on multi-stage data processing in an exemplary embodiment of the present application;
[0032] Figure 2 It is a multi-stage data processing for electronic handwriting recognition in this exemplary embodiment;
[0033] Figure 3 is a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present application. DETAILED DESCRIPTION
[0034] The present invention focuses on the problem of inaccurate handwriting recognition results caused by poor electronic signature quality, and proposes a multi-stage data processing method for improving handwriting recognition effect. It effectively deals with the non-standardization problem, noise problem and writing style transformation problem of signature data, and calculates more accurate recognition results according to the signature data quality in the handwriting verification stage, effectively improving the quality of electronic signature data and achieving more accurate electronic signature recognition.
[0035] Specifically, before processing the electronic signature data, the present invention performs three-stage processing on the collected electronic signature sequence data to obtain electronic sequence data suitable for the handwriting recognition task. That is, it performs pre-data processing on the original data sequence of the collected electronic signature, performs mid-data processing on the electronic signature features, and performs post-data processing on the electronic signature recognition process, which are hereinafter simply referred to as pre-data processing, mid-data processing and post-data processing.
[0036] Especially, in the pre-data processing process, a series of data preprocessings are carried out for the data non-standardization problem and noise problem to standardize the signature data and improve the robustness of the signature recognition model / method to noise.
[0037] In the mid-data processing process, for the signature style transformation problem, based on dynamic time warping (DTW), the paired data features are fused, which can generate more diverse paired features, making it easier for the signature recognition model / method to focus on the differences between paired data (such as style transformation), and can effectively improve the discrimination ability and robustness of the handwriting recognition model / method.
[0038] In the post-data processing process, when performing handwriting verification, the correction method for adaptively adjusting the sample-verification signature score is adjusted according to the signature sample quality to improve the final accuracy of handwriting recognition. Specifically, the quality score of each signature sample is calculated and a larger weight is adaptively given to the high-quality samples according to the quality score. At the same time, the most similar sample-verification signature pairs and the most different sample-verification signature pairs are dynamically corrected, so as to give a more reliable verification result.
[0039] In the handwriting recognition task, it is necessary to perform discrimination tasks through the signature data pairs of "sample signature-verification signature".
[0040] The present invention proposes a three-stage electronic sequence data processing method suitable for the handwriting recognition task,
[0041] performs pre-data processing on the original electronic signature data collected by the acquisition device, performs mid-data processing on the electronic signature features, and performs post-data processing on the electronic signature recognition process.
[0042] Such as Figure 1The following is a schematic flowchart of an electronic handwriting recognition method based on multi-stage data processing in an exemplary embodiment of the present application, including data preprocessing, collecting electronic signature sample data to obtain a series of signature sample data, and acquiring electronic signature verification data; standardizing a series of electronic signature sample data and signature verification data; data middle processing, extracting fused features of a series of electronic signature sample data pairs, respectively fusing them with signature verification data to obtain a series of sample-verification data pair fused features; the fused features pass through a handwriting recognition model to obtain scores of electronic signature data pairs, including a series of sample data pair quality scores and a series of sample-verification data pair verification scores; the data post-processing module verifies the scores of the electronic signature data pairs and outputs a judgment result.
[0043] In the data preprocessing module, a series of data preprocessing methods are adopted for the non-standardization problem and noise problem of the original electronic signature data to standardize the signature data and improve the robustness of the handwriting recognition model (method) to noise.
[0044] The data middle processing module proposes a paired data feature fusion method based on dynamic time warping (DTW) for the signature style transformation problem. First, standardize the manual features of the electronic signature data pairs, align and fuse them to generate more diverse paired fused features, making it easier for the model and method to focus on the differences between paired data (such as signature style transformation), effectively improving the discrimination ability and robustness of the handwriting recognition model and method. During the handwriting verification process, based on the existing handwriting recognition methods and models, the initial score of the verification signature is calculated from the paired fused features of the signature data pairs.
[0045] The data post-processing module adaptively adjusts the correction of the sample-verification signature score according to the signature sample quality to improve the final accuracy of handwriting recognition. This method calculates the quality score of each signature sample and adaptively gives a larger weight to high-quality samples, and at the same time dynamically corrects the most similar sample-verification signature pairs and the most different sample-verification signature pairs to obtain a more accurate handwriting verification result.
[0046] Figure 2The multi-stage data processing for electronic handwriting recognition in this exemplary embodiment is shown, including: the data acquisition module acquires a series of paired electronic signature data, including a series of signature sample data and signature verification data; the data pre-processing module processes the collected paired electronic signature data, which may include: scale normalization, direction normalization, breakpoint smoothing, Gaussian noise data augmentation, mask data augmentation; the data processing module further processes the data including: calculating manual features, DTW alignment, probability fusion of paired features and labels, inputting the electronic handwriting recognition model to obtain the verification signature score and sample quality score. The data post-processing module is used to calculate and correct the electronic signature recognition model verification score to obtain a more accurate electronic signature handwriting recognition judgment result, specifically including: calculating the maximum, minimum and average values of the verification score, calculating the maximum, minimum and average values of the quality score, fitting the quality score to the T distribution to calculate the sample difference, and correcting the verification score.
[0047] The present invention combines deep learning methods to facilitate the description of data processing at each stage. Given a set of electronic signature data pairs S = {(s1, s'1), (s2, s'2), ..., (s n ,s' n )}, where (s i ,s' i ) is a pair of signatures with the same content, s i With s' i They represent the electronic signature sample data and the electronic signature verification data respectively, and n is the number of electronic signatures. For a single electronic signature data s=[(x0,y0,p0,t0),(x1,y1,p1,t1),…,(x l ,y l ,p l ,t l )], where x i ,y i Represents the electronic signature trajectory data sequence at time t i The horizontal and vertical coordinates at the time, p i Indicates time t i When the pressure value.
[0048] The deep encoder model is used to obtain the deep features of the electronic signature, denoted as f(·). For the convenience of description, the deep (encoder) model is taken as an example here. The encoder can be any deep sequence model to process the manual features of the signature data pair (the deep model can also be replaced by a machine learning method). The deep model calculates the writing features of the electronic signature to obtain the deep feature results, which are used to calculate the scores of the corresponding signature pairs, and then correct the scores in data post-processing.
[0049] In the data processing stage: First, calculate the manual writing features in front of the signed data pair after preprocessing the data, and then align and probabilistically fuse the manual writing features of the data pair. Input the fused paired features into the deep model to calculate the paired deep features (deep representations). The paired deep features are input into the electronic signature handwriting recognition model to calculate the score of the data pair, which is used to calculate the loss and optimize the model during the model training stage, and used as the verification score to determine whether the verification passes during the signature verification stage.
[0050] (1) Data preprocessing is a set of methods for preprocessing the original electronic signature data to obtain normalized electronic signature data. On the one hand, preprocessing operations can eliminate noise interference from signature devices and remove problems such as data non-standardization caused by signature status. On the other hand, adding appropriate data augmentation operations can enhance data diversity to reduce model overfitting, improve the discriminative performance of the model, and enhance model robustness, thus better performing feature learning. Data preprocessing specifically includes one or more of the following data preprocessing operations.
[0051] Scale normalization, remap the collected electronic signature data pair to a standard canvas to ensure the consistency of the input data, and perform normalization processing on the electronic signature data. The specific methods that can be used include: for the i-th electronic signature data s i , according to the minimum and maximum coordinate points of the signature at this point, call the formula:
[0052] Perform standardization processing on the coordinate position (x i , y i ), where (x min , y min )(x max , y max ) respectively represent the minimum coordinate point and the maximum coordinate point in the signature s i . The same method can be used to perform processing such as writing speed normalization, pressure normalization, and direction angle normalization on other data in the electronic signature feature sequence, or other methods can be used to perform normalization processing on the electronic signature feature data to ensure the consistency of the signed handwriting data after processing.
[0053] For example, for direction normalization of electronic signature data, rotate the electronic signature data to the standard direction θ std to remove data differences caused by the writing angle. Translate the coordinate points (x i , y i ) of the collected original electronic signature data s i to achieve the overall rotation transformation of the signature, and calculate the current coordinate point angle difference R = θ std-θ, where the current electronic signature angle θ can be obtained from the center point - relative vector calculation, according to the formula:
[0054]
[0055] Calculate the position of the coordinate point after direction normalization where cos(R) and sin(R) represent the cosine value and sine value corresponding to the angle difference R, respectively.
[0056] Rotate the signature. Rotate according to the difference between the current signature angle and the set standard angle. Translate by calculating the offset corresponding to the angle difference R for each coordinate point, and finally achieve the overall rotation of the signature.
[0057] Breakpoint smoothing processing. Usually, there are discontinuous points on the signature curve caused by hand tremors, device errors, etc. The breakpoint is represented as high - frequency noise in the frequency domain of the data. Smooth the high - frequency components (corresponding to the breakpoints in the electronic signature data) in the frequency domain of the signature data, making the signature curve more continuous. For the original electronic signature s i Smooth the handwriting data. According to the frequency - domain representation (X(ω), Y(ω)) of the coordinate points and the frequency ω, according to the formula:
[0058]
[0059] Modify the frequency - domain representation (X(ω), Y(ω)) of the original signature data to a smoothed frequency - domain representation where X(ω) and Y(ω) represent the frequency - domain representations of the abscissa X and ordinate Y respectively, and ω represents the frequency; ω c represents the cut - off frequency, which determines the cut - off position of the filter in the frequency domain, and n represents the order of the filter, which determines the steepness of the filter.
[0060] The data pre - processing module includes a filter. The breakpoints in the signature data often appear as high - frequency noise in the frequency domain. In the data pre - processing stage, the filter suppresses and smooths the high - frequency components in the frequency domain, reducing the breakpoint problem of the signature data.
[0061] Gaussian noise data augmentation and masked data augmentation. Simulate the interference in the real scenario, improve the complexity of the data, make the model more robust to the interference in the real scenario, and improve the discriminative performance of the model. New point coordinates can be obtained through Gaussian noise data augmentation or masked data augmentation.
[0062] It can be implemented in the following way. Set the noise amplitude A, the frequency f i , y i ) corresponding to the i - th coordinate point (x i , the phase φ of the noise signal, and call the formula:
[0063] G(i)=A·cos(2πf i +φ),
[0064] Calculate the coordinates of the i-th point (x i ,y i ) corresponding to the Gaussian noise signal G(i), determine the coordinates of point i after Gaussian noise data augmentation
[0065]
[0066] Set the i-th coordinate point (x i ,y i ) corresponds to the mask state M(i), where 0 represents mask and 1 represents non-mask, and determines the point coordinates after mask data augmentation
[0067]
[0068] After completing the data pre-processing operation, you can get standardized, standard and more diverse electronic signature sequence data.
[0069] Data processing is to further process the standardized electronic signature data sequence to obtain aligned and fused handwritten signature features, which can make the manual features of the electronic signature pair more diversified, and make it easier for the model to focus on the differences between paired data (such as signature style changes) when entering the model training, effectively improving the identification ability and robustness of handwriting recognition models and methods.
[0070] In the handwriting recognition task, there are usually certain differences in the handwriting speed and style between the electronic sample signature and the electronic verification signature. The sample-verification signature pair can be aligned through the DTW (Dynamic Time Warping) operation, so that the handwriting recognition model can overcome the time offset and speed change between the signature pair features. However, using only DTW will have problems such as noise sensitivity and inability to handle local changes, making it difficult for the model to focus on the subtle differences between signature pairs and poor robustness. Therefore, the present invention adopts paired data feature fusion based on DTW, which can make the model easier to learn the differences between signature pairs and improve the robustness of the model. The data feature fusion method in the field of technology can also be used.
[0071] Specifically, after data pre-processing, the electronic signature data pair is obtained. Furthermore, the electronic signature manual features of the input model are calculated, such as the horizontal speed v x , vertical speed v y 、Path tangent angle θ=arctan(v y / v x) Cosine value cos(θ) of the path tangent angle, sine value sin(θ), etc., and this set of manual features of the electronic signature is briefly denoted as z. Use DTW to align the original manual features (z, z') of the electronic signature data pair to obtain
[0072] Further, randomly sample a random factor λ from the Beta distribution of the electronic signature stroke trajectory,
[0073]
[0074] where B(a, b) is the Beta function for normalizing the Beta distribution, and a, b are shape parameters,
[0075]
[0076] where Γ is the gamma function. Further, fuse the aligned manual features of the electronic signature and the corresponding labels (label, label') with the random factor λ to obtain the fused feature and the mixed label For example, the following formula can be used:
[0077]
[0078]
[0079] to obtain the fused feature and the mixed label Through the fused manual features of the electronic signature the handwriting recognition model can focus on the pairwise depth features with finer differences corresponding to the fused label so as to improve the discrimination performance of the model for signature pair data.
[0080] At the same time, the fused features have rich diversity, enabling the model to generalize better and improving the robustness of the model. Finally, generate the verification score score of this signature pair from the pairwise depth representation, which can be expressed as:
[0081]
[0082] The verification score score is used to calculate the loss to optimize the model during the training model stage and is used as the basis for determining whether the electronic signature to be verified is written by the user himself during the verification stage of the electronic signature.
[0083] Data post-processing is a correction method to improve the correct rate of handwriting recognition during the verification process of multi-retained handwriting samples. By adaptively adjusting the retained-verification signature score according to the quality of the signature retained samples, the final accuracy of handwriting recognition can be effectively improved.
[0084] Specifically, due to factors such as real-world interference in the signature scenario and the habits of the writer himself, there will be differences between multiple electronic retained signatures, and between the electronic verification signature and the electronic retained signature. Data preprocessing aims to correct and enhance the original data, while data postprocessing aims to further correct the verification result based on the quality of the retained samples and give a more accurate and reliable result.
[0085] Suppose there are three retained signatures (s1, s2, s3) for the electronic signature s' to be verified, then three electronic signature verification data pairs ((s1, s'), (s2, s'), (s3, s')) and three signature retention data pairs ((s1, s2), (s1, s3), (s2, s3)) can be formed. After the above data preprocessing and data midprocessing, three verification scores (score'1, score'2, score'3) and three retained sample quality scores (score1, score2, score3) can be obtained by the handwriting recognition model.
[0086] For the verification scores, respectively obtain the maximum score' max and the minimum score' min and calculate the average value of the verification scores score' mean , which can be specifically expressed as:
[0087] score' max = max(score'1, score'2, score'3),
[0088] score' min = min(score1, score2, score3),
[0089] score' mean = (score'1 + score'2 + score'3) / 3.
[0090] Similarly, the maximum score max , the minimum score min and the average value score mean of the retained sample quality scores can be obtained. Further, the verification scores are corrected according to the difference degree between the retained samples. In this exemplary implementation, the following method can be used for calculation:
[0091]
[0092]
[0093]
[0094] Among them, α, β, and γ are the weights of the quality of three retained samples. Specifically, the quality score score of the retained sample s1 is obtained by (score1 + score2) / 2 s1 , the quality score score of the retained sample s2 is obtained by (score1 + score3) / 2 s2 , the quality score score of the retained sample s3 is obtained by (score2 + score3) / 2 s3 , the higher the quality score indicates that the retained sample is more similar to the other two retained samples, that is, the better the quality. Fit the t-distribution t(μ, σ) with 1 degree of freedom, where its mean μ is the score of the retained sample with the best quality, and its standard deviation σ is:
[0095]
[0096] Furthermore, under this t(μ, σ) distribution, the probability values (α, β, γ) corresponding to the quality scores (score1, score2, score3) of the retained samples are obtained.
[0097] Finally, the final score score' of this verification signature is calculated. In this exemplary embodiment, the following method can be specifically used for calculation:
[0098] score' = (score' new_max + score' new_min + score' new_mean ) / 3.
[0099] This method comprehensively considers the quality of the retained samples to give greater weight to high-quality retained samples, and at the same time dynamically corrects the most similar retained sample-verification signature pairs and the most different retained sample-verification signature pairs, and finally obtains a more reliable score.
[0100] Data preprocessing performs a series of data preprocessing on the non-normalized electronic signature to obtain more standardized electronic signature data: correct problems such as different scales, inconsistent directions, and breakpoints in the electronic signature, and can improve the robustness of the model in the subsequent process of training the handwriting recognition model.
[0101] Data middle processing aligns and fuses the original manual features of the normalized electronic signature data pairs to obtain more diverse paired fusion features: making it easier for the model to focus on the differences between paired data (such as signature style transformation), effectively improving the discrimination ability and robustness of the handwriting recognition model.
[0102] Through the handwriting recognition method, the model calculates the initial score of the verification signature from the paired fusion features.
[0103] In data post-processing, the sample-verification signature score is adaptively adjusted according to the signature sample quality: high-quality samples are given a larger weight, and the most similar sample-verification signature pairs and the sample-verification signature pairs with the largest difference are dynamically corrected, ultimately obtaining a more accurate handwriting verification result.
[0104] like Figure 3 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the device 300 can also be stored. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0105] A plurality of components in the electronic device 300 are connected to the I / O interface 305, including: an input unit 306, an output unit 307, a storage unit 308, and a communication unit 309. The input unit 306 may be any type of device capable of inputting information to the electronic device 300, and the input unit 306 may receive input digital or character information, and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 307 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 308 may include, but is not limited to, a disk, an optical disk. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0106] The computing unit 301 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 executes the various methods and processes described above. For example, in some embodiments, the reconstruction and decomposition of the muscle movement trajectory according to the original trajectory of the signature stroke, and the decomposition of its logarithmic velocity curve, etc. can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 300 via the ROM 302 and / or the communication unit 309. In some embodiments, the computing unit 301 can be configured to execute the signature handwriting dynamic acquisition implementation method by any other suitable means (e.g., by means of firmware).
[0107] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0108] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0109] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, an optical disk, a memory, a programmable logic device (PLD)) that provides machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal that provides machine instructions and / or data to a programmable processor.
[0110] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0111] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0112] A computer system can include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
Claims
1. A multi-stage data processing method for handwriting recognition, characterized in that Including: Data pre - processing: Collecting electronic signature retention data, obtaining electronic signature verification data, and standardizing the electronic signature retention data and the electronic signature verification data; Data mid - processing: Extracting the fusion features of the electronic signature retention data and respectively fusing them with the electronic signature verification data to obtain the retention - verification data pair fusion features; Data post - processing: Correcting, verifying the scores of the electronic signature data pairs output by the handwriting recognition model, and outputting a judgment result.
2. The method according to claim 1, wherein The standardization of the data includes: Performing one or several pre - treatments on non - standardized data and the noise in the data, including but not limited to: scale normalization, direction normalization, breakpoint smoothing, Gaussian noise data augmentation, mask data augmentation, etc., to standardize the signature data; The data fusion includes: Fusing the data pair features based on dynamic time warping, completing the signature style transformation, and generating more diverse paired features; The correction of the scores of the electronic signature data pairs includes: Adaptively adjusting the correction value of the retention - verification signature scores according to the signature retention quality during handwriting verification.
3. The method according to claim 2, wherein The normalization includes: according to the minimum and maximum coordinates (x min , y min ) and (x max , y max ) of the original signature data trace points, call the formula: Obtain the coordinates of the trace points after scale normalization According to the angle θ of the original signature trace coordinate point and the standard rotation angle θ std , according to the formula: R = θ std - θ to calculate the angle difference of the original signature trace coordinate point, according to the formula: Calculate the coordinate position of the trajectory points after direction normalization Translate the offset corresponding to the angle difference of each coordinate point to achieve the rotation of the whole signature.
4. The method according to claim 2, wherein The breakpoint smoothing process includes: according to the frequency domain representations X(ω) and Y(ω) of the abscissa and ordinate, the frequency ω, and the cut-off frequency ω c of the filter, and the order n of the filter, calling the formula: Modify the coordinate frequency-domain representation (X(ω), Y(ω)) of the original signature data to a smoothed coordinate frequency-domain representation Complete the smoothing operation of the high-frequency noise components in the frequency-domain representation of the original electronic signature data.
5. The method according to claim 2, wherein The Gaussian noise data augmentation and masked data augmentation include: According to the noise amplitude A, noise signal frequency f i , y i ) and noise signal phase φ of the original electronic signature data trajectory point coordinates (x i , call the formula: G(i) = A·cos(2πf i + φ), Calculate the Gaussian noise signal G(i); determine the coordinates after Gaussian noise data augmentation which are: Determine the coordinates after masked data augmentation according to the mask state M(i). They are as follows:
6. The method according to claim 2, wherein The adaptive adjustment of the correction value of the retention - verification signature scores further includes: Calculating the quality scores of each retention and giving a higher weight to high - quality retentions than low - quality retentions according to the quality scores, and dynamically correcting the retention - verification signature pairs with the smallest difference and the retention - verification signature pairs with the largest difference according to the weights.
7. The method according to any one of claims 1 to 6, characterized in that The fusion of the data pair features further includes: Calculating the signature writing features, aligning the data pair features based on dynamic time warping, probabilistically fusing the data pair features and labels, and inputting into the electronic handwriting recognition model to obtain the retention - verification signature scores and the retention quality scores; The adaptive adjustment of the correction value of the retention - verification signature scores includes: Calculating the maximum, minimum, and average values of the retention - verification signature scores, calculating the maximum, minimum, and average values of the quality scores, and fitting the quality scores to the T - distribution to calculate the retention difference degree and correct the retention - verification signature scores.
8. The method according to any one of claims 1 to 6, characterized in that The probability fusion pair features and labels further include: randomly sampling a random factor λ from the Beta distribution of the electronic signature stroke trajectory, and fusing the aligned electronic signature handwriting features and the corresponding labels (label, label') with the random factor λ to obtain fused features and the mixed labels 9. A multi-stage data processing system for handwriting recognition, characterized in that, Including: Data pre - processing module: Used to collect electronic signature retention data, obtain electronic signature verification data, and standardize the electronic signature retention data and the electronic signature verification data; Data mid - processing module: Used to fuse the fusion features of the electronic signature retention data pairs with the signature verification data respectively to obtain the retention - verification data pair fusion features; Data post - processing module: Used to correct, verify the scores of the electronic signature data pairs output by the handwriting recognition model, and output a judgment result.
10. The system according to claim 9, wherein, The standardization of the data includes: Performing one or several pre - treatments on non - standardized data and the noise in the data, including but not limited to: scale normalization, direction normalization, breakpoint smoothing, Gaussian noise data augmentation, mask data augmentation, etc., to standardize the signature data; The data fusion includes: Fusing the data pair features based on dynamic time warping, completing the signature style transformation, and generating more diverse paired features; The correction of the scores of the electronic signature data pairs includes: Adaptively adjusting the correction value of the retention - verification signature scores according to the signature retention quality during handwriting verification.
11. The system according to claim 10, wherein The correction value for adaptively adjusting the sample-validation signature score further includes: calculating the quality score of each sample and giving a higher weight to high-quality samples than to low-quality samples according to the quality score, and dynamically correcting the sample-validation signature pair with the smallest difference and the sample-validation signature pair with the largest difference according to the weights; The fusion of data pair features further includes: calculating signature writing features, aligning data pair features based on dynamic time warping, probabilistically fusing data pair features and labels, and inputting into an electronic handwriting recognition model to obtain the sample-validation signature score and the sample quality score.
12. The system according to claim 10 or 11, characterized in that, The correction value for adaptively adjusting the sample retention-verification signature score includes: calculating the maximum, minimum, and average values of the sample retention-verification signature score, calculating the maximum, minimum, and average values of the quality score, and calculating the sample retention difference degree by fitting the quality score to the T distribution to correct the sample retention-verification signature score; The probability fusion of paired features and labels further includes: randomly sampling a random factor λ from the Beta distribution of the electronic signature stroke trajectory, and fusing the aligned electronic signature handwriting features and the corresponding labels (label, label') with the random factor λ to obtain the fused features and the mixed labels 13. An electronic device, comprising: A processor; And a memory storing a program, characterized in that the program includes instructions which, when executed by the processor, cause the processor to execute the multi-stage data processing method for handwriting recognition according to any one of claims 1-7.
14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, Wherein, The computer instructions are for causing the computer to execute the multi-stage data processing method for handwriting recognition according to any one of claims 1-7.
Citation Information
Patent Citations
Signature handwriting authentication method integrating point features, local features and global features, storage medium and electronic equipment
CN115862153A