Handwritten content authentication method and system based on image recognition and electronic equipment
Through the improved ResNet convolutional neural network and lightweight preprocessing technology, the problems of high efficiency, high precision and high concurrency in handwritten content recognition in consumer finance scenarios have been solved, and highly accurate recognition of complex and dynamic handwriting has been achieved, improving user experience and business process efficiency.
Patent Information
- Application Number
- CN202510831292.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies cannot balance the requirements of high efficiency, high precision, and high concurrency in consumer finance scenarios. Traditional machine learning algorithms are difficult to adapt to dynamic handwriting on mobile devices. Shallow networks have low accuracy in recognizing complex handwriting and are easily affected by the environment. The gradient vanishing problem of fully connected neural networks makes model training difficult and lacks a real-time optimization mechanism.
The ResNet convolutional neural network model is used, combined with lightweight preprocessing and dynamic optimization methods, and handwritten images are generated through a rasterization algorithm. Residual blocks and batch normalization layers are used to optimize model training. Kalman filtering and morphological operations are introduced to process handwriting. Retry and manual review processes are set up, and a cloud-based backup model is used to process uncollected characters.
It improves the accuracy and stability of handwritten content recognition, optimizes the model training process, enhances the adaptability to different environments, reduces the misjudgment rate, and provides a good user experience and business process flexibility.
Smart Images

Figure CN120673423A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of image recognition technology, and in particular to a handwritten content authentication method, system, electronic device, and storage medium based on image recognition. Background Art
[0002] In traditional banking, customers often need to handwrite their signatures or fill in digital information on specialized equipment at bank branches (such as signature pads and touch screen terminals) when applying for loans or signing contracts. Handwritten content recognition in these scenarios primarily relies on traditional machine learning algorithms (such as support vector machines (SVMs) and random forests) or shallow convolutional neural networks (such as AlexNet). These technical features include: Based on static image recognition: After the customer's signature is scanned or photographed as an image, it is matched and verified through feature extraction (such as HOG, SIFT) and classifiers.
[0003] Low-concurrency processing: In traditional banking scenarios, customer signing behavior is relatively dispersed, the system load is low, and the computational efficiency of existing algorithms can still meet the needs.
[0004] Fixed environment collection: Signatures are usually completed in a controlled environment (such as bank counter equipment), where factors such as lighting and writing angle are relatively stable, reducing the difficulty of recognition.
[0005] In addition, some solutions use fully connected neural networks for handwritten content recognition, but their training process requires processing a large number of parameters and is easily affected by gradient vanishing or gradient explosion problems.
[0006] In consumer finance scenarios (such as real-time mobile credit approval and online contract signing), these traditional technologies expose the following key issues: 1. Insufficient recognition efficiency. Traditional machine learning algorithms rely on manual feature engineering and are difficult to adapt to the dynamic handwriting (such as intermittent trajectory and pressure changes) of customers writing on mobile devices such as smartphones.
[0007] 2. In high-concurrency scenarios (such as a large number of users signing contracts simultaneously during a promotion), the GPU computing power requirements of traditional models increase exponentially, resulting in response delays.
[0008] 3. The model has poor generalization ability. Shallow networks (such as AlexNet) have low recognition accuracy for complex handwriting (such as cursive characters and rare characters) and are easily affected by the mobile terminal collection environment (such as screen size differences and lighting interference). Fully connected neural networks have difficulty training deep networks due to the gradient vanishing problem, resulting in limited ability of the model to capture subtle handwriting features.
[0009] 4. Lack of real-time optimization mechanism, Existing solutions do not perform dynamic pre-processing on the time series data (such as coordinate sequences and pen pressure) collected by mobile terminals, resulting in noise data directly affecting the recognition effect. Abnormal situations (such as low-quality signatures and unrecorded characters) are usually directly rejected, and there is a lack of hierarchical processing (such as retry guidance and manual review linkage), which affects the user experience.
[0010] In summary, existing technologies cannot balance the requirements of high efficiency, high precision, and high concurrency in consumer finance scenarios. There is an urgent need for a handwritten content recognition solution that combines lightweight preprocessing, deep residual networks, and dynamic optimization. Summary of the Invention
[0011] Embodiments of the present invention provide a handwritten content authentication method, system, electronic device, and storage medium based on image recognition to solve the problem that the existing system cannot balance the requirements of high efficiency, high precision, and high concurrency in consumer finance scenarios.
[0012] In a first aspect, an embodiment of the present invention provides a method for authenticating handwritten content based on image recognition, comprising: S1. Obtaining a coordinate sequence of user handwritten content, sorting the coordinate sequence by timestamp, distinguishing the start and end of a stroke based on pen pressure data, and generating a handwritten image using a rasterization algorithm; wherein the coordinate sequence includes xy coordinates, timestamp, and pen pressure data; S2. Preprocessing the handwritten image to generate a binary image of the handwritten image; S3. Input the binary image into a pre-trained recognition model to obtain the recognition result and confidence level. If the confidence level is higher than a preset threshold, the recognition result is verified with the customer's registered name. If the verification is successful, the business process continues.
[0013] Preferably, the step S3 further includes: If the confidence level is lower than the preset threshold, the customer is prompted to rewrite the signature; If the number of rewrites exceeds the preset number, the manual review process will be initiated.
[0014] Preferably, the step S3 further includes: If the verification fails, the customer is prompted to rewrite the signature; If the number of rewrites exceeds the preset number, the manual review process will be initiated.
[0015] Preferably, the recognition model is a ResNet convolutional neural network model; The ResNet convolutional neural network model includes several cascaded residual blocks, each of which includes a direct path to the input feature map X and a convolution operation path. The convolution operation path includes at least two convolution layers, and a ReLU activation function is set between the two convolution layers. When the dimension of the input feature map X is inconsistent with the output dimension of the convolution operation path, the input feature map is dimensionally matched through a 1×1 convolution layer. A batch normalization layer is set after the convolution operation path of each residual block. The batch normalization layer performs batch normalization on the input data of each layer. By calculating the mean and variance of neurons in each layer, the neuron values are standardized. This ensures that the numerical differences between neurons in each layer are not too large, thereby optimizing the model training process and improving the model convergence speed and generalization ability. It also includes a hyperparameter optimization module that uses the Bayesian optimization algorithm to iteratively optimize the number of residual blocks, convolution kernel size, learning rate decay strategy, and batch normalization regularization parameters.
[0016] Preferably, the step S2 specifically includes: S21, performing trajectory smoothing processing on the coordinate sequence using Kalman filtering; S22, normalizing the pen pressure data through adaptive pressure adjustment; S23, scaling the trajectory to the target size; S24, drawing an anti-aliased curve on a blank canvas to generate a binary image, wherein the line width is determined by multiplying the pen pressure value by 3; S25, performing a morphological closing operation on the binary image, with a 3×3 diamond kernel as the structural element, to connect the disconnected strokes; S26. Add uniformly distributed random noise at the image pixel level, and limit the jitter range to within 3 pixels around the stroke through a mask.
[0017] Preferably, the step S3 further includes: When the recognition result contains unrecorded characters, upload the character samples to the cloud and call the backup model to return a temporary recognition result; The backup model is a cloud-based auxiliary recognition model used to process unrecorded characters that cannot be recognized by the recognition model. The unrecorded characters include uncommon characters and rare symbols.
[0018] Preferably, the recognition result is matched with the characters pre-verified by the customer, and a secondary verification is performed in combination with the topological structure characteristics of the signature trajectory.
[0019] In a second aspect, an embodiment of the present invention provides a handwritten content authentication system based on image recognition, comprising: A handwriting data acquisition module, which acquires a coordinate sequence of the user's handwriting content, sorts the coordinate sequence by timestamp, distinguishes the start and end of the strokes based on the pen pressure data, and generates a handwritten image using a rasterization algorithm; wherein the coordinate sequence includes xy coordinates, timestamps, and pen pressure data; A preprocessing module, which preprocesses the handwritten image to generate a binary image of the handwritten image; The recognition module inputs the binary image into a pre-trained recognition model to obtain the recognition result and confidence level. If the confidence level is higher than a preset threshold, the recognition result is verified with the customer's registered name. If the verification is successful, the business process continues.
[0020] In a third aspect, an embodiment of the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the handwriting content authentication method based on image recognition as described in the embodiment of the first aspect of the present invention are implemented.
[0021] In a fourth aspect, an embodiment of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the handwriting content authentication method based on image recognition as described in the embodiment of the first aspect of the present invention are implemented.
[0022] The embodiments of the present invention provide a method, system, electronic device and storage medium for handwritten content authentication based on image recognition. The method enhances feature expression capability through an improved residual structure, adopts a ResNet convolutional neural network, and avoids the gradient vanishing or exploding problem in traditional convolutional neural networks through the propagation path of "residual blocks", thereby improving the model's recognition accuracy and stability for handwritten content. The multimodal feature interaction mechanism simultaneously utilizes spatiotemporal information, and the recognition effect of cursive characters and sloppy signatures is significantly better than traditional solutions. Batch normalization technology is introduced to standardize the input data of each layer, so that the model can dynamically adapt to input data of different scales, optimize the model training process, and improve the model convergence speed and generalization ability. A low-quality signature processing process is set. When the recognition confidence is lower than the preset threshold, the system automatically prompts the user to rewrite the signature, and starts the manual review process when the number of retries exceeds the preset number. The confidence calibration technology makes the output probability more reliable and reduces the misjudgment rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0024] Figure 1 is a flowchart of a handwritten content authentication method based on image recognition according to an embodiment of the present invention; Figure 2 This is a specific flow chart of a handwritten content authentication method based on image recognition according to an embodiment of the present invention.
[0025] Figure 3 is a block diagram of a handwriting content authentication system based on image recognition according to an embodiment of the present invention; Figure 4 Schematic diagram of the physical structure according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0027] In the embodiments of the present application, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent three situations: A exists alone, A and B exist at the same time, and B exists alone.
[0028] The terms "first" and "second" in the embodiments of the present application are only used for descriptive purposes and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusions. For example, a system, product or device comprising a series of components or units is not limited to the listed components or units, but may optionally also include components or units that are not listed, or may optionally also include other components or units that are inherent to these products or devices. In the description of the present application, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0029] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0030] In consumer finance scenarios (such as real-time mobile credit approval and online contract signing), these traditional technologies expose the following key issues: 1. Insufficient recognition efficiency. Traditional machine learning algorithms rely on manual feature engineering and are difficult to adapt to the dynamic handwriting (such as intermittent trajectory and pressure changes) of customers writing on mobile devices such as smartphones.
[0031] 2. In high-concurrency scenarios (such as a large number of users signing contracts simultaneously during a promotion), the GPU computing power requirements of traditional models increase exponentially, resulting in response delays.
[0032] 3. The model has poor generalization ability. Shallow networks (such as AlexNet) have low recognition accuracy for complex handwriting (such as cursive characters and rare characters) and are easily affected by the mobile terminal collection environment (such as screen size differences and lighting interference). Fully connected neural networks have difficulty training deep networks due to the gradient vanishing problem, resulting in limited ability of the model to capture subtle handwriting features.
[0033] 4. Lack of real-time optimization mechanism, Existing solutions do not perform dynamic pre-processing on the time series data (such as coordinate sequences and pen pressure) collected by mobile terminals, resulting in noise data directly affecting the recognition effect. Abnormal situations (such as low-quality signatures and unrecorded characters) are usually directly rejected, and there is a lack of hierarchical processing (such as retry guidance and manual review linkage), which affects the user experience.
[0034] In summary, existing technologies cannot balance the requirements of high efficiency, high precision, and high concurrency in consumer finance scenarios. There is an urgent need for a handwritten content recognition solution that combines lightweight preprocessing, deep residual networks, and dynamic optimization.
[0035] Therefore, the embodiment of the present invention provides a handwritten content authentication method based on image recognition, such as Figure 1 、 Figure 2 As shown, including: S1. Obtaining a coordinate sequence of user handwritten content, sorting the coordinate sequence by timestamp, distinguishing the start and end of a stroke based on pen pressure data, and generating a handwritten image using a rasterization algorithm; wherein the coordinate sequence includes xy coordinates, timestamp, and pen pressure data; Specifically, a coordinate sequence refers to the track points left by the user's handwriting on a mobile device or dedicated device, including a multidimensional data set of the xy coordinates, timestamp, and pen pressure data of each point. It is a digital record of the handwriting content. The timestamp is used to mark the order of strokes to ensure that the coordinate sequence is arranged in the order of the user's writing, so that subsequent processing can accurately reflect the user's writing trajectory; the pen pressure data can reflect the writing force; it can be used to distinguish different stages of the user's writing, such as the start, duration, and end of the stroke. During the writing process, the pen pressure may change. For example, the pen pressure is small at the beginning, gradually increases during the writing process, and decreases or disappears at the end. The start and end points of the strokes can be distinguished based on the pen pressure data. By analyzing the pen pressure data, the starting and ending points of each stroke can be accurately identified.
[0036] A rasterization algorithm is a method for converting a continuous handwritten trajectory into a discrete image. In this embodiment, the algorithm converts a coordinate sequence into a handwritten image, where each trajectory point corresponds to a pixel on the image. By connecting these pixels, the shape and structure of the handwritten content are formed.
[0037] Furthermore, when a user signs a credit contract using a smartphone, the smartphone application records the user's handwritten trace points, including each point's xy coordinates, timestamp, and pen pressure data. The coordinate sequence is sorted by timestamp, and the start and end of the stroke are distinguished based on the pen pressure data. Using a rasterization algorithm, the coordinate sequence is converted into a handwritten image with a shape and structure consistent with the user's handwriting.
[0038] S2. Preprocessing the handwritten image to generate a binary image of the handwritten image; Specifically, preprocessing refers to a series of optimization processes on handwritten images, including trajectory smoothing, pen pressure normalization, size standardization, binary image generation and image enhancement, to improve image quality and provide better input for subsequent recognition.
[0039] Specifically, the step S2 includes: S21, performing trajectory smoothing processing on the coordinate sequence using Kalman filtering; Specifically, the Kalman filter trajectory smoothing process includes: State definition: The state of each coordinate point is set as a four-dimensional vector (x coordinate, y coordinate, x-direction speed, y-direction speed), and the current position is predicted based on the state at the previous moment; Noise modeling: Build the covariance matrix of process noise (writing jitter) and observation noise (device acquisition error), and dynamically adjust the smoothing strength. Perform a "prediction-correction" loop for each coordinate point: Prediction: Calculate the current position and speed based on the previous state; Correction: Update the prediction results with the actual collected coordinates to reduce the impact of noise.
[0040] Output result: The original jittered coordinate sequence is converted into a smooth and continuous trajectory, and the error is reduced from ±5 pixels to ±1 pixel.
[0041] S22, normalizing the pen pressure data through adaptive pressure adjustment; Specifically, the minimum value of all pen pressure data of the current signature is counted ( P min ) and the maximum value ( P max ), through the formula P norm =[( P - P min ) / ( P max - P min )]×0.8+0.2, mapping the pen pressure range from the original interval (such as [0.5, 0.9]) to [0.2, 1.0] to avoid abnormal line width caused by extreme values. P Represents raw pen pressure data, that is, the real-time pressure value collected by the device as the user writes. This value reflects the writing force and is typically generated by touch devices (such as touchscreens and digital pens) based on the pressure applied by the pen tip or finger to the screen.
[0042] S23, scaling the trajectory to the target size; Compare the width and height of the original trajectory with the target size (400×64 pixels), and take the smaller value of the width and height scaling ratio (for example, if the original width is 200 and the height is 80, then the width ratio is 400 / 200=2, and the height ratio is 64 / 80=0.8, so 0.8 is finally taken); scale the trajectory size according to the selected ratio (for example, the original 200×80 is scaled to 160×64), keeping the aspect ratio unchanged.
[0043] S24, drawing an anti-aliased curve on a blank canvas to generate a binary image, wherein the line width is determined by multiplying the pen pressure value by 3; Directly drawing a line segment can produce jagged edges (stair-stepped edges), especially for diagonal lines. Therefore, to convert the preprocessed handwriting trajectory (coordinate sequence + pen pressure data) into a high-quality binary image, this embodiment employs an anti-aliasing algorithm (such as the Xiaolin Wu algorithm) to smooth edges through grayscale transitions. For pixels traversed by a line segment, the grayscale value is calculated based on the area covered. The coefficient of 3 is an empirical value intended to amplify differences in strokes under varying pressures. For example, if the pen pressure range is narrow (e.g., 0.1-0.3), multiplying by 3 will result in more pronounced line width changes (0.3-0.9 px), making it easier for the model to capture features. Binarization reduces data volume, making convolutional neural networks more efficient for processing black and white images, and also adapting to the processing requirements of subsequent convolutional networks such as ResNet. Anti-aliasing optimizes edge quality, not preserves grayscale information.
[0044] S25, performing a morphological closing operation on the binary image, with a 3×3 diamond kernel as the structural element, to connect the disconnected strokes; A 3×3 diamond kernel (with all 1s in the middle row and column and 0s in the corners) is used to detect and connect stroke breakpoints. Using a 3×3 diamond kernel as a structuring element is a key step in performing morphological closing on binary images. Morphological closing is an image processing technique that combines two basic operations, dilation and erosion, to fill small holes in an image, connect adjacent objects or regions, and smooth object outlines.
[0045] Dilation: The dilation operation expands the white areas (i.e., handwriting) in the image by sliding a structuring element across the image and assigning the maximum value of all pixels covered by the structuring element to the center pixel. This helps connect adjacent strokes and reduces disconnections caused by jitter or noise during handwriting.
[0046] Erosion: The erosion operation slides the structuring element across the image and assigns the minimum value of all pixels covered by the structuring element to the center pixel, thereby reducing the white area in the image. However, in the closing operation, the erosion operation is usually followed by the dilation operation to eliminate small noise points or poorly connected areas that may be generated during the dilation process, while maintaining the overall shape and size of the object.
[0047] 3×3 diamond kernel: The diamond kernel is a commonly used structure element shape that is particularly effective for connecting horizontal and vertical strokes. The 3×3 size means that the structure element covers a 3×3 pixel area, which is a suitable size for processing handwritten images with rich details.
[0048] Morphological closing is performed on binary images using a 3×3 diamond kernel as the structuring element. A dilation operation first expands the white areas in the image, connecting adjacent strokes. An erosion operation then removes any minor noise or poorly connected areas that may have been introduced during the dilation process, while maintaining the overall shape and size of the object.
[0049] S26. Add uniformly distributed random noise at the image pixel level, and limit the jitter range to within 3 pixels around the stroke through a mask.
[0050] To simulate the natural jitters of handwriting, enhance the diversity and robustness of handwritten images, and thus improve the generalization ability of subsequent recognition algorithms, this embodiment adds uniformly distributed random noise at the image pixel level. This simulates stroke discontinuities or slight variations caused by factors such as slight hand jitter and device precision limitations during handwriting. This noise helps train the recognition algorithm to better adapt to changes in the real handwriting environment.
[0051] To ensure that the added noise does not disrupt the overall structure and readability of the handwritten image, a mask is used to limit the range of noise addition. Specifically, the mask can be set to be effective only within 3 pixels around the stroke. This way, the noise is only added near the stroke and does not affect other parts of the image.
[0052] S3. Input the binary image into a pre-trained recognition model to obtain a recognition result and confidence level, wherein: If the confidence level is higher than the preset threshold, the recognition result will be verified with the customer's registered name. If the verification is successful, the business process will continue.
[0053] If the confidence level is lower than the preset threshold (0.65), the signature will be displayed as unclear, and the customer will be prompted to rewrite the signature, and the number of rewrites will be recorded; If the verification fails, the client is prompted to rewrite the signature and the number of rewrites is recorded; If the number of rewrites exceeds the preset number, the manual review process will be initiated and the signature will be reviewed by a manual reviewer.
[0054] The method provided by this invention effectively controls business processes based on handwritten image recognition, ensuring the authenticity and validity of signatures. Furthermore, through appropriate rewrite prompts and a manual review mechanism, it provides a good user experience and business process flexibility. Experimental results show that the application of this method for business process control significantly improves both customer satisfaction and business process efficiency.
[0055] Based on the above embodiment, the recognition model is a ResNet convolutional neural network model; The ResNet convolutional neural network model includes several cascaded residual blocks, each of which includes a direct path to the input feature map X and a convolution operation path. The convolution operation path includes at least two convolution layers, and a ReLU activation function is set between the two convolution layers. When the dimension of the input feature map X is inconsistent with the output dimension of the convolution operation path, the input feature map is dimensionally matched through a 1×1 convolution layer. A batch normalization layer is set after the convolution operation path of each residual block. The batch normalization layer performs batch normalization on the input data of each layer. By calculating the mean and variance of neurons in each layer, the neuron values are standardized. This ensures that the numerical differences between neurons in each layer are not too large, thereby optimizing the model training process and improving the model convergence speed and generalization ability. It also includes a hyperparameter optimization module that uses the Bayesian optimization algorithm to iteratively optimize the number of residual blocks, convolution kernel size, learning rate decay strategy, and batch normalization regularization parameters.
[0056] In this embodiment, environmental data (such as light intensity and device orientation) is also collected. Appropriate lighting conditions can ensure the clarity of the handwritten image and reduce image distortion caused by shadows or excessive brightness. By collecting light intensity data, targeted adjustments can be made during the image preprocessing stage, such as brightness enhancement or reduction, to improve image quality, thereby providing more reliable input for subsequent handwritten content recognition. The device orientation (such as whether the screen is placed horizontally) will affect the acquisition effect of the handwritten image. For example, if the device is tilted, it may cause the handwritten track to be distorted or deformed in the image. By collecting device orientation data, rotation or correction can be performed during the image preprocessing stage to restore the original form of the handwritten image and improve recognition accuracy.
[0057] Different lighting conditions and device orientations can cause changes in the visual characteristics of handwritten images. By collecting this environmental data and factoring it into model training, the model can learn a wider range of handwritten image features, improving its generalization capabilities. This means the model can maintain high recognition accuracy when faced with handwritten images under different lighting conditions and device orientations.
[0058] Under certain extreme lighting conditions or device orientations, handwritten image recognition accuracy may decrease. By collecting environmental data, real-time monitoring and judgment can be performed during the recognition process. When recognition accuracy falls below a preset threshold, the system can use environmental data to prompt the user to adjust collection conditions (such as improving lighting or adjusting device orientation), or initiate a manual review process to ensure business security.
[0059] Collected environmental data can also be used as part of the data record for subsequent data analysis and mining. For example, the distribution of handwritten image recognition accuracy under different lighting conditions and device orientations can be analyzed to provide a basis for further optimizing the model or adjusting the collection process.
[0060] Based on the above embodiment, step S3 further includes: When the recognition result contains unrecorded characters, upload the character samples to the cloud and call the pre-trained backup model to return a temporary recognition result; The backup model is a cloud-based auxiliary recognition model used to process unrecorded characters that cannot be recognized by the recognition model. The unrecorded characters include uncommon characters and rare symbols.
[0061] After receiving the character sample, the cloud server calls a pre-trained backup model for recognition. This backup model is specifically designed to handle unrecognized characters that the recognition model cannot recognize, offering broader character coverage and higher recognition accuracy. After the backup model completes recognition, it returns the provisional recognition result to the local system, allowing the business process to continue.
[0062] On the basis of the above embodiment, the recognition result is matched with the characters pre-verified by the customer, and a secondary verification is performed in combination with the topological structure characteristics of the signature trajectory.
[0063] Perform string matching on the recognition results (including the recognition results of the main model and the temporary recognition results returned by the backup model) with the characters pre-verified by the customer to preliminarily verify the accuracy of the recognition results.
[0064] For handwritten signatures or characters, extract their topological structural features, such as stroke order, stroke connection method, stroke direction, etc. These features can reflect the uniqueness and consistency of handwritten content.
[0065] Secondary verification combines string matching results and topological structure features to perform secondary verification on the recognition results. If the string match is successful and the topological structure features are consistent, the recognition result is considered accurate; if there is any inconsistency, the user is prompted to rewrite or initiate a manual review process.
[0066] In a second aspect, an embodiment of the present invention provides a handwritten content authentication system based on image recognition, comprising: The handwriting data acquisition module 310 acquires a coordinate sequence of the user's handwriting content, sorts the coordinate sequence by timestamp, distinguishes the start and end of the strokes based on the pen pressure data, and generates a handwriting image using a rasterization algorithm. The coordinate sequence includes xy coordinates, timestamps, and pen pressure data. A preprocessing module 320 preprocesses the handwritten image to generate a binary image of the handwritten image; The recognition module 330 inputs the binary image into a pre-trained recognition model to obtain the recognition result and confidence level. If it is determined that the confidence level is higher than a preset threshold, the recognition result is verified with the customer's registered name. If the verification is successful, the business process continues.
[0067] Based on the same concept, the embodiment of the present invention also provides a schematic diagram of an entity structure, such as Figure 4 As shown, the server may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. The processor 410, the communication interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the steps of the handwriting content authentication method based on image recognition as described in the above embodiments.
[0068] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0069] Based on the same concept, an embodiment of the present invention also provides a non-transitory computer-readable storage medium, which stores a computer program. The computer program includes at least one section of code, which can be executed by a main control device to control the main control device to implement the steps of the handwriting content authentication method based on image recognition as described in the above embodiments.
[0070] Based on the same technical concept, an embodiment of the present application also provides a computer program, which, when executed by a main control device, is used to implement the above method embodiment.
[0071] The program may be stored in whole or in part on a storage medium packaged with the processor, or may be stored in whole or in part on a memory not packaged with the processor.
[0072] Based on the same technical concept, the embodiment of the present application further provides a processor, which is used to implement the above method embodiment. The above processor can be a chip.
[0073] The various embodiments of the present invention can be combined arbitrarily to achieve different technical effects.
[0074] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in this application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).
[0075] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A handwritten content authentication method based on image recognition, characterized in that: include: S1. Obtaining a coordinate sequence of user handwritten content, sorting the coordinate sequence by timestamp, distinguishing the start and end of a stroke based on pen pressure data, and generating a handwritten image using a rasterization algorithm; wherein the coordinate sequence includes xy coordinates, timestamp, and pen pressure data; S2. Preprocessing the handwritten image to generate a binary image of the handwritten image; S3. Input the binary image into a pre-trained recognition model to obtain the recognition result and confidence level. If the confidence level is higher than a preset threshold, the recognition result is verified with the customer's registered name. If the verification is successful, the business process continues.
2. The handwritten content authentication method based on image recognition according to claim 1, characterized in that: The step S3 further includes: If the confidence level is lower than the preset threshold, the customer is prompted to rewrite the signature; If the number of rewrites exceeds the preset number, the manual review process will be initiated.
3. The handwriting content authentication method based on image recognition according to claim 1, characterized in that: The step S3 further includes: If the verification fails, the customer is prompted to rewrite the signature; If the number of rewrites exceeds the preset number, the manual review process will be initiated.
4. The handwritten content authentication method based on image recognition according to claim 1, characterized in that: The recognition model is a ResNet convolutional neural network model; The ResNet convolutional neural network model includes several cascaded residual blocks, each of which includes a direct path to the input feature map X and a convolution operation path. The convolution operation path includes at least two convolution layers, and a ReLU activation function is set between the two convolution layers. When the dimension of the input feature map X is inconsistent with the output dimension of the convolution operation path, the input feature map is dimensionally matched through a 1×1 convolution layer. A batch normalization layer is set after the convolution operation path of each residual block. The batch normalization layer performs batch normalization on the input data of each layer and standardizes the neuron values by calculating the mean and variance of the neurons in each layer; It also includes a hyperparameter optimization module that uses the Bayesian optimization algorithm to iteratively optimize the number of residual blocks, convolution kernel size, learning rate decay strategy, and batch normalization regularization parameters.
5. The handwritten content authentication method based on image recognition according to claim 1, characterized in that: The step S2 specifically includes: S21, performing trajectory smoothing processing on the coordinate sequence using Kalman filtering; S22, normalizing the pen pressure data through adaptive pressure adjustment; S23, scaling the trajectory to the target size; S24, drawing an anti-aliased curve on a blank canvas to generate a binary image, wherein the line width is determined by multiplying the pen pressure value by 3; S25, performing a morphological closing operation on the binary image, with a 3×3 diamond kernel as the structural element, to connect the disconnected strokes; S26. Add uniformly distributed random noise at the image pixel level, and limit the jitter range to within 3 pixels around the stroke through a mask.
6. The handwriting content authentication method based on image recognition according to claim 1, characterized in that: The step S3 further includes: When the recognition result contains unrecorded characters, upload the character samples to the cloud and call the backup model to return a temporary recognition result; The backup model is a cloud-based auxiliary recognition model used to process unrecorded characters that cannot be recognized by the recognition model. The unrecorded characters include uncommon characters and rare symbols.
7. The handwriting content authentication method based on image recognition according to claim 1, characterized in that: The recognition result is matched with the characters pre-verified by the customer, and a secondary verification is performed based on the topological structure characteristics of the signature trajectory.
8. A handwritten content authentication system based on image recognition, characterized in that: include: A handwriting data acquisition module, which acquires a coordinate sequence of the user's handwriting content, sorts the coordinate sequence by timestamp, distinguishes the start and end of the strokes based on the pen pressure data, and generates a handwritten image using a rasterization algorithm; wherein the coordinate sequence includes xy coordinates, timestamps, and pen pressure data; A preprocessing module, which preprocesses the handwritten image to generate a binary image of the handwritten image; The recognition module inputs the binary image into a pre-trained recognition model to obtain the recognition result and confidence level. If the confidence level is higher than a preset threshold, the recognition result is verified with the customer's registered name. If the verification is successful, the business process continues.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the handwritten content authentication method based on image recognition are implemented as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the handwriting content authentication method based on image recognition as claimed in any one of claims 1 to 7 are implemented.