Identity, behavior and intent authentication triad digital penmanship video signing method and system

CN122369126BActive Publication Date: 2026-09-22CHONGQING AOXIONG INFORMATION TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202610822507.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-22
Estimated Expiration
2046-06-09

AI Technical Summary

Technical Problem

[0005]针对上述问题,本申请提出一种身份、行为、意愿三合一认证的数字笔迹视频签署方法及系统,解决现有技术中签署流程中身份认证、签署行为、意愿表达三个环节时空同步割裂、完整证据链分散依赖于CA等多方分别举证,无法由单一数据包、单一证据载体独立实现闭环验证,以及现有数字证书方案中使用签署流程冗杂,签署人体验繁重的问题

Benefits of technology

[0050]应用本申请的技术方案,在进行文件签署的过程中同时完成签署人身份认证、签署行为以及签署人意愿表达的可鉴定数字笔迹数据采集,保证三者具有时空同步性,并在签署行为完成时,将身份认证结果、签署行为以及意愿表达的可鉴定数字笔迹数据三者统一封装于一个不可分割的数字数据包中,并通过隐写技术嵌入至该签署行为的笔迹签字图片内,使得隐写签名图片一旦生成,即成为承载身份、行为、意愿三重认证信息的唯一、完整、自足的证据实体。当需要验证时,从签署文件中提取隐写签名图片,解析隐写在其中数字笔迹签署数据包并提取笔迹点位序列数据,进而通过数字笔迹鉴定技术精准确认书写者身份与真实意愿的一致性,同时关联比对同步提取的人脸身份认证结果与完整签署行为跟踪日志,使身份认证、签署行为与意愿表达三者在时空维度上形成强关联、可追溯、独立可验证的证据链,从而彻底消除传统电子签名中三者分离、证据分散、无法闭环验证的缺陷,达到在无需调取外部数据库、视频文件、无第三方介入条件下,仅凭单一隐写签名图片(证据载体),其中的单一数据包即可完整还原签署全过程,独立证明签署人身份真实性、签署行为真实性与意思表示自愿性的技术效果,显著提升数字签署的法律效力与司法采信度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369126B_ABST
    Figure CN122369126B_ABST
Patent Text Reader

Abstract

The application discloses a three-in-one digital handwriting video signature method and system for identity, behavior and willingness authentication, belongs to the technical field of image and video recognition, and simultaneously completes data collection of identity authentication, signature behavior and willingness expression of a signer in a file signature process, guarantees that the three have space-time synchronism, and unifies identity authentication results, signature behavior and willingness expression data in an inseparable digital handwriting signature data package when the signature behavior is completed, and embeds the digital handwriting signature data package into a handwriting signature picture through steganography to obtain a steganographic signature picture, the steganographic signature picture becomes a unique, complete and self-sufficient evidence carrier for bearing three authentication information of identity, behavior and willingness, external data and video files do not need to be called, and only by extracting the data package in the picture, a closed loop proof of identity comparison, willingness expression and behavior consistency verification can be independently completed, the difficulty of cross-examination and the cost of identification are significantly reduced, and the signature safety is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of image and video recognition and electronic digital data processing technology, and specifically relates to a digital handwriting video signing technology that integrates identity, behavior, and intent authentication. Background Technology

[0002] As a core component of digital transactions, electronic signatures must simultaneously meet the requirements of identifiable identity, expressible intent, tamper-proof content, and traceable process for legal validity. According to relevant laws, a reliable electronic signature should meet the following conditions: the signature creation data is the exclusive property of the signatory; the signatory has exclusive control over the signature at the time of signing; any alteration to the signature after signing can be detected; and any alteration to the content and form of the data message can be detected.

[0003] In traditional electronic signature schemes centered on digital certificates (CA certificates), the signer must go through multiple steps to complete the signing process, including digital certificate application, real-name authentication, digital certificate issuance, document receipt, identity authentication during the signing process, and secondary confirmation of intent. Because these steps are independent of each other, the overall process is complex and has a high barrier to entry. In particular, core steps such as "identity authentication," "document signing," and "intent verification," although each can be completed independently by the CA, are separated and independent in the signer's workflow, resulting in an isolated and fragmented state in terms of time and space synchronization. Therefore, existing schemes suffer from low efficiency, poor convenience for signers, and security risks such as "authenticating oneself while signing for others" or "signing for others while authenticating one's own intent," leading to identity theft and other problems.

[0004] To address these issues, a dual-recording video solution combining video recording and document signing has emerged in the industry. For example, patent application CN202610025116.3 proposes a video recording-based intention authentication method for electronic contract signing. This method involves verifying identity before the signing process begins and simultaneously collecting biometric video, audio of the intention statement, and information about the signing operation during continuous video recording. The video file and contract document are hashed separately, then fused with digital signatures and task identifiers for blockchain or trusted timestamp storage, achieving a strong data-level correlation between audiovisual evidence and the electronic contract. However, this technical solution still fails to solve the problems of spatiotemporal synchronization and self-verification of evidence at the behavioral level among identity authentication, signing behavior, and intention expression. First, the identity authentication in the comparison documents only involves a static comparison before recording begins, failing to achieve continuous identity monitoring and real-time anomaly blocking throughout the signing process. This cannot eliminate risks such as switching people midway through the signing or having someone else renew the signature after the camera is obstructed. Second, the signing operation is only recorded as "interaction process information," such as clicks and coordinates, without collecting a unique biometric sequence of handwritten handwriting points (including pressure, timing, and pen stroke state). This makes it impossible to achieve objective technical verification of "expression of intent" through forensic handwriting identification. Third, the evidence chain still relies on external notary platforms or third-party storage systems. Although data such as video files, contract hashes, digital signatures, and task identifiers are integrated and stored, they cannot be independently extracted from a single carrier to complete the closed-loop verification of "identity-behavior-intention" triple authentication. Multiple heterogeneous data sources still need to be retrieved for cross-comparison, resulting in high forensic identification costs and complex procedures. Especially during the evidence extraction stage, if the original video is cropped, compressed, or transcoded, key frames may be lost, leading to inaccurate alignment between identity and behavior, thus failing to meet the requirement that "any alteration after signing can be detected" for electronic signatures. Therefore, although the above technologies have made progress in binding audiovisual evidence with contracts, they have not yet been able to build a digital signature system based on biometric behavioral characteristics (digital handwriting) that achieves spatiotemporal synchronization of identity authentication, behavior recording, and expression of intent, and can independently complete judicial appraisal and closed-loop verification of evidence with a single data packet and a single evidence carrier. There is an urgent need for a more systematic, secure, and more consistent solution with the essential requirements of electronic signatures. Summary of the Invention

[0005] To address the aforementioned issues, this application proposes a digital handwriting video signing method and system that integrates identity, behavior, and intention authentication. This solves the problems in existing technologies where the signing process involves spatiotemporal disconnect between identity authentication, signing behavior, and intention expression; the complete evidence chain is scattered and relies on multiple parties, such as CAs, to provide evidence separately; closed-loop verification cannot be achieved independently by a single data packet or a single evidence carrier; and the signing process in existing digital certificate schemes is cumbersome and the signer experience is cumbersome.

[0006] This invention abandons the traditional digital certificate signing process and designs a digital signature system based on digital handwriting signature data packages (containing handwriting point sequence data of biometric behavioral characteristics). This system achieves spatiotemporal synchronization of identity authentication, signing behavior recording, and expression of intent, and can independently complete forensic identification and evidence closed-loop verification with a single data package and a single evidence carrier. First, during the document signing process, it simultaneously completes the signing of the signatory's identity, records the signing behavior, and collects identifiable handwriting point sequence data expressing the signatory's intent, ensuring the spatiotemporal synchronization of these three aspects. Furthermore, after signing, information including, but not limited to, the handwriting point sequence data is hidden in the signature image. This hidden signature image, as the single evidence carrier, is used in the signed document. Therefore, the entire evidence chain is directly linked by the digital handwriting signature data package, achieving a closed-loop evidence process for the signing process with a single data package and a single evidence carrier.

[0007] The technical solution of the present invention is as follows:

[0008] This invention proposes a method for digital handwriting video signing that integrates identity, behavior, and intent authentication. The method includes:

[0009] In response to a document signing instruction, the system displays a visual representation of the document to be signed, along with its hash value, and initiates camera recording to generate a real-time video stream.

[0010] Key video frames are extracted from the real-time video stream, and face detection and liveness detection are performed using the key video frames containing faces to complete the facial identity authentication of the signatory.

[0011] After authentication, the system collects the handwriting position sequence data of the signer's signature. During the signing process, it continuously tracks faces and monitors the presence, number, and identity consistency of faces in the video stream. When no face, multiple faces, or inconsistent identities are detected, an anomaly warning is triggered and the signature submission is blocked. After the anomaly is recovered, the system re-extracts frames from the real-time video stream for face identity authentication until authentication is successful.

[0012] The handwritten signature stroke position sequence data includes the original parameters such as the coordinates (x, y) of each stroke position during the writing process, timestamp t information, pressure value p, and stroke state value s.

[0013] In response to the signature submission command, frame extraction of the real-time video stream is triggered for facial recognition authentication.

[0014] After successful authentication, the handwriting point sequence data, face identity authentication result, face tracking abnormal event log, key video frames obtained by frame extraction, hash value of the document to be signed, trusted timestamp, device fingerprint information, and signing business information are encrypted and encoded to form a digital handwriting signing data package that supports digital handwriting identification.

[0015] The handwriting point sequence data is rendered to generate a visualized handwriting signature image, and the digital handwriting signature data package is steganographically embedded in the handwriting signature image to generate a steganographic signature image carrying complete handwriting data. This steganographic signature image forms a single evidence carrier that combines identity authentication, signing behavior, and expression of intent, which can be used for closed-loop evidence verification.

[0016] The steganographic signature image and the digital handwriting signature data packet are hashed to generate a digital fingerprint and bound with a trusted timestamp to form a certificate of evidence.

[0017] The steganographic signature image is embedded into the document to be signed, generating a signed electronic document.

[0018] The frame extraction process involves selecting key video frames from important nodes in the real-time video stream. These key video frames include, but are not limited to: the first frame at the start of the video, the first frame showing a face, the first frame showing the beginning of a stroke, the last frame showing the end of a stroke, the first frame after recovery from a face tracking anomaly, and the frame where the signer triggers the submission command. The important nodes are nodes where security changes, signing actions, or handwriting changes occur in the business scenario, serving as time anchors for the identity authentication status. The selection of these important nodes must meet the following requirements: covering the signer's signing behavior, covering the writing behavior, and including nodes where identity can be reconfirmed after an anomaly has occurred and recovered.

[0019] The video stream frame extraction, digital handwriting acquisition, and face tracking are all time-synchronized based on the same time base, so that the handwriting point sequence data and the key video frames obtained by frame extraction are synchronized in time and space.

[0020] Furthermore, the facial identity authentication involves comparing the key video frames containing faces obtained from frame extraction with the facial features in the identity information database to complete identity verification; the identity consistency involves comparing the feature similarity between the current face and the face template that has passed facial identity authentication, and if the similarity is greater than a set threshold, the identity is considered consistent.

[0021] Furthermore, time synchronization is achieved using timestamp technology, including: using the video stream start time as the base timestamp T0; recording the relative time Δt1 when extracting each frame of the video stream and calculating the timestamp T1 = T0 + Δt1; recording the relative time Δt2 when acquiring each handwriting point and calculating the timestamp T2 = T0 + Δt2; and achieving spatiotemporal binding of a specific video frame with the corresponding handwriting point at millisecond-level precision by matching the corresponding T1 and T2.

[0022] Furthermore, the encryption involves encrypting the data in two levels: a core level and a related level.

[0023] The core-level data includes: handwriting point sequence data, facial recognition results, facial tracking anomaly event logs, key video frames obtained from frame extraction, and device fingerprint information. The associated-level data includes: signing business information (including signature image generation time and signing task number), hash value of the document to be signed (SHA-256 hash value), and trusted timestamp.

[0024] Furthermore, the encryption also includes auxiliary-level data, forming a three-level data encryption system: core-level, association-level, and auxiliary-level. The auxiliary-level data includes: the fingerprint of the signing device terminal, the signing device model, the data collection SDK version number, and the cloud signing service version number.

[0025] Furthermore, the steganography involves pixel-level steganography of data at each level. Specifically, it is a hierarchical embedding based on a combination of the least significant bit and non-least significant bit in the spatial domain. Core-level data is preferentially embedded in the intermediate significant bit planes of the R and G channels of the signature handwriting area pixels. Correlation-level data is embedded in the least significant bit plane of the B channel of the signature handwriting area pixels, as well as the least significant bit plane of the R channel of the non-handwriting background area. Auxiliary-level data is embedded in the least significant bit planes of the G and B channels of the non-handwriting background area pixels. The embedding position sequence is generated by a linear congruence generator that uses the signature task number as a seed.

[0026] Furthermore, the evidence closed-loop verification involves extracting the steganographic signature image from the signed document, parsing the digital handwriting signature data packet hidden within it, extracting the handwriting point sequence data, confirming the writer's identity and intent consistency through digital handwriting identification, and combining the facial identity authentication results with facial tracking abnormal event logs to form an evidence closed loop.

[0027] Furthermore, the evidence loop verification includes:

[0028] (1) Obtain the steganographic signature image (i.e. the original signature image) from the signed document to prove the business relevance between the signed document and the signature image.

[0029] (2) From the steganographic data of the steganographic signature image, the digital handwriting signature data packet is parsed and the handwriting position sequence data is extracted. Through digital handwriting identification, the identity at the time of signing and the intention at the time of signing are confirmed.

[0030] (3) By analyzing the identity authentication results of video stream frame extraction in the digital handwriting signature data packet, the continuous judgment results of the authentication status of the face tracking module, and the corresponding video stream frame extraction, a dual identity authentication is formed with the handwriting identification results.

[0031] (4) By detecting whether the steganographic signature image supports the extraction of digital handwriting signature data packets, the integrity of the steganographic data packets and whether they have been tampered with can be verified.

[0032] This invention also provides a digital handwriting video signing system that integrates identity, behavior, and intent authentication, and the system includes the following functional modules:

[0033] The camera module is used to simultaneously activate video recording in response to document signing instructions, and to continuously record video during the signing process, generating a real-time video stream.

[0034] The handwriting acquisition module is used to collect handwriting point sequence data when the signer writes a signature. The handwriting point sequence data includes the coordinates (x, y) of each handwriting point during the writing process, timestamp information t, pressure value p, pen touch state value s and other raw parameters.

[0035] The display module is used to display the real-time video stream, the signer's handwritten signature, and the hash value of the document to be signed.

[0036] The video stream frame extraction module is used to extract key video frames from the real-time video stream at important signing nodes. Key video frames include, but are not limited to: the first frame at the start of the video, the first frame of a face, the first frame of the pen stroke, the last frame of the pen stroke, the first frame after the face tracking anomaly is recovered, and the frame when the signer triggers the submission command.

[0037] The identity authentication module compares key video frames containing faces extracted from the real-time video stream with facial features in the identity database to complete identity verification. It performs facial authentication based on the extracted key video frames before signing begins, performs facial authentication based on the extracted key video frames when submitting after signing, and re-executes facial authentication by extracting frames from the real-time video stream after recovering from any abnormal state during signing. This forms a complete closed-loop identity authentication process from before signing, during signing, to signing completion, allowing access to subsequent processes only after successful authentication.

[0038] The face tracking module is used to continuously track faces during the signing process, monitor the presence, number, and identity consistency of faces in the video stream, and trigger an abnormal warning and block the signature submission when no face, multiple faces, or inconsistent identities are detected.

[0039] The data encapsulation module is used to uniformly encrypt and encode the handwriting point sequence data, face identity authentication results, face tracking abnormal event logs, key video frames obtained by frame extraction, hash values ​​of documents to be signed, trusted timestamps, terminal device fingerprint information, and signing business information into a digital handwriting signing data packet.

[0040] The signature image generation and steganography module is used to render and generate a handwriting signature image based on the handwriting point sequence data, and to use steganography technology to embed the digital handwriting signature data packet into the RGB channel pixel data of the handwriting signature image to generate a steganography signature image carrying complete handwriting data, forming a single evidence carrier that combines identity authentication, signing behavior, and expression of intent, for closed-loop evidence verification.

[0041] The evidence storage module is used to perform hash calculations on the steganographic signature image and the digital handwriting signature data packet, generate a digital fingerprint and bind a trusted timestamp to form a signature evidence storage certificate, which is then uploaded to a blockchain or trusted evidence storage platform.

[0042] The document synthesis module is used to embed the steganographic signature image into the document to be signed, thereby generating a signed electronic document.

[0043] The video stream frame extraction module, handwriting acquisition module, and face tracking module are all time-synchronized based on the same terminal system clock, realizing the spatiotemporal anchoring of the data handwriting point sequence with the corresponding key video frames.

[0044] Furthermore, the time synchronization mark is implemented through timestamp technology, including: using the video stream start time as the base timestamp T0; recording the relative time Δt1 when extracting each frame of the video stream, and calculating the timestamp T1 = T0 + Δt1; recording the relative time Δt2 when acquiring each handwriting point, and calculating the timestamp T2 = T0 + Δt2; and achieving spatiotemporal binding of a specific video frame and its corresponding handwriting point with millisecond-level precision by matching the corresponding T1 and T2. Specifically, when a video stream is started for data acquisition, the start time is recorded and its timestamp is calculated as a time reference; when frames are extracted from the video stream for face tracking, the extraction time is recorded, the relative time is calculated, and a timestamp is calculated; when handwriting data points are acquired, the acquisition time is recorded, the relative time with the time reference is calculated, and a timestamp is calculated; using the relative timestamps of the handwriting data and the extracted video stream frames, the extracted video stream frame information is mapped to the handwriting point data, achieving spatiotemporal binding of handwriting points and video frames with millisecond-level precision, ensuring their spatiotemporal synchronization.

[0045] Furthermore, the data encapsulation module encrypts data at two levels: core level and association level. The core level data includes: handwriting point sequence data, face authentication results, face tracking abnormal event logs, key video frames obtained by frame extraction, and device fingerprint information. The association level data includes: signing business information (including signature image generation time and signing task number), hash value of the document to be signed (SHA-256 hash value), and trusted timestamp. The two levels of ciphertext are concatenated into a single bit stream to be embedded according to a preset format.

[0046] Furthermore, the signature image generation and steganography module adopts a hierarchical embedding method based on a combination of least significant bits and non-least significant bits in the spatial domain. Core-level data is preferentially embedded in the intermediate significant bit planes of the R and G channels of the signature handwriting area pixels. Correlation-level data is embedded in the least significant bit plane of the B channel of the signature handwriting area pixels, as well as the least significant bit plane of the R channel of the non-handwriting background area. Auxiliary-level data is embedded in the least significant bit planes of the G and B channels of the non-handwriting background area pixels. The embedding position is determined by a pseudo-random sequence generated by a linear congruent generator (LCG) using the signature task number as a seed.

[0047] Furthermore, it also includes an evidence closed-loop verification module, which is used to extract the steganographic signature image from the signed document, parse the digital handwriting signature data packet hidden in it and extract the handwriting point sequence data, confirm the consistency of the writer's identity and intention through digital handwriting identification, and form an evidence closed loop by combining the facial identity authentication result and the facial tracking abnormal event log.

[0048] Furthermore, the camera module, handwriting capture module, display module, video stream frame extraction module, face tracking module, data encapsulation module, and signature image generation and steganography module are configured in the signing terminal, while the identity authentication module, evidence storage module, and document synthesis module are configured in a cloud server. Alternatively, the camera module, handwriting capture module, display module, data encapsulation module, and signature image generation and steganography module are configured in the signing terminal, while the video stream frame extraction module, face tracking module, identity authentication module, evidence storage module, and document synthesis module are configured in a cloud server.

[0049] Furthermore, the signing terminal is a smartphone, tablet computer, or a smart signing device integrating a front-facing camera and a signing panel; the handwriting acquisition module and the display module are integrated into the signing panel of the terminal device. The signing panel includes an upper video display layer, a middle handwriting acquisition layer, and a lower handwriting generation layer, all of which cover the entire signing and writing area of ​​the signing panel and display the hash value of the document to be signed in the non-writing area. The middle handwriting acquisition layer acquires the handwriting input signal of the signer, collects handwriting point data, and generates handwriting in the lower handwriting generation layer.

[0050] By applying the technical solution of this application, the signing process simultaneously completes the identification of the signer, the signing behavior, and the collection of verifiable digital handwriting data expressing the signer's intentions, ensuring that the three are synchronized in time and space. When the signing behavior is completed, the identification results, the signing behavior, and the verifiable digital handwriting data expressing the intentions are uniformly encapsulated in an indivisible digital data packet, and embedded into the handwriting signature image of the signing behavior through steganography. Once the steganography signature image is generated, it becomes a unique, complete, and self-sufficient evidentiary entity carrying the triple authentication information of identity, behavior, and intentions. When verification is required, the steganographic signature image is extracted from the signed document, the digital handwriting signature data packet hidden within it is parsed, and the handwriting point sequence data is extracted. Then, through digital handwriting identification technology, the consistency between the writer's identity and true intention is accurately confirmed. At the same time, the extracted facial identity authentication results and the complete signing behavior tracking log are correlated and compared, so that identity authentication, signing behavior, and expression of intention form a strongly correlated, traceable, and independently verifiable evidence chain in the spatiotemporal dimension. This completely eliminates the defects of traditional electronic signatures, such as the separation of the three elements, the dispersion of evidence, and the inability to close the loop for verification. It achieves the technical effect that, without the need to access external databases, video files, or third-party intervention, the entire signing process can be completely reconstructed based solely on a single steganographic signature image (evidence carrier) and the single data packet within it. This independently proves the authenticity of the signer's identity, the authenticity of the signing behavior, and the voluntariness of the expression of intent, significantly improving the legal validity and judicial acceptance of digital signatures.

[0051] As can be seen, this solution organically integrates identity authentication, document signing, and intent verification into a single process. It innovatively uses a closed-loop evidence system comprised of handwriting biometrics, identity authentication status, and behavioral trajectories to replace the traditional digital certificate signing process. This not only solves the problem of lengthy processes in existing solutions and significantly improves the signing experience for signatories, but also ensures the spatiotemporal synchronization of identity authentication, document signing, and intent verification, further enhancing signing security. Simultaneously, the method of constructing the evidence chain has changed from requiring multiple parties to extract and splice evidence separately to directly forming a complete evidence chain based on handwriting data packets, significantly reducing the difficulty of cross-examination and the cost of authentication in judicial practice. Furthermore, the video stream frame extraction method used in this solution greatly reduces system consumption, and the facial recognition authentication method forms a complete identity authentication closed loop from pre-signing, during signing, to completion, further significantly improving signing security. Attached Figure Description

[0052] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0053] Figure 1A system architecture diagram of one implementation of the digital handwriting video signing system of the present invention;

[0054] Figure 2 A diagram showing the display status of the signing panel of the aforementioned digital handwriting video signing system;

[0055] Figure 3 A flowchart illustrating one implementation of the digital handwriting video signing method described in this invention. Detailed Implementation

[0056] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0057] It should be noted that the terms "comprising" and "having" and any variations thereof in the specification, claims and accompanying drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such processes, methods, products or devices.

[0058] In traditional digital certificate-based signing processes, signatories need to separately complete two main tasks: applying for a digital certificate and signing based on that certificate. The certificate application process involves first requesting a digital certificate, then undergoing real-name authentication by a Certificate Authority (CA). Only after successful authentication is the digital certificate issued to the applicant. The signing process requires further identity authentication before signing contracts or other documents, followed by a second round of intent authentication. This process involves six completely independent steps before signing can be completed. Clearly, traditional electronic signature schemes, relying on digital certificate systems, separate identity authentication, signing behavior, and intent expression into multiple independent and asynchronous steps. This results in a fragmented signing process in terms of time and space, with the evidence chain scattered among multiple entities such as CAs, platforms, and timestamp service providers. Judicial evidence requires manual piecing together, leading to high costs and low credibility in evidence verification.

[0059] While existing technologies have introduced video recording and task identification binding mechanisms, they still only conduct one-time identity verification before signing. They do not achieve continuous identity monitoring throughout the signing process, nor do they collect handwritten handwriting point sequences with unique biometric characteristics. Therefore, they cannot verify the "expression of intent" through judicial handwriting identification, and the evidence still relies on external system storage. It is impossible for a single carrier to independently complete the closed-loop verification of identity, behavior, and intent.

[0060] The applicant's research revealed that the aforementioned predicament stems from the fact that existing technologies fail to synchronously record and bind the three legal elements—identity authentication, signing, and expression of intent—as a single, inseparable data entity in both time and space. This results in a situation where, while logically they should be integrated, they remain loosely connected and separable data fragments in technical implementation. Particularly in judicial verification, if identity authentication relies solely on a single front-end comparison, signing on click coordinates, and expression of intent on voice, any one of these elements can be forged, replaced, or subsequently denied. Existing solutions cannot provide an evidentiary unit that is solidified upon signing and can prove the shared origin, synchronization, and identity of the three elements solely through the data entity itself.

[0061] Therefore, the applicant proposed a digital handwriting video signing method and system that integrates identity, behavior, and intention authentication. The overall concept is as follows: at the moment the signing is completed, the identity authentication result, signing behavior data, and the verifiable evidence of intention expression, namely the handwriting point sequence data (biometric data), are uniformly encapsulated in an indivisible digital handwriting signing data package. This data is then embedded into the visual presentation of the signing behavior in the handwriting signature image using steganography technology, resulting in a steganographic signature image. Once generated, the steganographic signature image becomes a unique, complete, and self-sufficient evidentiary entity carrying the triple authentication information of identity, behavior, and intention. It does not require access to external databases, video files, or third-party evidence storage platforms. By simply extracting the steganographic data package from the image, a closed-loop proof of identity comparison, intention expression, and behavior consistency verification can be independently completed.

[0062] The following examples, combined with typical electronic signature signing processes, provide a detailed explanation of a digital handwriting video signing solution that integrates identity, behavior, and intent authentication.

[0063] For ease of understanding, the following embodiments illustrate the digital handwriting video signing system and method in combination.

[0064] like Figure 1As shown, the system architecture of the digital handwriting video signing system is as follows: The signing system includes an interaction layer, a service layer, and a data layer. The signing system interacts with the user (signer) through devices such as mobile phones, tablets, or signing boards in the interaction layer. In the service layer, the system includes a camera module, a handwriting acquisition module, a display module, a video stream frame extraction module, an identity authentication module, a face tracking module, a data encapsulation module, a signature image generation and steganography module, a document synthesis module, and an evidence storage module. The data layer stores various data generated by the above modules, such as handwriting sequence point data, identity authentication and face tracking log data, video stream frame extraction image data, signing business data (including the hash of the document to be signed), signature image data, digital handwriting signing data packets, evidence storage result data, device fingerprint information, and trusted timestamps.

[0065] Based on actual business needs, different modules of the service layer are deployed in the signing terminal and the cloud server. In this embodiment, the signing terminal is configured with a camera module, a handwriting capture module, a display module, a video stream frame extraction module, a face tracking module, a data encapsulation module, and a signature image generation and steganography module. The cloud server is configured with an identity authentication module, an evidence storage module, and a document synthesis module. The signing terminal and the cloud server exchange data via a network.

[0066] At the signing terminal, the camera module responds to the document signing command and simultaneously starts recording video, continuously recording during the signing process to generate a real-time video stream. The handwriting capture module collects handwriting point sequence data as the signer writes their signature, including handwriting point coordinates, timestamps, pressure values, and pen stroke status. The display module shows the real-time video stream, the signer's handwritten signature, and the hash value of the document to be signed.

[0067] Depending on the specific application scenario, the signing terminal can be a smartphone, tablet, or a smart signing device integrating a front-facing camera and a signing panel. In this embodiment, a smart signing device integrating a front-facing camera and a signing panel is used. The front-facing camera is the video module, and the handwriting capture module and display module are integrated into the signing panel of the terminal device. When the signer signs, both the signing panel and the front-facing camera are simultaneously activated. The signing panel can display the video stream information captured by the front-facing camera in real time. Furthermore, in terms of visual presentation and writing experience, the video stream image does not affect the handwriting itself. Figure 2 As shown, the signing panel can be designed as follows:

[0068] Panel hierarchy: The upper layer is the video stream display layer, used to display video information from the front-facing camera; the middle layer is the handwriting capture layer, used to collect handwriting data point information; and the lower layer is the handwriting generation layer. The middle handwriting capture layer acquires the handwriting input signal of the signer, collects handwriting point data, and generates handwriting in the lower handwriting generation layer. Existing mature screen interaction event click-through technology can be used. The principle is that the signer writes on the screen with a pen, and the pen tip coordinates penetrate the video stream display layer, reaching the middle capture layer, and finally generating handwriting in real time on the lower Canvas.

[0069] Panel Page Design: The upper video display layer covers the entire signature area of ​​the signing panel. The middle handwriting capture layer and the lower handwriting generation layer also cover the entire signature area. The hash value of the document to be signed is displayed in the non-signature area of ​​the signing panel. The signer must complete the signature without being obstructed by the video content. This design must ensure that the user can observe the information captured by the front-facing camera and the hash value of the document to be signed throughout the signing process.

[0070] The above structure and design will record all information related to signing. For example, when the signer completes writing and submits the signature, the video stream is extracted, that is, the information on the current signing panel is captured. The screenshot includes (1) the face image when the signature is submitted, (2) the signature image showing the signer's writing trajectory, and (3) the hash value of the document to be signed. In this way, the video frame information, the signature data packet echo image, and the business information of the document to be signed can be organically combined, laying the foundation for the spatiotemporal synchronization of the three-in-one scheme of identity authentication, handwriting signing, and intention authentication. Its actual technical implementation can be adjusted according to the needs of real business scenarios to meet the spatiotemporal synchronization of the scheme.

[0071] In the signing system, the handwriting acquisition module is used to collect handwriting point sequence data when the signer writes a signature. This data includes the coordinates (x, y) of each handwriting point during the writing process, timestamp t, pressure value p, and pen touch state value s. This handwriting point sequence data can be used for digital handwriting forensic identification afterward. By confirming the similarity between the handwriting and the handwriting style and habits of the handwriting and the original handwriting sample, it can be determined whether it is the original person's true and unique handwriting habit, thereby completing the expression of intent and identity authentication.

[0072] In this embodiment, the handwriting dot sequence data can be the data sequence of electronic signature handwriting dot information supporting handwriting identification as described in the Chinese patent application filed by the Chongqing Western Handwriting Big Data Research Institute with publication number CN117437699A.

[0073] At the signing terminal, the video stream frame extraction module is used to extract key video frames from the real-time video stream generated by the camera module at critical signing nodes. These image frames can be used for facial recognition authentication. The video stream frame extraction scheme may include, but is not limited to:

[0074] 1) Full-frame video solution

[0075] 2) Frame extraction scheme at fixed time intervals

[0076] 3) Selection based on keyframes in the video

[0077] 4) Frame extraction is performed according to rules based on actual business needs, with specific rules designed specifically for business security requirements.

[0078] In practical applications, it is preferable to determine the key nodes for video stream frame extraction based on business security requirements. Specifically, key nodes are those where security changes occur in the business scenario, signing actions occur, and handwriting changes occur. These serve as time anchors for establishing the identity authentication status. Frame extraction is performed on key nodes in the signing process, and the extracted key video frames include, but are not limited to:

[0079] 1) The first frame of the video start, also known as the first frame of the video, refers to the first image captured when the signing process starts and the camera begins to capture the video stream. It serves as the starting point for initial environment detection and identity authentication, confirming whether the signing environment meets the conditions for legal facial input.

[0080] 2) The first face frame, which is the video frame in which a face appears for the first time in the camera, refers to the frame in the video stream where a valid face image is detected for the first time. It is used to initiate the first face identity authentication process, ensuring that the signer is present and in an authenticable state. It is the first key anchor point in the identity authentication chain.

[0081] 3) The first pen placement frame, i.e. the video frame of the first pen placement point, refers to the video frame extracted at the corresponding time point when the handwriting acquisition module detects the signer's pen placement action on the touch screen. It is used to establish the spatiotemporal binding between "identity verified" and "writing behavior started", confirming that the signing action was initiated by the verified person.

[0082] 4) The last pen lift frame, which is the video frame of the last pen lift in the handwriting, refers to the video frame extracted at the corresponding time point when the handwriting acquisition module detects that the signer has finally lifted the pen to finish writing. It is used to confirm the complete end of the signing behavior and ensure that the identity status is still legal and consistent before the writing ends. It is the endpoint anchor point of the identity continuity in the writing process.

[0083] 5) The first frame after the face tracking anomaly is resolved, that is, the first video frame after the face tracking anomaly is resolved. This refers to the first video image that the system immediately extracts after the face tracking module detects an anomaly (such as face disappearance, multiple people appearing, or identity mismatch) and the abnormal state is eliminated and restored to a single face that matches the initial authentication. This frame is used to re-verify the identity of the signer after the restoration and to prevent the identity from being replaced during the anomaly period.

[0084] 6) The frame that triggers the submission instruction by the signer, i.e. the video frame when the signer clicks the submit button, refers to the current video frame that the system extracts in real time when the signer clicks the "Submit Signature" button. It is used to perform final facial identity authentication at the end of the signing process to ensure that the submission is completed by the authenticated person. It is the final verification node of the identity authentication closed loop.

[0085] In summary, the selection of important nodes by the video stream frame extraction module needs to meet the following requirements: covering the signing behavior of the signer, covering the writing behavior, and including nodes that require reconfirmation of identity when an anomaly occurs and recovery occurs.

[0086] Because this signing system incorporates a video stream frame extraction module, it eliminates the need to store the complete video stream. This avoids the problem in other solutions where dual video recording requires saving the entire video stream for later retrieval and evidence collection. This solution only needs to retain a set of key frames from the video stream that meet the scenario requirements and security requirements to ensure the continuity of identity authentication throughout the process, greatly saving system resources.

[0087] In addition, a face tracking module will continuously track faces throughout the signing process, constantly monitoring the video stream from the front-facing camera to check the presence, quantity, and consistency of faces in the video stream. If a security risk is detected in the video frame: 1) no face is present, 2) multiple faces are present, or 3) a single face is present but it is not the face of the person used for the signing face authentication, the system will trigger an error message and block the handwriting submission. The alert strategy includes, but is not limited to, the following:

[0088] 1) Notify the signer that the current signatory's face detection is abnormal, and provide the specific reason for the abnormality;

[0089] 2) Immediately stop the signer from writing until the error message is cleared;

[0090] 3) It allows writing, but prevents the signatory from submitting the signature to avoid the use of risky handwriting for document signing;

[0091] 4) You can continue writing only after the face detection anomalies have been corrected and the face tracking module has stopped reporting anomalies;

[0092] 5) If an error occurs and the signing fails, the signatory must sign again.

[0093] The above security levels can be adjusted according to business scenarios, security requirements, and actual needs.

[0094] This embodiment adopts the following approach: after triggering an anomaly warning, the user can still continue writing, but the current signature cannot be submitted until the anomaly is resolved and the user's true identity is confirmed.

[0095] The face tracking module can continuously ensure the identity authentication status throughout the entire process through the following four rules:

[0096] 1. Rule 1: The face tracking module will enable face detection throughout the process. When the writing panel shows no face or multiple faces, it will directly prompt the relevant abnormality.

[0097] 2. Rule 2: After the face identity authentication based on video stream frame extraction is successful, the face will be continuously tracked. If the face remains on the screen and no face is moved out (resulting in no face) or multiple faces are detected, the identity authentication status will be maintained until submission.

[0098] 3. Rule 3: When the face removal and multiple face strategies are triggered and then return to normal, a face identity authentication based on video stream frame extraction will need to be performed again to ensure that the face still meets the identity authentication requirements after the recovery.

[0099] 4. Rule 4: When the signer submits their signature after completing the handwriting, facial recognition based on frame extraction from the video stream will be performed again.

[0100] The specific rules and details of the face tracking module are subject to actual business needs and security standards. Under the corresponding business needs and security standards, the face tracking module can ensure the continuity of identity authentication throughout the process.

[0101] The face tracking module can be implemented using existing technologies. For example, face-api.js, a pure front-end face recognition library based on TensorFlow.js, can deploy a lightweight deep learning model to the browser using WebAssembly technology, enabling real-time face detection and analysis with zero dependencies and low latency.

[0102] Throughout the signing process, the signer's facial identity verification is completed through the identity verification module. This module performs facial identity verification at least three times: first, before the signing begins, facial identity verification is performed based on key video frames extracted from the real-time video stream; second, when the signature is completed and the submit button is clicked, facial identity verification is performed based on key video frames extracted from the real-time video stream; and third, after the abnormal state is restored during the signing process, facial identity verification is re-executed by extracting frames from the real-time video stream. This forms a complete closed loop of identity verification from before the signing, during the signing, to the completion of the signing process, allowing entry into subsequent processes only after successful verification.

[0103] Facial identity authentication compares key video frames extracted from a real-time video stream with facial features from an authoritative and reliable identity information database to determine the authenticity and consistency of the identity.

[0104] After completing identity authentication and handwriting collection through the above modules, this signing system uses a data encapsulation module to uniformly encrypt and encode the handwriting point sequence data collected by the handwriting collection module, the face identity authentication result from the identity authentication module, the face tracking abnormal event log from the face tracking module, the key video frames obtained by the video stream frame extraction module, the hash value of the document to be signed, the trusted timestamp, the fingerprint information of the terminal device, and the signing business information into a digital handwriting signing data packet.

[0105] Of course, the encrypted information can include more information as needed, such as: information on the data collection canvas area, details of the signing device, identification information of the signer, handwriting data version information, and geographical location information of the handwriting.

[0106] Simultaneously, based on the handwriting point sequence data collected by the handwriting acquisition module, the system also uses the signature image generation and steganography module to render the handwriting point sequence data into a handwriting signature image. Using layered pixel-level steganography technology, the digital handwriting signature data package is embedded into the RGB channel pixel data of the handwriting signature image to generate a steganographic signature image carrying complete handwriting data. The format can be PNG. This steganographic signature image forms a single evidence carrier that combines identity authentication, signing behavior, and expression of intent, which can be used for closed-loop evidence verification.

[0107] Then, the system uses the evidence storage module to perform hash calculations on the steganographic signature image and the digital handwriting signature data packet, generate a digital fingerprint and bind it to a trusted timestamp to form a signature evidence storage certificate. This evidence storage certificate can then be uploaded to a blockchain or trusted evidence storage platform, such as AntChain or Zhixin Chain, or a traditional timestamp service such as a trusted timestamp center.

[0108] On the other hand, the system's document synthesis module also embeds the steganographic signature image containing the digital handwriting signature data packet into the document to be signed, generating a signed electronic document. The digital handwriting signature data packet contains handwriting position sequence data, including basic position information and identity authentication information. The data packet can be extracted from the signature image of the signed document, thereby completing the one-time construction of a complete evidence chain based on the handwriting data packet.

[0109] The digital handwriting video signing system provided above integrates identity, behavior, and intent authentication. Through a camera module, handwriting acquisition module, and video stream frame extraction module, it achieves millisecond-level spatiotemporal synchronization of handwriting point sequences and key video frames under a unified terminal clock reference, ensuring precise anchoring of writing behavior and identity status in the time dimension. Combined with a face tracking module, it continuously monitors the presence, quantity, and identity consistency of faces throughout the signing process, triggering re-authentication after anomaly recovery, forming a closed-loop identity authentication chain from before signing, during signing, to after submission. A data encapsulation module encrypts and integrates multi-dimensional information such as handwriting data, identity authentication results, tracking logs, keyframes, file hashes, timestamps, and device fingerprints, and a signature image generation and steganography module embeds this information layer by layer into the RGB channels of the handwriting signature image, achieving physical binding of evidence data and visual signature. An evidence storage module synchronously generates hash fingerprints and binds trusted timestamps to the signature image and digital handwriting signature data package, completing tamper-proof on-chain evidence storage. A file synthesis module embeds the steganized signature image into the original file, directly linking the signing result with the business content. Through hardware collaboration and data closed-loop design, the system achieves integrated recording of identity authentication, signing behavior and expression of intent in time and space, avoiding the risk of authentication disconnect caused by the fragmentation of multiple links in traditional solutions. At the same time, it supports the complete extraction and verification of all signing evidence through a single signature image after the fact, significantly improving the integrity and traceability of judicial evidence collection.

[0110] Of course, depending on the needs, the configuration of each functional module in the system can also take other forms. For example, the camera module, handwriting capture module, display module, data encapsulation module, and signature image generation and steganography module can be configured on the signing terminal, while the video stream frame extraction module, face tracking module, identity authentication module, evidence storage module, and document synthesis module can be configured on the cloud server. The signing terminal and the cloud server exchange data via the network.

[0111] The following further embodiments, in conjunction with the above-described signing system, detail the specific implementation process of the entire signing method.

[0112] See Figure 3 The digital handwriting video signing method, which integrates identity, behavior, and intent authentication, includes the following implementation steps:

[0113] I. Start Signing: Initiate the signing task

[0114] The signing system uploads the contract documents to be signed to the cloud server. After successful authentication, a signing task is generated with a signing task number. The signing task specifies the signer's identity information (including name, ID number, and other basic information required for the signing task).

[0115] In response to the document signing instruction initiated by the signing system, the signing terminal pulls up the signing panel, presents the contract document to be signed and its hash value in a visual page on the signing panel, and simultaneously activates the camera module (such as the front-facing camera) to capture video and generate a real-time video stream.

[0116] A comparison with the traditional CA certificate signing process shows that when initiating a signing task, no digital certificate needs to be generated or issued. The signer does not need to apply for a management certificate in advance or carry intermediate media such as a UKey. Moreover, compared with the traditional CA certificate method, the signature data is exclusive to the signer. The signature during signing is controlled by the signer. The expression of identity verification and intent verification is all satisfied by the three-in-one solution, which complies with the requirements of relevant electronic signature regulations.

[0117] II. Facial Recognition

[0118] After the contract signing page is loaded on the signing terminal, the front-facing camera will be automatically activated to obtain a real-time video stream. Frames will be extracted to capture key video frames containing the signatory's facial information. Simultaneously, the facial recognition authentication module will perform face detection and liveness detection on the extracted video stream frames, comparing facial features with features from an authoritative identity database (such as a public security facial database) to complete the initial facial identity authentication. If identity authentication fails, the subsequent signing process will be terminated directly.

[0119] III. Document Signing and Continuous Facial Tracking

[0120] When the signer clicks the signature button to sign the digital handwriting document, the handwriting acquisition module collects the handwriting point sequence data of the signer's handwritten signature on the signing panel. The data includes raw data such as handwriting point coordinates, pressure values, timestamps, and pen stroke states. During the signing process, the face tracking module continuously tracks faces, monitors the presence and number of faces in the video stream, and verifies the consistency of face identities. When no face, multiple faces, or inconsistent identities are detected, an anomaly warning is triggered and handwriting submission is blocked. After the anomaly is recovered, frames are extracted from the real-time video stream for face identity authentication again. Simultaneously, the handwriting point sequence and the key video frames obtained from the frame extraction are time-synchronized based on the same time base to achieve temporal and spatial synchronization.

[0121] Here, verifying the consistency of facial identity involves comparing the feature similarity between the current face and the facial template that passed the initial facial identity authentication. If the similarity is lower than a set threshold, the identity is considered inconsistent.

[0122] In this embodiment, by collecting raw parameters such as the coordinates (x, y), timestamp t, pressure value p, and pen stroke state value s of each handwriting point during the writing process, a complete handwriting point sequence data containing dynamic writing behavior characteristics is constructed. This handwriting point sequence data not only records the trajectory of the signature but also accurately reflects the unique pen pressure change pattern and pen stroke start-stop sequence characteristics of the signer. Combined with the time synchronization marker based on the same time base as the key video frame, a strong correlation between writing behavior and identity authentication in the spatiotemporal dimension is achieved. This supports the comparison of the identity of the signer's behavioral habits and the judgment of the authenticity of intention in digital handwriting identification. It can effectively make up for the problem of insufficient representation of intention caused by relying solely on static trajectory data, and improve the collaborative reliability of identity, behavior, and intention triple authentication and the integrity of the judicial evidence chain.

[0123] IV. Facial Identity Authentication: In response to the signature submission command, the system triggers frame extraction of the real-time video stream again to perform facial identity authentication. Once authentication is successful, the system proceeds to the next step.

[0124] Specifically, one of the core aspects of steps two, three, and four above is that the face tracking module continuously tracks faces, and real-time video stream frame extraction is also used by the face tracking module to perform targeted tracking of the target face information. According to the face tracking module's activation mechanism and rules (specific rules are described in the aforementioned description of the signing system), the verifiability of facial identity is ensured during the signing process. After the initial face authentication is successful, the face authentication module provides the specific face after authentication to the face tracking module. The face tracking module uses facial feature analysis to lock onto and track that specific face after authentication. As long as the tracked object does not move off the screen, it means that the face used for face authentication has not left. This face tracking mechanism solves problems such as authentication departure, mid-process personnel change, and multiple people signing together in actual business: once anomalies are detected during the signing process, such as no face, multiple faces, or the current face not being the same as the initial identity authentication face, the system will trigger a warning, prevent handwriting submission, and force a return to a state of "only the original person." Therefore, the identity authentication status will cover the entire process from the first completion of identity authentication to the generation of the first pen stroke, to the end of the last pen stroke, until the signer finally submits the document. This truly realizes the full-process constraint of "who is signing and who is always signing". Furthermore, it does not adopt a solution of dual recording of complete video stream and handwriting, which saves implementation costs and ensures the spatiotemporal synchronization of identity authentication and handwriting signing. The implementation of spatiotemporal synchronization will be discussed in detail later.

[0125] At this time, the signing data obtained through facial recognition complies with the relevant provisions of the Electronic Signature Law: when electronic signature creation data is used for electronic signature, it belongs exclusively to the electronic signer; when signing, the electronic signature creation data is controlled only by the electronic signer, that is, the electronic signature is dedicated and under exclusive control.

[0126] Since the identity authentication status covers the entire signing process, the correspondence between the extracted frame images from the video stream and the handwriting data points forms the basic data source for the chain of evidence. This proves that the signing behavior is spatiotemporally synchronized with the identity authentication status and the state of intention expression, anchoring the association between the signer's identity and the handwriting data packet.

[0127] The second core of the above steps lies in the spatiotemporal synchronization of identity authentication, document signing, and intent authentication. That is, the spatiotemporal synchronization of handwriting point sequence data, identity authentication and face tracking status, and intent authentication of data packets is achieved through timestamp technology.

[0128] (1) Handwriting dot sequence data

[0129] The signer uses their finger or a stylus to handwrite their signature on the signing panel. The terminal collects raw parameters in real time, including the (x, y) coordinates, timestamp t, pressure value p, and pen touch state value s, for each stroke point during the writing process, forming a sequence of handwriting point data. This sequence completely records the entire dynamic process of the signer's writing behavior. In this handwriting point information, the signer's signing behavior is firmly linked to the time information, forming a one-to-one correspondence between the handwriting point trajectory information and the time point.

[0130] (2) Real-time video stream frame extraction and face tracking

[0131] The front-facing camera video stream frame extraction and handwriting point sequence data are strictly synchronized and marked with the same time base (e.g., the terminal system clock). In the video stream image frame information, whether the facial information of the signer in the current video stream frame meets the verification result of the face tracking module is firmly bound to the time information, forming a correspondence between video stream frame extraction and time point. A precise backtracking relationship is constructed: the state of the facial image captured in each frame has a strict anchoring relationship with the start, movement, and lifting moments of the stroke at the corresponding time point. In the actual evidence presentation process, if someone claims "signature was replaced midway," it is only necessary to retrieve the state information of the face tracking module at the time of the corresponding video stream frame before and after the disputed stroke moment, as well as the video stream frame itself, to clearly determine the true identity of the writer at that moment.

[0132] (3) Further optional, frame extraction of the video stream in the terminal environment can also be performed.

[0133] When the terminal device supports dual cameras, in addition to the front-facing camera recording the signer's facial information, another camera supports video streaming of the terminal environment. Frames of the terminal environment video stream, frame samples of the signer's video, and the corresponding handwriting position sequence information at the same time point are all strictly synchronized and marked using the same time reference, forming a multi-channel spatiotemporal synchronous data acquisition system.

[0134] The correspondence between the above video stream frame extraction and handwriting point sequence data acquisition is guaranteed by timestamp technology, including: using the video stream start time as the base timestamp T0; recording the relative time Δt1 when extracting each frame of the video stream, and calculating the timestamp T1 = T0 + Δt1; recording the relative time Δt2 when acquiring each handwriting point, and calculating the timestamp T2 = T0 + Δt2; by matching the corresponding T1 and T2, spatiotemporal binding between specific video frames and corresponding handwriting points is achieved with millisecond-level precision. Through this millisecond-level precision timestamp matching mechanism, specific handwriting points are precisely aligned with the corresponding key video frames on the timeline, thereby achieving a strong spatiotemporal correlation between the handwriting trajectory of the signing behavior and the facial state of the signer. This effectively solves the technical defects caused by the inability to accurately synchronize handwriting and video frames due to ambiguous time bases or inconsistent clock sources, ensuring that identity authentication, signing behavior, and expression of intent form a verifiable, traceable, and reproducible closed-loop evidence chain within a single data packet, improving the authenticity and legal validity of digital signatures in judicial evidence preservation scenarios.

[0135] Time synchronization markers can be implemented as follows:

[0136] (1) When the video stream is started for data acquisition, record the start time and calculate its timestamp as the time reference;

[0137] (2) When the video stream is frame-stripped due to the need of the face tracking model, record the extraction time, calculate the relative time and timestamp;

[0138] (3) When collecting handwriting data points, record the collection time, calculate the relative time with the time base, and calculate the timestamp;

[0139] (4) Use the relative timestamps of handwriting data and video stream frame extraction to match the video stream frame extraction information with the handwriting data to ensure their spatiotemporal synchronization.

[0140] Therefore, the temporal and spatial synchronization between video stream frame extraction and the signing behavior can clearly present an inseparable chain of evidence that "the same person, at the same time, and on the same device completed the writing of a specific handwriting." If the signatory wants to claim that "it is not their signature" or "the signature was replaced," they need to prove that both the video frame and the handwriting data have been tampered with, which greatly enhances the non-repudiation capability.

[0141] V. Data Encryption:

[0142] The cloud-based signing service compresses and encodes relevant data containing handwriting point sequence data to form encrypted information, resulting in a digital handwriting signing data package that supports digital handwriting identification.

[0143] The relevant data includes the handwriting point sequence data, face identity authentication results, face tracking abnormal event logs, key video frames obtained by frame extraction (i.e., the set of extracted frames of video stream images), hash values ​​of documents to be signed, trusted timestamps, device fingerprint information, and signing business information, etc., which are encrypted and encoded to generate ciphertext information.

[0144] VI. Signature Image Generation and Signature Data Steganography

[0145] The handwriting dot sequence data is rendered to generate a visualized handwriting signature image. Using image steganography, the digital handwriting signature data package is steganographically embedded into the RGB channels of the original handwriting signature image, generating a steganographic signature image carrying complete handwriting data. Specific steps:

[0146] (1) A handwriting signature image is generated based on the handwriting point sequence data. The image is visually presented as a handwritten signature.

[0147] (2) Using a pixel-level embedding method, the encrypted information is written into the specified positions of each pixel layer of the RGB channel of the handwriting signature image according to a certain hidden position sequence, so as to generate a steganographic signature image carrying complete handwriting data.

[0148] In practical applications, since the basic handwriting point information fields, the encoding and encryption methods of the original complete handwriting are usually unknown to the outside world, and the specific pixel-level embedding method used is also unknown to the outside world, the data of the final embedded signature image can independently and completely support handwriting identification and signature fact determination, and evidence extraction and verification can be completed without additional connection to the cloud database.

[0149] The most important data to be steganized is the core data that expresses the biometric characteristics of handwriting, confirms identity, and confirms intent. This core data includes: handwriting point sequence data, facial recognition results, facial tracking anomaly event logs, key video frames obtained through frame extraction, and device fingerprint information. Specifically, its content includes, but is not limited to, basic point information: X and Y axis coordinate point sequences, time sequence information, pressure information, pen stroke state information, writing principle information, writing area and echo area information, signing device information, and necessary additional information, as well as identity verification results from frame extraction of the front-facing camera video stream, continuous judgment results of the facial tracking module's authentication status, and corresponding video frame extraction information. This level of data is the fundamental basis for supporting subsequent digital handwriting forensic identification and has the highest priority protection level.

[0150] Secondly, it supports steganography of association-level information, which includes, but is not limited to: signing business information (signature image generation time, signing task number), hash value of the document to be signed (SHA-256 hash value), trusted timestamp, etc. This level of data strictly links the signing behavior with specific documents and characteristic time information to prove the legal relevance of the signing behavior.

[0151] Finally, it can also support steganography of auxiliary-level data, which includes, but is not limited to, information such as the fingerprint of the signing device terminal, the frame extraction feature value of the video stream of the terminal environment, the data acquisition SDK version number, and the cloud signing service version number. This level of information is used to assist in proving the hardware and software environment at the time of signing, further enhancing the integrity of the evidence chain.

[0152] The image steganalysis sampling pixel-level embedding method in this step, one optional processing flow includes:

[0153] 1) In the previous step, the data for each steganography level is preprocessed separately.

[0154] a. Core level: This level of data has a large volume. In order to ensure the accuracy of identification, lossless compression encoding will be used, and the compressed data will be encrypted independently.

[0155] b. Association level: The data packets at this level are smaller and are encoded using a formatted encoding method with error checking (e.g., CRC cyclic redundancy check + JSON structure) and encrypted independently.

[0156] c. Auxiliary level: Similar to the associated level, it performs formatted encoding and independent encryption.

[0157] In this way, two- or three-level ciphertexts are finally concatenated into a complete bit stream to be embedded according to a preset format.

[0158] 2) Pixel-level layered embedding and position selection

[0159] A hierarchical embedding method based on combining the least significant bit (LSB) and non-least significant bit (NSB) of the spatial domain is adopted to fully utilize the visual redundancy differences of each pixel in the RGB three channels of the signature image to achieve differentiated and robust embedding. Core-level data is preferentially embedded in the intermediate significant bit planes of the R and G channels of the signature handwriting region pixels; correlation-level data is embedded in the LSB plane of the B channel of the signature handwriting region pixels, as well as the LSB plane of the R channel of the non-handwriting background region; auxiliary-level data is embedded in the LSB planes of the G and B channels of the non-handwriting background region pixels; the embedding position sequence is generated by a linear congruence generator using the signature task number as a seed.

[0160] The specific embedding location selection strategy is as follows:

[0161] A. Core-level data: Core-level data is crucial to the validity of evidence and must possess the highest resistance to interference. Therefore, core-level ciphertext is preferentially embedded in the intermediate-significant bit planes (such as the 3rd to 5th bit planes) of the R and G channels of the signature area pixels. The signature area itself has a dark color and complex texture, so modifying the pixel values ​​of the intermediate bit planes is unlikely to cause visual differences, and compared to the least significant bit, it is more robust to common image processing operations such as JPEG compression.

[0162] B. Association-level data: Association-level data is embedded in the least significant bit plane of the B channel of the signature area pixels and the least significant bit plane of the R channel of the non-signature background area. The background area is large, providing sufficient embedding capacity, and the B channel contributes the least to human visual perception, so modifying the least significant bit does not affect the visual effect.

[0163] C. Auxiliary Level Data: The auxiliary layer has the smallest data volume and is embedded in the least significant bit plane of the G and B channels of pixels in the non-handwriting background area as a supplement.

[0164] In terms of embedding order, the core and related layers of main data are embedded first in the handwriting area, and then the remaining related and auxiliary layer data are embedded in the background area. The specific embedding position sequence is generated by a pseudo-random sequence generator bound to the signing task number, ensuring that the bit distribution pattern is different for each embedding, thus increasing concealment and resistance to analysis.

[0165] 3) Pixel embedding, the specific embedding process is as follows:

[0166] A. Bit-plane unpacking: Unpack the R, G, and B channels of each pixel in the handwriting signature image into 8 bit planes. Let I_C_k(x,y) be the k-bit plane of channel C at pixel (x,y), where k=0 is the least significant bit and k=7 is the most significant bit.

[0167] B. Mask generation: Generate a mask for the handwriting area based on the pixel coordinates of the signature.

[0168] C. Initial bit plane selection and host arrangement

[0169] (1) Core-level host arrangement: traverse all pixel positions belonging to signature writing, select bit planes for each position according to the priority of R channel k=3→k=4→k=5→G channel k=3→k=4→k=5, until the total number of bits to be embedded in the core layer is reached.

[0170] (2) Associated host arrangement: After the core host arrangement is completed, continue to select the bit plane by R channel k=1→B channel k=0 in all remaining pixel positions belonging to signature writing until the capacity is sufficient; if it is insufficient, select the bit plane by R channel k=0→B channel k=0 in non-signature writing positions.

[0171] (3) Auxiliary host arrangement: After completing the associated host arrangement, select the bit plane at the optional non-signature writing position by G channel k=0→B channel k=0.

[0172] (4) Bit substitution embedding: The generated bit stream to be embedded is sequentially replaced with the original bit values ​​of the corresponding host bit plane according to the embedding order determined by the pixel embedding logic. During the embedding process, the actual embedded channels, bit planes and coordinate ranges of each layer of data are recorded synchronously to generate header structure information.

[0173] (5) Header information embedding: The header structure information (including the data length of each layer, encryption algorithm identifier, embedding range coordinates, etc.) is written into the specified pixel area of ​​the corner of the handwriting signature image (such as the R channel k=0 bit plane of the upper left corner of 10×10 pixels) in a fixed, least significant bit replacement method that is independent of the task number, so that it can be located and parsed first during extraction.

[0174] (6) Image reconstruction output: After all embedding operations are completed, each plane and channel is re-merged to generate an RGB signature image file carrying complete layered evidence information.

[0175] The steganographic signature image carrying complete handwriting data can be extracted for handwriting identification.

[0176] In summary, by classifying data, a structured, layered encapsulation of evidentiary elements is achieved. This method ensures that core-level data receives priority access to highly robust embedding positions during the steganography process, guaranteeing the integrity and tamper-resistance of handwriting biometrics and identity authentication status. Related-level data, through independent encryption and structured encoding, maintains the legal connection between the signing act and the specific document and time. Thus, without relying on external databases, the data packets steganographically embedded in the signature image possess the independent capability to support the integrity verification of identity recognition, expression of intent, and document binding, improving the efficiency and reliability of data consistency verification during evidence extraction and forensic identification.

[0177] Regarding steganography, besides the adaptive LSB replacement algorithm based on handwriting region masks mentioned above, other algorithms can also be used. For example, adaptive embedding algorithms based on pixel block complexity sorting can be employed.

[0178] VII. Hash Solidification and Notarization: Solidification and notarization of the entire chain of information related to the signing process.

[0179] The digital handwriting signature data packet and the steganographic signature image are hashed together on a cloud signing server, for example, using SHA-256, to generate a unique digital fingerprint. A timestamp is then applied for from a trusted timestamp service provider. The digital fingerprint and timestamp are bound together to form a signature and evidence storage certificate, which is then stored in an evidence storage database or blockchain for end-to-end evidence preservation.

[0180] IX. Compilation of Signed Documents

[0181] The steganographic signature image containing the steganographic signature data packet is embedded into the document to be signed, and the signed document is generated by combining the images.

[0182] Through the above steps, identity authentication, signing behavior, and signing intention can be organically integrated into the same digital handwriting video signing process. The signed documents obtained through this signing method can be used for forensic identification, forming a closed loop of evidence chain.

[0183] During evidence verification, the embedded steganographic signature image is extracted from the signed document, and the digital handwriting signature data packet hidden in it is extracted. The handwriting point sequence data is obtained. Combined with the digital handwriting identification results, the full-process facial recognition results and tracking logs, the consistency verification of the signer's identity and signing intention is achieved, forming a closed-loop evidence system composed of handwriting biometrics, identity authentication status and behavioral trajectory. This closed loop does not need to rely on external systems or splicing data from multiple parties.

[0184] As described above, since the handwriting point sequence data supports handwriting identification, only the digital handwriting signature data package is needed to connect the complete chain of evidence, without the need for a digital certificate signing scheme, which requires multiple parties to present evidence to build the chain of evidence.

[0185] When digital handwriting signatures are required for evidence presentation or post-judicial services, the following process can be used to create a closed-loop evidence system:

[0186] (1) It supports obtaining the original steganographic signature image from the signed document, proving the business relationship between the signed document and the signature image, that is, "the business is handled by the signer of the signed document".

[0187] (2) Digital handwriting signature data packets can be obtained from the steganographic data of the signature image. The handwriting point sequence data in the signature data packets supports digital handwriting forensic identification, which can identify the identity at the time of signing and the authentication of the signatory's intention, that is, "the handwriting at the time of signing the document was written by the person himself in his true writing style, which can confirm the signatory's identity and intention." Digital handwriting identification can determine whether it was written by the same person and whether it expressed the party's intention to sign in a single identification. If the handwriting identification is successful, it means that the handwriting does not hide the signatory's writing habits, and the signatory completed the signing according to his own intention and writing habits. If the handwriting cannot express "willingness to sign", then the handwriting identification will fail.

[0188] (3) The digital handwriting signature data package also includes the identity authentication result of the front camera video stream frame extraction, the continuous judgment result of the authentication status of the face tracking module, and the corresponding video stream frame extraction. It can form a dual identity authentication with the handwriting identification result, that is, "the signer who passes the face identity authentication in the whole process and the signer who is continuously ensured by face tracking is the signer himself, which further confirms the identity of the signer."

[0189] (4) Only the original file data of the signed image that has completed data steganography can support the extraction of the signature data package. Any operation such as copying, taking pictures, using PS, or cutting out images cannot support the correct extraction.

[0190] The above steps only require obtaining a single signed document, and the entire chain of evidence can be linked by the digital handwriting signature data package within it, which greatly reduces the difficulty of cross-examination and the cost of identification in judicial practice of digital handwriting identification.

[0191] The above complies with the specific requirements of relevant regulations that any alteration to the signature after signing, as well as any alteration to the handwriting data package, must be detected.

[0192] This evidence loop achieves a fully verifiable closed loop from front-end acquisition—spatiotemporal binding—layered steganography—hash-based evidence storage—independent extraction—handwriting identification. Using digital handwriting signature data packets as its core, it not only outlines a complete evidence chain but also supports the following four mutually corroborating and independently verifiable evidence chains:

[0193] (1) Identity chain: Core-level handwriting biometric data is bound to core-level identity authentication results. The same data packet in the same image can simultaneously support the identification of "who signed" and the identification of handwriting identity of "whether it is the person who signed", thus locking the true identity of the signer.

[0194] (2) Time chain: The trusted timestamps in the associated data and the handwriting point timestamp t sequence in the core data corroborate each other at the millisecond level, forming a dual time anchor of "the precise moment when the signing behavior occurred" and "the trusted time of a third party", thus eliminating time forgery.

[0195] (3) Behavioral chain: The video stream frame extraction results of core-level data, the face tracking abnormal event log and the handwriting point sequence of core-level data are precisely anchored with the same time benchmark, providing proof of the continuity of behavior of "the same person, at the same time period, completing the writing", so that the defense such as "changing people in the middle" must be overturned at the same time to be valid.

[0196] Scenario Chain: The fingerprint of the auxiliary data device and the frame-sampling feature value of the environmental video record the hardware and physical environment at the time of signing. Combined with the spatiotemporal anchor point information of the front-end acquisition terminal, the completeness of the scenario of "what device and what environment the signing was completed" can be restored.

[0197] As can be seen, by extracting the signature image containing the embedded digital handwriting signature data packet from the signed document, parsing the handwriting point sequence data hidden within, and combining it with the digital handwriting identification results, cross-verification is performed with the facial identity authentication records and facial tracking anomaly event logs throughout the process. This achieves dual confirmation of the signer's identity authenticity and consistency of signing intent, forming a closed-loop evidence system composed of handwriting biometrics, identity authentication status, and behavioral trajectory. This mechanism relies on the spatiotemporal anchoring relationship between the handwriting data packet and video frame extraction under timestamp synchronization, enabling the judicial identification process to independently complete the complete backtracking from handwriting biometrics, identity authentication status to behavioral continuity evidence. It does not rely on third-party systems or scattered data sources; identity confirmation, behavioral tracing, and anti-tampering verification can be completed using only a single signature image. This significantly improves the extractability, integrity, and self-evidentity of electronic signature evidence in judicial scenarios, effectively reducing the complexity of evidence chain integration and the cost of cross-examination. At the same time, through the integrity verification mechanism of the hidden data, it ensures that the signature image has not been copied, tampered with, or intercepted, thereby enhancing the admissibility and non-repudiation capability of electronic signature behavior in legal practice.

[0198] In summary, the digital handwriting video signing scheme disclosed in this invention, which integrates identity, behavior, and intent authentication, uses a closed-loop evidence system composed of handwriting biometrics, identity authentication status, and behavioral trajectory to replace the traditional entire digital certificate signing process. This solves the problem of lengthy processes in existing schemes and greatly improves the signing experience for signatories. It also ensures the spatiotemporal synchronization of identity authentication, document signing, and intent authentication in the signing process, further ensuring its security. Furthermore, the method of constructing the evidence chain has changed from requiring multiple parties to extract and splice evidence chains separately to directly forming a complete evidence chain based on the handwriting data package. This "one data package self-proof" approach greatly reduces the difficulty of cross-examination and the cost of identification in judicial practice for digital handwriting identification.

[0199] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0200] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A digital handwriting video signing method integrating identity, behavior, and intent authentication, characterized in that, In response to a document signing instruction, the system displays the visual content of the document to be signed and its hash value, and starts recording to generate a real-time video stream. It extracts key video frames from the real-time video stream and uses these frames, which contain faces, for face detection and liveness detection to complete the facial identity authentication of the signer. After facial identity authentication, it collects the handwriting point sequence data of the signer's signature, including handwriting point coordinates (x, y), timestamp t, pressure value p, and pen touch state value s. During the signing process, it continuously tracks faces, monitoring the presence, number, and consistency of faces in the video stream. If no face, multiple faces, or inconsistent identities are detected, an anomaly warning is triggered and the signature submission is blocked. After the anomaly is resolved, the system re-extracts frames from the real-time video stream for facial identity authentication. In response to the signature submission command, real-time video stream frame extraction is triggered for facial identity authentication; After facial recognition authentication, the handwriting point sequence data, facial recognition authentication result, facial tracking anomaly event log, key video frames obtained by frame extraction, hash value of the document to be signed, trusted timestamp, device fingerprint information, and signing business information are encrypted and encoded to obtain a digital handwriting signing data packet. A visualized handwriting signature image is generated based on the handwriting point sequence data, and the digital handwriting signing data packet is steganographically embedded into the handwriting signature image to generate a steganographic signature image carrying complete handwriting data. This forms a single evidence carrier integrating identity authentication, signing behavior, and expression of intent. The entire evidence chain is directly linked by the digital handwriting signing data packet, achieving a closed-loop evidence process for signing using a single data packet and a single evidence carrier, providing evidence loop verification. A hash calculation is performed on the steganographic signature image and the digital handwriting signing data packet to generate a digital fingerprint and bind a trusted timestamp, forming a storage certificate. The steganographic signature image is embedded into the document to be signed to generate a signed electronic document. The frame extraction process involves selecting key video frames from important nodes in a real-time video stream, including at least: the first frame at the start of the video, the first frame showing a face, the first frame showing the pen stroke, the last frame showing the pen lift, the first frame after the face tracking anomaly is recovered, and the frame where the signer triggers the submission command. The video stream frame extraction, digital handwriting acquisition, and face tracking are all time-synchronized based on the same time reference, ensuring that the handwriting point sequence data is synchronized with the key video frames obtained from the frame extraction process in time and space. The encryption involves encrypting the data in two levels: a core level and a related level. The core level data includes: handwriting point sequence data, facial recognition results, facial tracking anomaly event logs, key video frames obtained by frame extraction, and device fingerprint information. The related level data includes: signed business information, hash values ​​of documents to be signed, and trusted timestamps. The steganography is pixel-level steganography of data at all levels. Specifically, it is a hierarchical embedding based on the combination of least significant bits and non-least significant bits in the spatial domain. Core-level data is preferentially embedded in the intermediate significant bit planes of the R and G channels of the signature handwriting area pixels, while correlation-level data is embedded in the least significant bit plane of the B channel of the signature handwriting area pixels, as well as the least significant bit plane of the R channel of the non-handwriting background area.

2. The method as described in claim 1, characterized in that, The facial recognition authentication involves comparing key video frames containing faces obtained from frame extraction with facial features in the identity information database to complete identity verification; the identity consistency involves comparing the feature similarity between the current face and the face template that has passed facial recognition authentication, and if the similarity is greater than a set threshold, the identity is considered consistent.

3. The method according to claim 1 or 2, characterized in that, The time synchronization marking is achieved using timestamp technology, including: using the video stream start time as the base timestamp T0; recording the relative time Δt1 when extracting each frame of the video stream and calculating the timestamp T1 = T0 + Δt1; recording the relative time Δt2 when acquiring each handwriting point and calculating the time T2 = T0 + Δt2; and achieving spatiotemporal binding of specific video frames and corresponding handwriting point sequence data with millisecond-level precision by matching the corresponding T1 and T2.

4. The method according to claim 1, characterized in that, It also includes auxiliary-level data, forming a three-level data encryption system of core level, association level and auxiliary level. The auxiliary-level data includes: fingerprint of the signing terminal device, model of the signing terminal device, version number of data collection SDK, and version number of cloud signing service. When steganography, the auxiliary-level data is embedded in the least significant bit plane of the G channel and B channel of the pixels in the non-handwriting background area.

5. The method according to claim 3, characterized in that, The evidence closed-loop verification involves extracting the steganographic signature image from the signed document, parsing the digital handwriting signature data packet hidden within it, extracting the handwriting point sequence data, confirming the writer's identity and intent consistency through digital handwriting identification, and combining the facial identity authentication results with facial tracking abnormal event logs to form an evidence closed loop.

6. A digital handwriting video signing system integrating identity, behavior, and intent authentication, characterized in that, include: The camera module is used to simultaneously start recording video in response to document signing instructions, and to continuously record video during the signing process to generate a real-time video stream; The handwriting acquisition module is used to collect handwriting point sequence data when a signer writes a signature, including handwriting point coordinates (x, y), timestamp t, pressure value p, and pen touch state value s; The display module is used to display the real-time video stream, the signer's handwritten signature, and the hash value of the document to be signed; The video stream frame extraction module is used to extract key video frames from the real-time video stream at important signing nodes. The key video frames include at least: the first frame when the video starts, the first frame when a face appears, the first frame when the pen touches the ground, the last frame when the pen lifts off the ground, the first frame after the face tracking anomaly is recovered, and the frame when the signer triggers the submission command. The identity authentication module compares key video frames containing faces extracted from the real-time video stream with facial features in the identity information database to complete identity verification. Before signing begins, it performs facial identity authentication based on the extracted key video frames. When submitting after signing is completed, it performs facial identity authentication based on the key video frames extracted from the real-time video stream. After recovering from an abnormal state during signing, it re-executes facial identity authentication by extracting frames from the real-time video stream, forming a complete identity authentication closed loop from before signing, during signing to signing completion. It only allows entry into subsequent processes after successful authentication. The face tracking module is used to continuously track faces during the signing process, monitor the presence, number, and identity consistency of faces in the video stream, and trigger an anomaly warning and block the signature submission when no face, multiple faces, or inconsistent identities are detected. The data encapsulation module is used to uniformly encrypt and encode the handwriting point sequence data, face identity authentication results, face tracking abnormal event logs, key video frames obtained by frame extraction, hash values ​​of documents to be signed, trusted timestamps, terminal device fingerprint information, and signing business information into a digital handwriting signing data packet. The signature image generation and steganography module is used to render and generate a handwriting signature image based on the handwriting point sequence data, and to use steganography technology to embed the digital handwriting signature data packet into the RGB channel pixel data of the handwriting signature image to generate a steganography signature image carrying complete handwriting data. This forms a single evidence carrier that combines identity authentication, signing behavior, and expression of intent. The entire evidence chain is directly linked by the digital handwriting signature data packet, realizing a single data packet and a single evidence carrier to complete the signing process evidence loop for evidence loop verification. The evidence storage module is used to perform hash calculations on the steganographic signature image and the digital handwriting signature data packet, generate a digital fingerprint and bind a trusted timestamp to form a signature evidence storage certificate, and upload it to the blockchain or a trusted evidence storage platform. The document synthesis module is used to embed the steganographic signature image into the document to be signed, thereby generating a signed electronic document; Among them, the video stream frame extraction module, handwriting acquisition module and face tracking module are time-synchronized and marked based on the same terminal system clock to realize the spatiotemporal anchoring of handwriting point sequence data and corresponding key video frames; The data encapsulation module performs two-level data encryption: core-level and association-level. The core-level data includes: handwriting point sequence data, face authentication results, face tracking abnormal event logs, key video frames obtained by frame extraction, and device fingerprint information. The association-level data includes: signed business information, hash value of the document to be signed, and trusted timestamp. The two levels of ciphertext are concatenated into a single bit stream to be embedded according to a preset format. The signature image generation and steganography module adopts a hierarchical embedding method based on the combination of least significant bits and non-least significant bits in the spatial domain. Core-level data is preferentially embedded in the intermediate significant bit planes of the R and G channels of the signature handwriting area pixels, while correlation-level data is embedded in the least significant bit plane of the B channel of the signature handwriting area pixels and the least significant bit plane of the R channel of the non-handwriting background area.

7. The system as described in claim 6, characterized in that, The time synchronization marker is implemented using timestamp technology, including: using the video stream start time as the base timestamp T0; recording the relative time Δt1 when extracting each frame of the video stream and calculating the timestamp T1 = T0 + Δt1; recording the relative time Δt2 when acquiring each handwriting point and calculating the timestamp T2 = T0 + Δt2; and achieving spatiotemporal binding of a specific video frame with the corresponding handwriting point at millisecond-level precision by matching the corresponding T1 and T2.

8. The system according to any one of claims 6-7, characterized in that, The camera module, handwriting capture module, display module, video stream frame extraction module, face tracking module, data encapsulation module, and signature image generation and steganography module are configured on the signing terminal, while the identity authentication module, evidence storage module, and document synthesis module are configured on a cloud server; or, the camera module, handwriting capture module, display module, data encapsulation module, and signature image generation and steganography module are configured on the signing terminal, while the video stream frame extraction module, face tracking module, identity authentication module, evidence storage module, and document synthesis module are configured on a cloud server.

9. The system as described in claim 8, characterized in that, The signing terminal is a smartphone, tablet, or a smart signing device that integrates a front-facing camera and a signing panel. The handwriting acquisition module and the display module are integrated into the signing panel of the terminal device. The signing panel includes an upper video display layer, a middle handwriting acquisition layer, and a lower handwriting generation layer, all of which cover the entire signing area of ​​the signing panel. The Hash value of the document to be signed is displayed in the non-writing area. The middle handwriting acquisition layer acquires the handwriting input signal of the signer, collects the handwriting point data, and generates the handwriting in the lower handwriting generation layer.

Citation Information

Patent Citations

  • Electronic handwriting auxiliary identification method, system and device and storage medium

    CN117437699A

  • Method, system and device for generating and decrypting secure electronic file and medium

    CN116108502A

  • Image steganography method

    CN120259066A

  • Non-contact double-recording signature method and device based on terminal intelligence

    CN121768085A

  • Multi-factor identity authentication method, device and equipment

    CN121884466A