Multi-party protocol remote signing system and signing method based on blockchain technology

By combining optical flow analysis and audio analysis with a blockchain-based evidence storage platform, the challenges of identity verification and video authenticity in multi-party remote signing have been solved, enabling efficient remote signing of multi-party agreements and improving the security and transparency of the signing process.

CN121000365BActive Publication Date: 2026-02-17BEIJING POWER LAW INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511109175.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2026-02-17
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

In remote agreement signing scenarios involving multiple parties, existing blockchain electronic signature technology struggles to effectively verify the true identities of the signatories and ensure the authenticity of video signing information, leading to a lack of trust, especially when signing non-standard contracts.

Method used

A video authenticity verification mechanism is introduced. The optical flow features of video frames are extracted by the optical flow analysis module and combined with audio analysis to generate optical flow distribution maps and audio change features. This determines the real-time performance and authenticity of the video information. The blockchain evidence storage platform ensures that the information is tamper-proof.

Benefits of technology

It significantly improves the security and credibility of multi-party remote signing, effectively defends against video forgery, provides reliable assurance of the authenticity of dynamic content, and enhances the transparency and credibility of the signing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121000365B_ABST
    Figure CN121000365B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of blockchains, in particular to a multi-party protocol remote signing system and method based on a blockchain technology, which comprises a blockchain storage platform, a video acquisition module and a video distribution module.The blockchain storage platform is used for collecting and storing uploaded information and ensuring that the information is not tampered with through a time stamp chain structure.The video acquisition module is used for receiving video signing information uploaded by each signing party and sending the video signing information to the blockchain storage platform for storage.The video distribution module is used for downloading video signing information from the blockchain storage platform and presenting the video signing information to a current signing user in real time.The technical solution innovatively introduces a video authenticity verification mechanism.The mechanism can accurately identify whether a video has been tampered with through cutting, splicing and other tampering operations by analyzing the optical flow characteristics of each frame of the video to generate an optical flow distribution diagram and combining the change characteristics of audio information.This effectively guarantees the authenticity and real-time performance of uploaded video information and significantly improves the security and reliability of a multi-party remote signing environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of blockchain technology, and more specifically, to a multi-party remote signing system and signing method based on blockchain technology. Background Technology

[0002] The content in this section provides only background information related to this application and may not constitute prior art.

[0003] Blockchain technology, with its immutable and traceable characteristics, provides a solid foundation of trust for electronic signatures. Its core lies in encrypting and distributively storing key contract information (such as digital fingerprints, signing timestamps, and identity verification records) on the blockchain. This ensures that the contract cannot be unilaterally altered from the moment it is signed, and that the entire process (sending, viewing, and signing) is backed by verifiable timestamps and identity verification. This significantly enhances the security and transparency of the signing process, forms a complete and credible chain of legal evidence, and strengthens the admissibility of electronic contracts in judicial disputes, effectively solving the trust and evidence preservation challenges of traditional electronic signatures.

[0004] Currently, blockchain electronic signature technology is widely used in remote agreement signing scenarios involving multiple parties. However, practice shows that this technology is mainly applied to models between individual users and commercial platforms, where the platform provides standardized contract templates for users to sign. When the signing parties are equal entities and need to enter into non-standardized contracts based on temporarily negotiated content, existing technological solutions are often difficult to apply.

[0005] The key obstacle lies in the reliability of remote identity verification: during remote conference negotiations, it is difficult to effectively verify the true identities of participants; at the same time, the authenticity of the video evidence of the signing process ultimately uploaded to the blockchain cannot be guaranteed, meaning there is a risk of video forgery. This problem is particularly prominent in complex signing scenarios involving multiple collaborating entities, leading to a general lack of trust among all parties in the remote signing method itself and the resulting blockchain-based evidence. Summary of the Invention

[0006] The purpose of this application is to provide a remote signing system for multi-party agreements based on blockchain technology to solve the technical problems raised in the background.

[0007] A remote signing system for multi-party agreements based on blockchain technology, comprising:

[0008] Blockchain-based evidence storage platform: used to collect and store uploaded information, ensuring that the information cannot be tampered with through a timestamp chain structure;

[0009] Video capture module: used to receive video signing information uploaded by each signatory and send it to the blockchain evidence storage platform for storage;

[0010] Video distribution module: used to download video contract information from the blockchain evidence storage platform and present it to the currently contracted user in real time;

[0011] Optical flow analysis module: used to extract the optical flow features of each frame in the video information and generate the corresponding optical flow distribution map;

[0012] Audio analysis module: used to collect audio information corresponding to video information and generate video information change indicators based on audio change characteristics;

[0013] Video verification module: used to determine whether the video information is a real-time, tamper-free video stream based on the video information change index and the optical flow distribution map of each frame.

[0014] This technical solution addresses the issue in multi-party blockchain signing scenarios where dishonest parties may upload false video information through methods such as video splicing and forgery to gain trust and seek illicit gains. It innovatively introduces a video authenticity verification mechanism. This mechanism generates an optical flow distribution map by analyzing the optical flow characteristics of each frame of the video and combines this with the changing characteristics of audio information to accurately identify whether the video has been tampered with through cutting, splicing, or other manipulations. This effectively ensures the authenticity and real-time nature of the uploaded video information, significantly improving the security and credibility of the multi-party remote signing environment.

[0015] The optical flow analysis module includes:

[0016] The video frame segmentation module is used to segment video information into video frames that are related to the time series.

[0017] The optical flow vector analysis module analyzes the optical flow vector information of adjacent video frames to generate optical flow features;

[0018] The optical flow analysis module generates optical flow distribution maps for video frames based on the optical flow distribution maps.

[0019] This solution introduces optical flow analysis (extracting inter-frame motion vectors to generate optical flow distribution maps) and combines it with audio variation analysis to construct a spatiotemporal consistency verification mechanism. This mechanism can accurately identify dynamic tampering traces in videos (such as discontinuous motion and asynchronous lip-sync audio) and verify their real-time nature, thereby effectively defending against advanced forgery methods, significantly improving the authenticity of dynamic content and the overall credibility of the environment in the process of multi-party remote signing, and solving core trust barriers.

[0020] Existing optical flow-based video verification methods lack sufficient detection accuracy for complex dynamic tampering (such as subtle deformations and local motion anomalies in deepfakes). They are unable to effectively capture the correlation between spatial deformation features and global pixel features in video frames, as well as abnormal patterns, resulting in limited recognition rates against advanced forgery techniques.

[0021] Furthermore, the optical flow analysis module includes:

[0022] The optical flow image preprocessor preprocesses the optical flow image to generate the original image;

[0023] The deformation feature extraction network performs pixel convolution and deformation convolution on the original image in sequence, and outputs a deformation feature map.

[0024] The pixel feature extraction network performs deformation convolution and pixel convolution on the original image in sequence, and outputs a pixel feature map.

[0025] The feature fusion network concatenates the deformation feature map and the pixel feature map and then performs pooling to generate a joint feature map. Channel attention is then added to the joint feature map to generate an attention feature map.

[0026] The feature transformation network generates optical flow distribution maps related to dynamic changes in the video based on attention feature maps.

[0027] This solution innovatively designs a dual-branch feature extraction and attention fusion network in the optical flow analysis module: a deformation feature extraction network (pixel convolution + deformation convolution) focuses on capturing local motion deformation features, while a pixel feature extraction network (deformation convolution + pixel convolution) extracts global pixel-related features; the feature fusion network concatenates and pools these two networks, then adaptively weights key information through a channel attention mechanism to generate a highly discriminative attention feature map; finally, the feature transformation network outputs an optical flow distribution map that accurately reflects the dynamic realism of the video. This structure significantly improves the ability to capture subtle deformation artifacts, local motion inconsistencies, and complex spatiotemporal tampering traces, thereby greatly enhancing the system's accuracy and robustness in defending against advanced attacks such as deepfakes, providing a more reliable guarantee of dynamic video authenticity for blockchain remote signing.

[0028] Standard convolution operations primarily rely on convolution kernels with fixed geometric structures for feature extraction. However, this rigid sampling mechanism is prone to losing crucial spatial information when dealing with complex, non-rigid deformations.

[0029] Furthermore, the deformation convolution operation is defined as:

[0030] N represents the number of pixels in the input feature map, n represents the index of the pixel, y(p0) represents the feature value of the output feature map at position p0; x() represents the input feature map, w n p represents the weight of the convolution kernel at position n. n Δp represents the fixed coordinates of the nth position in the standard convolution kernel. n p0 represents the spatial offset and the index of the output feature map.

[0031] Deformation convolution introduces a learnable offset at the sampling position of the standard convolution kernel, enabling the kernel to adaptively focus on task-relevant irregular regions in the input feature map.

[0032] In the technical solution provided in this application, the offset mechanism enables the convolution kernel to dynamically "deform" and actively adapt to changes in the bending of the optical path or the scattering mode (such as the stretching or twisting of the light spot), thereby extracting features related to the macroscopic characteristics of the image change distribution.

[0033] Convolution operations typically involve setting a convolution kernel, traversing the feature map, and then performing dimensionality reduction. This approach struggles to focus on feature extraction at lower scales. Therefore, this application provides the following technical solution:

[0034] The pixel convolution operation is defined as follows:

[0035] y c' (i,j) represents the value of the c'th channel at position (i,j) in the output feature map;

[0036] x c (i,j) represents the value of the c-th channel at position (i,j) in the input feature map;

[0037] w c',c This represents the weights in the convolution kernel that connect the input channel c and the output channel c'.

[0038] c represents the channel index;

[0039] C represents the total number of channels;

[0040] Pixel convolution linearly combines the channel dimensions of the input features.

[0041] In this application, the focus is on the optical flow response of a single pixel, quantifying the influence of local optical flow information and directly relating it to microscopic information.

[0042] The steps for generating attention feature maps include:

[0043] S1: Concatenate the deformation feature map and pixel feature map according to channels to generate a joint feature map;

[0044]

[0045] Among them, X J (i,j,c) is the joint feature map, X def For deformation feature map, X pix For pixel feature maps, i, j, and c represent the feature map spatial height index, feature map spatial width index, and feature map channel index, respectively, and C represents the number of channels in the deformation feature map;

[0046] S2: Perform global average pooling on the joint feature map;

[0047] N represents the size of the joint feature map, z c This represents the feature value of the c-th channel after global average pooling.

[0048] S3: For each channel, add attention weights to generate an attention feature map;

[0049] s c =σ(W2·ReLU(W1·z+b1)+b2) c ;

[0050]

[0051] in, For attention feature maps, s c σ represents the attention weights, ReLU is the linear activation function, W1 is the weight matrix of the first fully connected layer, W2 is the weight matrix of the second fully connected layer, b1 and b2 are the bias vectors of the first and second fully connected layers, respectively, and σ represents the Sigmoid activation function.

[0052] This scheme employs a channel attention mechanism to optimize feature fusion: First, the deformation feature map and pixel feature map channels are concatenated to generate a joint feature map (S1); then, global average pooling is performed to obtain global statistical information for each channel (S2); finally, a small neural network (containing fully connected layers, ReLU, and sigmoid activation) is used to calculate channel attention weights (S3), and the weighted joint feature map is used to generate an attention feature map. This mechanism can adaptively learn and enhance channel features that are crucial for tamper detection (such as channels corresponding to abnormal deformations or motion patterns), while suppressing unimportant channels, significantly improving the discriminative power and robustness of the fused features. This lays a solid foundation for the subsequent accurate generation of optical flow distribution maps that reflect the dynamic realism of the video, effectively improving the detection accuracy of advanced video forgery.

[0053] Further, determining whether the video information is a real-time, untampered video stream includes the following steps:

[0054] Z1: Average pooling is performed on the optical flow distribution map to generate a pooled feature map;

[0055]

[0056] Among them, g c The feature vector is composed of the spatial average values ​​of channel c, where c represents the channel index, i and j represent the spatial location indices, and N represents the size of the attention feature map.

[0057] Z2: Based on pooling feature map gc Generate predicted concentration y k ;

[0058]

[0059] Among them, W cls This represents the classifier weight matrix, where k represents the risk level index, and b... cls y represents the classifier bias term. k This represents the output of the probability of tampering, where K represents the number of risk levels.

[0060] In this application, the average pooling operation on the optical flow distribution map can accurately compress the three-dimensional tensor (height * width * number of channels) into a one-dimensional feature vector, and then accurately compress the attention features into a 2C-dimensional feature vector, thus achieving higher accuracy in generating the predicted concentration.

[0061] Furthermore, obtaining the location of the abnormal interruption includes the following steps:

[0062] Y1: Obtain audio information from the received video stream and generate time-series audio segments by dividing the video stream into frames using the Hanning window.

[0063] Y2: Calculate the short-time energy E of each audio frame. i and zero-crossing rate ZCR i , where i represents the index of the audio segment;

[0064]

[0065] Where n represents the index of the sampling point, N represents the total number of sampling points, and x i Represents the sequence of sample points for an audio frame;

[0066] Y3: Define the silent frame determination condition: when energy E i Less than the preset energy value, zero crossing rate ZCR i If the value is greater than the predicted zero-crossing rate threshold, i is marked as a silent frame.

[0067] This audio analysis module accurately detects short-term abnormal audio interruptions (intervals less than a preset threshold, such as 200ms) and maps them to a high-confidence video tampering risk indicator, effectively addressing the core vulnerability where forgers use minute audio cracks to cover up video editing. Its dynamic threshold adjustment and noise filtering mechanisms can distinguish between ambient silence and malicious tampering. Combined with blockchain timestamps, it achieves frame-level alignment of audio and video, forming a spatiotemporal dual-dimensional tampering evidence linked to optical flow analysis. This significantly enhances the system's real-time defense capabilities against advanced forgery attacks (such as key statement deletion and lip-syncing), providing a verifiable chain of evidence with microscopic anomaly localization capabilities for judicial evidence preservation.

[0068] A method for remote signing of multi-party agreements based on blockchain technology, which uses the aforementioned remote signing system for multi-party agreements based on blockchain technology to sign agreements. Attached Figure Description

[0069] Figure 1 The present invention is a schematic diagram of a multi-party protocol remote signing system based on blockchain technology. Detailed Implementation

[0070] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments. The same reference numerals in the accompanying drawings represent the same components. It should be noted that the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the described embodiments of this application without creative effort are within the scope of protection of this application.

[0071] Compared to the embodiments shown in the accompanying drawings, feasible embodiments within the scope of this application may have fewer components, other components not shown in the drawings, different components, differently arranged components, or components with different connections, etc. Furthermore, two or more components in the drawings may be implemented in a single component, or a single component shown in the drawings may be implemented as multiple separate components.

[0072] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” and similar terms used in this specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not necessarily indicate a quantity limitation. Terms such as “upper” and “lower” are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described object changes.

[0073] refer to Figure 1 This application discloses a multi-party agreement remote signing system based on blockchain technology, including a blockchain evidence storage platform, a video acquisition module, a video distribution module, an optical flow analysis module, an audio analysis module, and a video verification module.

[0074] Blockchain-based evidence storage platform: used to collect and store uploaded information, ensuring that the information cannot be tampered with through a timestamp chain structure.

[0075] The blockchain-based evidence storage platform employs distributed ledger technology (such as a consortium blockchain architecture), deploying multiple consensus nodes to jointly maintain data consistency. The platform receives video streams and metadata (such as timestamps and digital signatures) from the video capture module, generates unique data fingerprints using hash algorithms (such as SHA-256), and writes them sequentially, along with the timestamps, into immutable blocks. Each new block links to the hash values ​​of the previous block through a Merkle tree structure, forming a chain of evidence storage. The platform supports smart contracts to automatically verify the uploader's identity and permissions, ensuring the integrity and traceability of all stored content (including video stream keyframe hashes and audio fingerprints), providing a trusted anchor for judicial evidence storage.

[0076] Video capture module: Used to receive video signing information uploaded by each signatory and send it to the blockchain evidence storage platform for storage.

[0077] The video capture module is integrated into the user terminal (such as a web browser or mobile app), capturing the video and audio streams of the signatory in real time through the device's camera. The module uses the WebRTC protocol for peer-to-peer transmission, performing frame-by-frame compression (e.g., H.264 encoding) and adding digital watermarks (e.g., invisible watermarks based on blockchain IDs) locally on the original video. The processed video stream is then uploaded to the blockchain evidence storage platform via an encrypted channel (e.g., TLS 1.3), along with the user's identity certificate (e.g., X.509) and session timestamp. The module supports breakpoint resumption and integrity verification, ensuring that critical signing actions (such as gesture signatures and verbal confirmations) are stored without omission.

[0078] Video distribution module: Used to download video contract information from the blockchain evidence storage platform and present it to the currently contracted user in real time.

[0079] The video distribution module operates on a subscription-push mechanism: when a party enters the signing process, the module automatically initiates a video stream retrieval request to the blockchain evidence storage platform. After verifying the requester's permissions through a smart contract, it obtains the evidence-storing video stream from the peer node in real time. The module employs low-latency transmission technologies (such as WebTransport or the QUIC protocol) combined with a front-end dynamic buffer pool to achieve millisecond-level video loading. In the user interface, the video stream and the electronic contract are displayed side-by-side (picture-in-picture mode), overlaid with real-time optical flow analysis results to assist users in verifying the authenticity of the other party's behavior. All distributed videos carry blockchain evidence hashes for verification at any time.

[0080] Optical flow analysis module: used to extract the optical flow features of each frame in the video information and generate the corresponding optical flow distribution map.

[0081] The optical flow analysis module includes:

[0082] The video frame segmentation module is used to segment video information into video frames that are related to the time series.

[0083] The video frame segmentation module is implemented based on a time-series parsing engine. Upon receiving the encrypted video stream from the blockchain evidence storage platform, it first performs real-time decoding using a hardware-accelerated decoder. Then, a dynamic frame rate adaptation algorithm is employed (e.g., automatically reducing the frame rate to 15fps to maintain clarity in low-light conditions, and maintaining 30fps under normal conditions). The module binds each video frame to an absolute time anchor point using a high-precision timestamp lock, generating a frame sequence with time-series markers. For critical signing actions (such as the instant a signature is made), the module triggers millisecond-level frame capture (supporting frame interpolation up to 120fps) to ensure no dynamic actions are missed. Abnormal frames (such as corrupted I-frames) are marked and triggered for retransmission by the blockchain evidence storage platform.

[0084] The optical flow vector analysis module analyzes the optical flow vector information of adjacent video frames to generate optical flow features;

[0085] The optical flow analysis module generates optical flow distribution maps for video frames based on the optical flow distribution maps.

[0086] The optical flow analysis module includes:

[0087] The optical flow image preprocessor preprocesses the optical flow image to generate the original image;

[0088] The deformation feature extraction network performs pixel convolution and deformation convolution on the original image in sequence, and outputs a deformation feature map.

[0089] The deformation convolution operation is defined as follows:

[0090] N represents the number of pixels in the input feature map, n represents the index of the pixel, y(p0) represents the feature value of the output feature map at position p0; x() represents the input feature map, w n p represents the weight of the convolution kernel at position n. n Δp represents the fixed coordinates of the nth position in the standard convolution kernel. n p0 represents the spatial offset and the index of the output feature map.

[0091] The pixel feature extraction network performs deformation convolution and pixel convolution on the original image in sequence, and outputs a pixel feature map.

[0092] The pixel convolution operation is defined as follows:

[0093] y c' (i,j) represents the value of the c'th channel at position (i,j) in the output feature map;

[0094] x c (i,j) represents the value of the c-th channel at position (i,j) in the input feature map;

[0095] w c',c This represents the weights in the convolution kernel that connect the input channel c and the output channel c'.

[0096] c represents the channel index;

[0097] C represents the total number of channels;

[0098] Pixel convolution linearly combines the channel dimensions of the input features.

[0099] The feature fusion network concatenates the deformation feature map and the pixel feature map and then performs pooling to generate a joint feature map. Channel attention is then added to the joint feature map to generate an attention feature map.

[0100] The steps involved in creating an attention feature map include:

[0101] S1: Concatenate the deformation feature map and pixel feature map according to channels to generate a joint feature map;

[0102]

[0103] Among them, X J (i,j,c) is the joint feature map, X def For deformation feature map, X pix For pixel feature maps, i, j, and c represent the feature map spatial height index, feature map spatial width index, and feature map channel index, respectively, and C represents the number of channels in the deformation feature map;

[0104] S2: Perform global average pooling on the joint feature map;

[0105] N represents the size of the joint feature map, z c This represents the feature value of the c-th channel after global average pooling.

[0106] S3: For each channel, add attention weights to generate an attention feature map;

[0107] s c =σ(W2·ReLU(W1·z+b1)+b2) c ;

[0108]

[0109] in, For attention feature maps, s c σ represents the attention weights, ReLU is the linear activation function, W1 is the weight matrix of the first fully connected layer, W2 is the weight matrix of the second fully connected layer, b1 and b2 are the bias vectors of the first and second fully connected layers, respectively, and σ represents the Sigmoid activation function.

[0110] The feature transformation network generates optical flow distribution maps related to dynamic changes in the video based on attention feature maps.

[0111] Attention feature map is actually an optical flow distribution map, but the attention feature map has been normalized for format conversion.

[0112] Audio analysis module: used to collect audio information corresponding to video information and generate video information change indicators based on audio change characteristics;

[0113] The change index of video information is the abnormal interruption position of audio information; the abnormal interruption position is the interval where the audio interruption interval is less than a preset threshold.

[0114] Obtaining the location of the abnormal interruption includes the following steps:

[0115] Y1: Obtain audio information from the received video stream and generate time-series audio segments by dividing the video stream into frames using the Hanning window.

[0116] Y2: Calculate the short-time energy E of each audio frame. i and zero-crossing rate ZCR i , where i represents the index of the audio segment;

[0117]

[0118] Where n represents the index of the sampling point, N represents the total number of sampling points, and x i Represents the sequence of sample points for an audio frame;

[0119] Y3: Define the silent frame determination condition: when energy E i Less than the preset energy value, zero crossing rate ZCR i If the value is greater than the predicted zero-crossing rate threshold, i is marked as a silent frame.

[0120] Video verification module: used to determine whether the video information is a real-time, tamper-free video stream based on the video information change index and the optical flow distribution map of each frame.

[0121] Determining whether the video information is a real-time, untampered video stream includes the following steps:

[0122] Z1: Average pooling is performed on the optical flow distribution map to generate a pooled feature map;

[0123]

[0124] Among them, g c The feature vector is composed of the spatial average values ​​of channel c, where c represents the channel index, i and j represent the spatial location indices, and N represents the size of the attention feature map.

[0125] Z2: Based on pooling feature map g c Generate predicted concentration y k ;

[0126]

[0127] Among them, W cls This represents the classifier weight matrix, where k represents the risk level index, and b... cls y represents the classifier bias term. k This represents the output probability of tampering risk, K represents the number of risk levels, and the classifier weight matrix is ​​set according to whether it is a silent frame.

[0128] Example 2: A method for remote signing of multi-party agreements based on blockchain technology, using the aforementioned remote signing system for multi-party agreements based on blockchain technology for agreement signing.

[0129] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A remote signing system for multi-party agreements based on blockchain technology, characterized in that, include: Blockchain-based evidence storage platform: Used to collect and store uploaded information, ensuring the information is immutable through a timestamp chain structure; Video acquisition module: Used to receive video signing information uploaded by each signatory and send it to the blockchain evidence storage platform for storage; Video distribution module: Used to download video contract information from the blockchain evidence storage platform and present it to the currently contracted user in real time; Optical flow analysis module: Used to extract the optical flow features of each frame in the video information and generate the corresponding optical flow distribution map; Audio analysis module: Used to collect audio information corresponding to video information and generate video information change indicators based on audio change characteristics; Video verification module: Used to determine whether the video information is a real-time, tamper-free video stream based on the video information change index and the optical flow distribution map of each frame; The optical flow analysis module includes: The video frame segmentation module is used to segment video information into video frames that are related to the time series. The optical flow vector analysis module analyzes the optical flow vector information of adjacent video frames to generate optical flow features; The optical flow analysis module generates optical flow distribution maps for video frames based on the optical flow distribution maps. The optical flow image preprocessor preprocesses the optical flow image to generate the original image; The deformation feature extraction network performs pixel convolution and deformation convolution on the original image in sequence, and outputs a deformation feature map. The pixel feature extraction network performs deformation convolution and pixel convolution on the original image in sequence, and outputs a pixel feature map. The feature fusion network concatenates the deformation feature map and the pixel feature map and then performs pooling to generate a joint feature map. Channel attention is then added to the joint feature map to generate an attention feature map. A feature transformation network generates optical flow distribution maps related to dynamic changes in videos based on attention feature maps; ; N represents the number of pixels in the input feature map, and n represents the index of the pixel. This represents the feature value of the output feature map at position p0; x() represents the input feature map. This represents the weight of the convolution kernel at position n. Represents the fixed coordinates of the nth position in the standard convolution kernel; p0 represents the index of the output feature map, where p represents the spatial offset. ; Indicates the position of the output feature map The The value of the channel; Indicates the location of the input feature map The value of the c-th channel; This indicates that the convolution kernel connects the input channel c and the output channel c. The weights; c represents the channel index; C represents the total number of channels; Pixel convolution linearly combines the channel dimensions of the input features.

2. The multi-party agreement remote signing system based on blockchain technology according to claim 1, characterized in that, The steps for generating attention feature maps include: S1: Concatenate the deformation feature map and pixel feature map according to channels to generate a joint feature map; ; in, For joint feature maps, This is a deformation feature diagram. For pixel feature maps, i, j, and c represent the feature map spatial height index, feature map spatial width index, and feature map channel index, respectively, and C represents the number of channels in the deformation feature map; S2: Perform global average pooling on the joint feature map; N represents the size of the joint feature map. This represents the feature value of the c-th channel after global average pooling. S3: For each channel, add attention weights to generate an attention feature map; ; ; ; in, For attention feature maps, Here, ReLU represents the attention weights, and ReLU is a linear activation function. This is the weight matrix of the first fully connected layer. Let b1 be the weight matrix of the second fully connected layer, and b2 be the bias vectors of the first and second fully connected layers, respectively. This represents the Sigmoid activation function.

3. The multi-party agreement remote signing system based on blockchain technology according to claim 2, characterized in that, Determining whether the video information is a real-time, untampered video stream includes the following steps: Z1: Average pooling is performed on the optical flow distribution map to generate a pooled feature map; ; Among them, g c The feature vector is composed of the spatial average values ​​of channel c, where c represents the channel index, i and j represent the spatial location indices, and N represents the size of the attention feature map. Z2: Based on pooling feature maps Generate the probability y of predicting tampering risk k ; ; in, This represents the classifier weight matrix, where k represents the risk level index. y represents the classifier bias term. k This represents the output of the probability of tampering, where K represents the number of risk levels.

4. The multi-party agreement remote signing system based on blockchain technology according to claim 1, characterized in that, The change indicator for video information is the location of abnormal interruption in audio information; The abnormal interruption location is the interval where the audio interruption interval is less than a preset threshold.

5. The multi-party agreement remote signing system based on blockchain technology according to claim 4, characterized in that, Obtaining the location of the abnormal interruption includes the following steps: Y1: Obtain audio information from the received video stream and generate time-series audio segments by dividing the video stream into frames using the Hanning window; Y2: Calculate the short-time energy E of each audio frame. i and zero-crossing rate ZCR i , where i represents the index of the audio segment; ; ; Where n represents the index of the sampling point, N represents the total number of sampling points, and x i Represents the sequence of sample points for an audio frame; Y3: Define the silent frame determination condition: when energy E i Less than the preset energy value, zero crossing rate ZCR i If the value is greater than the predicted zero-crossing rate threshold, i is marked as a silent frame.

6. A method for remote signing of multi-party agreements based on blockchain technology, characterized in that, The agreement is signed using the multi-party remote signing system based on blockchain technology as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Block chain-based video evidence storage method, verification method and system

    CN114065255A

  • Video stream dynamic fragment encryption and block chain evidence storage method

    CN120416543A