Video encryption method and system based on dynamic fragmentation and multilayer confusion

Through the video encryption method of dynamic segmentation and multi-layer obfuscation, combined with watermark embedding and distributed key storage, the problem of the existing technology that is unable to track the abuse of authorized users is solved, the security and anti-attack capability of video encryption are enhanced, and it is suitable for the protection of high-resolution video data in rail transit.

CN120711252APending Publication Date: 2025-09-26ZHENGZHOU THINK FREELY HI TECH

Patent Information

Application Number
CN202510809267.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing technologies can only prevent unauthorized access, but cannot track the abuse of authorized users. In addition, the encryption system lacks a collaborative protection mechanism with watermarking technology and is easily cracked and tampered with by attackers.

Method used

A video encryption method based on dynamic segmentation and multi-layer obfuscation is adopted. By identifying sensitive areas in video frames, dynamically dividing the segments, and embedding watermark information in the frequency domain, combined with multi-layer encryption technology and distributed key storage, a watermark self-repair mechanism is set up to achieve key protection of sensitive areas and anti-attack capabilities.

Benefits of technology

It improves the anti-attack capability of video encryption, can track the abuse of authorized users, reduce the probability of key leakage, ensure the security and copyright protection of video content, and is suitable for real-time processing needs of high resolution and large bit streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120711252A_ABST
    Figure CN120711252A_ABST
Patent Text Reader

Abstract

The invention discloses a video encryption method and system based on dynamic fragmentation and multilayer confusion. Identifying all sensitive areas in each frame of original image and calculating a frame-level sensitivity score of the original image; setting a fragmentation mechanism, and dynamically dividing and fragmenting the target video according to the fragmentation mechanism; embedding watermark information into the segmented video frame frequency domain, and converting the video frame frequency domain embedded with the watermark information into a spatial domain; encrypting the video frame embedded with the hidden watermark information by adopting a multi-layer encryption technology; and a watermark self-repairing mechanism is set, and a decryption algorithm is automatically matched according to the terminal type. The encryption calculation amount of a non-sensitive area can be reduced, the anti-attack capability is enhanced, and the key leakage probability is reduced; watermarks are embedded according to characteristic differentiation of different regions, it is ensured that copyright identifiers of key content cannot be deleted, and meanwhile robustness and processing efficiency are balanced; the decryption algorithm supports real-time video processing, an authorized user does not need to manually input a key, and the key management cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data encryption, and in particular to a video encryption method and system based on dynamic sharding and multi-layer obfuscation. Background Art

[0002] Currently, in the rail transit sector, with the widespread use of audio and video data in crew and vehicle operations, such as real-time monitoring, remote dispatch, training and education, and safety inspections, video data contains a large amount of important information involving personal privacy, commercial secrets, and copyrighted content. To effectively protect this sensitive information and prevent unauthorized access, leakage, tampering, and illegal copying and dissemination, strong protection measures are required.

[0003] The invention patent, "A Method and Apparatus for Generating an Encrypted Video File Library," with publication number CN112188308B, discloses a method and apparatus for generating an encrypted video file library, as well as a method and apparatus for distributing and decrypting videos. The method for generating the encrypted video file library includes: splitting a video file into at least one video segment file of a fixed length based on a preset HLS protocol; copying the video segment file to obtain at least two identical target video segment files; assigning a unique key to the target video segment file, and using the key to encrypt the corresponding target video segment file to obtain an encrypted video segment file library. The segment encryption in this technology uses a fixed segment size, making it easily cracked by attackers through pattern analysis. Furthermore, the encryption system can only prevent unauthorized access but cannot track abuse by authorized users. If the encrypted video is cracked, it is difficult to locate the leaker after the video is leaked, and it is impossible to distinguish different transmission paths through technical means. There is a lack of a collaborative protection mechanism with encryption technology. Existing watermarking technology is easily destroyed by denoising and cropping attacks, and affects the video quality after embedding. Watermark embedding may destroy the encrypted fragment structure, resulting in decryption failure or loss of watermark information. Summary of the Invention

[0004] The purpose of the present invention is to provide a video encryption method and system based on dynamic segmentation and multi-layer obfuscation to solve the problem that the existing technology can only prevent unauthorized access but cannot track the abuse of authorized users, and to solve the problem of lack of a collaborative protection mechanism with encryption technology by adding an anti-attack hidden watermark.

[0005] The present invention adopts the following technical solutions:

[0006] The video encryption method based on dynamic segmentation and multi-layer obfuscation includes the following steps:

[0007] S1: Identify all sensitive areas in each frame of the original image in the original video and calculate the frame-level sensitivity score for the original image;

[0008] S2: Calculate the average sensitivity score of the frame images in the unit window, set a fragmentation mechanism based on the average sensitivity score, and dynamically divide the target video into fragments according to the fragmentation mechanism;

[0009] S3: Embed the acquired watermark information into the frequency domain of the fragmented video frame, then convert the frequency domain of the video frame with embedded watermark information into the spatial domain and perform watermark hiding processing;

[0010] S4: Encrypt the video frames embedded with hidden watermark information using a multi-layer encryption technology including slice content encryption, slice structure encryption, and distributed key storage;

[0011] S5: Set up a watermark self-repair mechanism and automatically match the decryption algorithm according to the terminal type.

[0012] Furthermore, the average sensitivity score of the s-frame images in the unit window is calculated. The first-level sensitivity condition is that the average sensitivity score is greater than the second scoring threshold; the second-level sensitivity condition is that the average sensitivity score is less than or equal to the second scoring threshold and greater than or equal to the first scoring threshold; if the window states of the current window and the next window are different, it is considered that a sensitivity jump has occurred; if the average sensitivity score does not reach the first scoring threshold, it is a normal state, and the first-level candidate state is started when the first-level sensitivity condition is met, and the second-level candidate state is started when the second-level sensitivity condition is met.

[0013] Furthermore, the sharding mechanism is as follows:

[0014] The window starts in normal state;

[0015] When the secondary candidate state is started, the time counter is reset;

[0016] When the first-level candidate state is started, the time counter is reset;

[0017] If the time in the secondary candidate state is greater than the preset minimum fragmentation duration, the secondary candidate state is cancelled. When a sensitivity jump occurs or the time is equal to the preset maximum fragmentation duration, a secondary sensitive fragment is generated;

[0018] If the time in the first-level candidate state is greater than the preset minimum fragmentation duration, the first-level candidate state is cancelled. When a sensitivity jump occurs or the time is equal to the preset maximum fragmentation duration, a first-level sensitive fragment is generated;

[0019] The adjacent windows of the first-level candidate state without generated fragments, the second-level candidate state without generated fragments, and the normal state are merged into a normal fragment with a total duration less than or equal to the preset maximum fragment duration.

[0020] Furthermore, step S3 includes:

[0021] S301: Obtain a frequency domain coefficient matrix F(u,v) through discrete cosine transform;

[0022] S302: Adjust the watermark strength based on the fragment sensitivity and embed the watermark information into the frequency domain of the fragmented video frame;

[0023] S303: Convert the video frame embedded with the watermark from the frequency domain to the spatial domain, and hide the watermark by setting a visual perception threshold ∈.

[0024] Furthermore, step S302 adjusts the embedding mode according to the fragment sensitivity:

[0025] When the video frame segment is a first-level sensitive segment, a modified frequency domain coefficient matrix is ​​obtained by superimposing the watermark signal on F(u,v) according to the first intensity gain, the weight matrix related to the frequency domain position, and the watermark signal;

[0026] When the video frame is segmented into secondary sensitive segmentation and common segmentation, the modified frequency domain coefficient matrix after quantizing the watermark signal based on F(u,v) is obtained according to the quantization step size, the second intensity gain, the weight matrix set by the frequency domain energy distribution and the watermark signal.

[0027] Furthermore, step S4 includes:

[0028] S401: Generate a random key stream based on Logistic mapping to perform chaotic encryption on the content of the sliced ​​video frames embedded with watermark information;

[0029] S402: Hash the metadata of the shards, including the location and timestamp.

[0030] S403: Set up a dynamic key management mechanism and store the key fragments in multiple nodes through the distributed key storage of the blockchain.

[0031] Furthermore, the S403 triggers dynamic key rotation according to the playback environment; the playback environment can be an IP address or a device ID; the dynamic key management mechanism is as follows:

[0032] S403a: Obtain key information about the playback environment, obtain the current IP address of the device through the network protocol stack, and read the device ID from the device system;

[0033] S403b: When a change in the IP address or device ID is detected, the system triggers a key update;

[0034] S403c: Obtain a new encryption key, including recalculating the video content hash and device fingerprint at the current playback position.

[0035] Furthermore, the watermark self-repair mechanism is as follows: the terminal detects the integrity of the watermark after receiving it, and after the watermark is embedded in the video,

[0036] When the watermark information content is completely consistent with the originally generated watermark information, it indicates that it has not been tampered with or replaced;

[0037] When the watermark information content is not completely consistent with the originally generated watermark information, indicating that it has been tampered with, the watermark self-repair is achieved by reversely restoring part of the watermark information through the hash value.

[0038] The video encryption system based on dynamic segmentation and multi-layer obfuscation includes: an identification and rating module, a dynamic segmentation module, a watermark embedding unit, a multi-layer encryption unit and a decryption adaptation module;

[0039] The recognition and rating module is used to identify all sensitive areas in each frame of the original image and perform an overall frame-level sensitivity rating.

[0040] Dynamic fragmentation module: used to dynamically divide the target video into fragments according to the fragmentation mechanism. The fragments after pollen include primary sensitive fragments, secondary sensitive fragments and common fragments;

[0041] Watermark embedding unit: used to embed the acquired watermark information into the frequency domain of the video frame after segmentation, and then convert the frequency domain of the video frame with embedded watermark information into the spatial domain, and realize the hiding of watermark by setting the visual perception threshold ∈;

[0042] Multi-layer encryption unit: used to encrypt video frames embedded with hidden watermark information using multi-layer encryption technology including slice content encryption, slice structure encryption and distributed key storage, and then slice and store the key required for encrypting the video; the multi-layer encryption unit includes a slice content encryption module, a slice structure encryption module and a key slice storage module;

[0043] Decryption adaptation module: used to set the watermark self-repair mechanism during decryption, and automatically matches the decryption algorithm according to the terminal type. The watermark self-repair mechanism is set to implement watermark self-repair after detecting incomplete watermark information during decryption.

[0044] Furthermore, the watermark embedding unit includes a frequency domain conversion module, a watermark embedding module and a space domain conversion module; the multi-layer encryption unit includes a slice content encryption module, a slice structure encryption module and a key slice storage module; wherein,

[0045] Frequency domain conversion module: used to convert the original image from the spatial domain to the frequency domain through discrete cosine transform to obtain the frequency domain coefficient matrix F(u,v);

[0046] Watermark embedding module: adjusts the watermark strength based on the fragment sensitivity and is used to embed the watermark information into the frequency domain of the fragmented video frame;

[0047] Spatial domain conversion module: converts the video frame embedded with watermark from frequency domain to spatial domain, and hides the watermark by setting the visual perception threshold ∈;

[0048] Sliced ​​content encryption module: Chaotic encryption is the first layer of multi-layer obfuscation encryption technology. It generates a random key stream based on Logistic mapping and performs chaotic encryption on the sliced ​​video frames embedded with hidden watermark information.

[0049] Segment structure encryption module: used to hash and obfuscate the metadata of the segments. The metadata includes the location and timestamp representing the structural information. The location refers to the position of the segment in the video, and the timestamp refers to the time when the segment was generated or created.

[0050] Key sharding storage module: Set up a dynamic key management mechanism and store key shards in multiple nodes through the distributed key storage of the blockchain.

[0051] The present invention can reduce the encryption calculation amount in non-sensitive areas through dynamic sharding, and achieve precise optimization of encryption calculation amount by focusing on protection of sensitive areas and lightweight processing of non-sensitive areas, thereby improving processing speed and encryption efficiency.

[0052] Watermarks are embedded differently according to the characteristics of different regions. The watermark embedding strength is increased for the first-level sensitive fragments to ensure that the copyright identification of key content cannot be deleted; the embedding strength is reduced for the second-level sensitive fragments and ordinary fragments to reduce the amount of calculation and balance robustness and processing efficiency.

[0053] By superimposing chaotic encryption on video frame fragments, the cost of video cracking is increased and the ability to resist attacks is enhanced. The metadata of video frame fragments is encrypted through hash obfuscation technology to effectively resist fragment reassembly attacks. The risk of centralized servers being attacked is reduced through distributed key storage, and the probability of key leakage is reduced. After the watermark is embedded, the video is encrypted, forming a dual security mechanism of watermark identification and video encryption.

[0054] The parallel encryption algorithm supports real-time video processing, improves computing efficiency, optimizes resource utilization, reduces latency, and enhances system scalability. It is particularly suitable for real-time processing needs of high resolution, high frame rate, and large bit rate. Authorized users do not need to manually enter keys, reducing key management costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 Flowchart of the video encryption method based on dynamic segmentation and multi-layer obfuscation provided by the present invention;

[0056] Figure 2 A schematic diagram of the functional modules of the video encryption system based on dynamic sharding and multi-layer obfuscation provided by the present invention;

[0057] Figure 3 A schematic diagram of the detailed functional modules of the watermark embedding unit in the video encryption system based on dynamic sharding and multi-layer obfuscation provided by the present invention;

[0058] Figure 4 This is a schematic diagram of the detailed functional modules of the multi-layer encryption unit in the video encryption system based on dynamic segmentation and multi-layer obfuscation provided by the present invention. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of the embodiments of this document clearer, the technical solutions in the embodiments of this document will be clearly and completely described below in conjunction with the drawings in the embodiments of this document. Obviously, the described embodiments are part of the embodiments of this document, not all of the embodiments. Based on the embodiments of this document, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this document. It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of this document can be combined with each other in any way.

[0060] The present invention is described in detail below with reference to the accompanying drawings and embodiments:

[0061] As attached Figure 1 As shown, an exemplary embodiment of the present invention provides a video encryption method based on dynamic segmentation and multi-layer obfuscation, the steps are as follows:

[0062] S1: The original video includes multiple frames of original images, identifying all sensitive areas in each frame of the original image and calculating a frame-level sensitivity score for the original image;

[0063] According to an exemplary embodiment, the sensitive area containing the sensitive target is identified by the YOLOv7 target detection model and normalized to [0,1], and the total number of sensitive areas B and the ratio of the number of pixels in the sensitive area to the total number of pixels in the frame image are obtained. i , i∈(0,B), the OCR algorithm is used to identify the content of the license plate and text area, and the corresponding sensitive target weight is matched to each sensitive target; among them, the sensitive area is the image area containing sensitive targets, and sensitive targets include faces, license plates and text areas; frame-level sensitivity score Score f The formula is as follows:

[0064]

[0065] Among them, w i Represents the weight of sensitive targets, w i ∈[0,1], and Area i Indicates the ratio of the number of pixels in the sensitive area to the total number of pixels in the original image of the frame; Conf i Indicates the confidence level of sensitive area detection;

[0066] S2: Calculate the average sensitivity score of the frame images in the unit window, set a fragmentation mechanism based on the average sensitivity score, and dynamically divide the target video into fragments according to the fragmentation mechanism;

[0067] According to an exemplary embodiment, the target video segments are divided into segments according to their sensitivity based on the average sensitivity score, and are divided into first-level sensitive segments with high sensitivity, second-level sensitive segments with low sensitivity, and ordinary segments that can be regarded as not containing sensitive targets. Segments refer to media stream blocks arranged in chronological order.

[0068] Using a unit window of length s frames, calculate the average sensitivity score of the frame image within the unit window The first-level sensitive conditions and the second-level sensitive conditions are set according to the average sensitivity score and the score threshold.

[0069] According to an exemplary embodiment, Q1 is the first scoring threshold, Q2 is the second scoring threshold, Q2>Q1, and the first-level sensitive condition is The secondary sensitive condition is

[0070] Set the window status classification: Normal state, the window average sensitivity score does not reach the first score threshold, that is In the first-level candidate state, the average sensitivity score of the window meets the first-level sensitivity conditions; in the second-level candidate state, the average sensitivity score of the window meets the second-level sensitivity conditions; if the window states of the current window and the next window are different, it is considered that a sensitivity jump has occurred.

[0071] The sharding mechanism is as follows:

[0072] The window starts in normal state, that is,

[0073] when When , the window starts the secondary candidate state and resets the time counter;

[0074] when When , the window starts the first-level candidate state and resets the time counter;

[0075] If the time in the secondary candidate state is greater than the preset minimum fragmentation duration, the secondary candidate state is cancelled. When a sensitivity jump occurs or the time is equal to the preset maximum fragmentation duration, a secondary sensitive fragment is generated;

[0076] If the time in the first-level candidate state is greater than the preset minimum fragmentation duration, the first-level candidate state is cancelled. When a sensitivity jump occurs or the time is equal to the preset maximum fragmentation duration, a first-level sensitive fragment is generated;

[0077] The adjacent windows of the first-level candidate state without generated fragments, the second-level candidate state without generated fragments, and the normal state are merged into a normal fragment with a total duration less than or equal to the preset maximum fragment duration.

[0078] According to an exemplary embodiment, the unit window length is set to 15 frames and the time is 0.5 seconds; the minimum segment length is preset to 2 seconds, and the maximum segment length is preset to 7 seconds. If a 2-second normal state and a 3-second first-level candidate state are adjacent, the 2-second normal state and the 3-second first-level candidate state are merged to ensure that the segment contains the complete sensitive event rather than fragmented instantaneous sensitive images;

[0079] S3: Embed the acquired watermark information into the frequency domain of the fragmented video frame, then convert the frequency domain of the video frame with embedded watermark information into the spatial domain and perform watermark hiding processing;

[0080] According to an exemplary embodiment, the watermark information includes an authorized user ID and a timestamp; the authorized user ID and timestamp are converted into a binary sequence through UTF-8 encoding, the binary sequence is merged and a check code is added to perform CRC check, and the binary stream is spread spectrum using a pseudo-random sequence Gold code to enhance anti-attack capabilities.

[0081] According to an exemplary embodiment, a video file is decoded to obtain video source data in RGB format, and the RGB format is converted into YUV format, or the video source data in YUV format is directly obtained. The Y component is the luminance component, indicating the brightness of the image, and the U and V components are the chrominance components, indicating the blue and red chrominance of the image. The video source data is read frame by frame, and the video frame is divided into several 8x8 pixel blocks, each pixel block corresponding to the luminance component of the original image. Since the human eye has a low sensitivity to brightness, a watermark is embedded in the Y component of the luminance channel, and the Y component of each frame is digitally watermarked, while the U and V components are not processed.

[0082] S301: Obtain a frequency domain coefficient matrix F(u,v) through discrete cosine transform (DCT);

[0083] Perform two-dimensional DCT transformation on each pixel in each 8x8 pixel block in turn to obtain the frequency domain coefficient matrix F(u,v):

[0084] F(u,v)=DCT(block(x,y)) (2)

[0085] Among them, block(x,y) represents the pixel point with coordinates (x,y) in the pixel block block, x represents the horizontal coordinate of the pixel from left to right in the block, x∈[0,7], y represents the vertical coordinate of the pixel from top to bottom in the block, y∈[0,7]; F(u,v) represents the frequency domain coefficient matrix obtained by two-dimensional DCT transformation; u represents the horizontal frequency component from low frequency to high frequency in the frequency domain, u∈[0,7]; v represents the vertical frequency component from low frequency to high frequency in the frequency domain, v∈[0,7]; DCT() represents the two-dimensional DCT transformation function.

[0086] S302: Adjust the watermark strength based on the fragment sensitivity and embed the watermark information into the frequency domain of the fragmented video frame;

[0087] By quantizing or superimposing the watermark signal to adjust the value of the specific F(u,v) and embedding the watermark signal W(u,v), the watermark information can be embedded without significantly affecting the visual quality, ensuring a balance between visual quality and robustness.

[0088] The DCT coefficient position and frequency domain embedding position are selected according to the slice sensitivity: when the video frame slice is a first-level sensitive slice, the watermark signal is embedded in the low-frequency area; when the video frame slice is a second-level sensitive slice and a normal slice, the watermark signal is embedded in the high-frequency area.

[0089] According to an exemplary embodiment, the watermark signal is embedded in the medium and low frequency region to avoid high frequency noise interference and low frequency visual sensitivity, and the watermark signal is embedded in the medium and high frequency region to improve the robustness of the image, the low frequency region is u and v both belong to [2,4); the high frequency region is u and v both belong to [4,6].

[0090] The first intensity (low intensity) watermark is used for the first-level sensitive fragments to avoid image quality loss, and the second intensity (high intensity) watermark is used for the second-level sensitive fragments and ordinary fragments to improve robustness; the watermark signal W(u,v) is binarized to generate a binary watermark sequence {w k}, convert the original watermark into a discrete sequence and adapt the discrete characteristics of quantization embedding to overcome the disadvantage that directly embedding the original watermark signal will cause the frequency domain coefficients to be modified too much and cause serious distortion.

[0091] Adjust embedding method based on sharding sensitivity:

[0092] When the video frame segment is a first-level sensitive segment, the modified frequency domain coefficient matrix F′(u,v) after superimposing the watermark signal on F(u,v) is obtained by the following formula:

[0093] F′(u,v)=F(u,v)+α·w k ·M(u,v) (3)

[0094] Among them, α represents the first intensity gain, which is used to control the invisibility of the watermark after embedding. M(u,v) represents the weight matrix set according to the frequency domain position. The larger the value, the higher the tolerance of the frequency domain coefficient to watermark embedding.

[0095] When the video frame is divided into two segments, namely, secondary sensitive segments and common segments, the modified frequency domain coefficient matrix F′(u,v) after quantizing the watermark signal based on F(u,v) is obtained by the following formula:

[0096]

[0097] Among them, round() represents the rounding function, which rounds the number to a given number of digits; η represents the quantization step size. The larger the step size, the stronger the robustness but the greater the risk of distortion; β represents the second intensity gain; Q(u,v) represents the weight matrix set according to the frequency domain energy distribution;

[0098] S303: Convert the video frame embedded with the watermark from the frequency domain to the spatial domain, and hide the watermark by setting a visual perception threshold ∈.

[0099] Perform inverse DCT (IDCT) on the modified frequency domain coefficient matrix F′(u,v) to convert the modified frequency domain coefficient matrix of DCT coefficients after embedding watermark back to spatial domain slices:

[0100] block′(x,y)=IDCT(F′(u,v)) (5)

[0101] Wherein, block′(x,y) represents an 8×8 pixel block in the spatial domain of the video frame after the watermark is embedded; IDCT() represents the two-dimensional inverse DCT transform function.

[0102] The modified frequency domain matrix is ​​converted back to the spatial domain slice, so that the watermark information is finally integrated into the pixel values ​​of the image or video, completing the complete technical closed loop of watermark embedding;

[0103] Set the visual perception threshold ∈, which refers to the minimum value of the image change that the human eye can perceive. Control the modification amplitude in the video frame to not exceed the visual perception threshold to ensure that the human eye cannot perceive the image quality change; limit the pixel value modification amplitude according to the fragment sensitivity,

[0104] ∣block′(x,y)-block(x,y)∣≤∈ (6)

[0105] According to an exemplary embodiment, the visual perception threshold ∈1=3 for the first-level sensitive slices, and the visual perception threshold ∈2=8 for the second-level sensitive slices and ordinary slices; the allowed modification range is larger, the first-level sensitive slices are flat or low-texture areas such as facial skin, because there is a lack of visual masking effect in the spatial domain, the human eye is sensitive to slight distortion in flat areas, so large modifications are allowed to be smaller, otherwise the visual consistency will be destroyed, and the second-level sensitive slices and ordinary slices are generally high-texture or complex detail areas with architectural textures or noise backgrounds, and the distortion can be masked by the visual masking effect (Masking Effect), so large modifications are allowed, and the robustness of the video slices can be improved at the same time; in addition, if the RGB format frame is converted to YUV format before embedding, the YUV needs to be converted back to RGB format after reconstruction to ensure that the video frame can be displayed directly.

[0106] S4: Encrypt the video frames embedded with hidden watermark information using a multi-layer encryption technology including slice content encryption, slice structure encryption, and distributed key storage;

[0107] The encryption priority is that the first-level sensitive shards take precedence over the second-level sensitive shards and over the ordinary shards; through multi-dimensional and multi-level security protection, multi-layer obfuscation encryption superimposes heterogeneous encryption mechanisms, so that the output of each layer of encryption becomes the input of the next layer, forming a "chain protection". Each layer of technology solves different security problems, avoiding the complete collapse of the disk due to the cracking of a single encryption method.

[0108] S401: Generate a random key stream based on Logistic mapping to perform chaotic encryption on the content of the sliced ​​video frames embedded with watermark information;

[0109] Set the initial value a0 and the control parameter μ; determine the key stream length L = m*N, where m is the number of video frames and N is the number of pixels in a single frame.

[0110] Iterative Logistic mapping generates a chaotic sequence {a1, a2, ..., a L}

[0111] a n+1 =μ·a n ·(1-a n ) (7)

[0112] Among them, μ is the chaos control parameter, μ∈[0,1], which is used to determine the randomness and complexity of the key stream; a n is the current chaos state value, a n ∈[0,1],a n+1 The next chaos state value.

[0113] The chaotic sequence is normalized and converted into a binary bit stream {k1, k2, ..., k j …,kL}, j∈[1,L]; Finally, we get the random key stream K={k1,k2,…,k L}.

[0114] According to an exemplary embodiment, the generated key stream is XORed with the fragmented video frame data to implement chaotic encryption of the video frame, so that the encrypted video frame data presents disordered and random characteristics, and an encrypted video frame sequence is output; the difficulty of cracking is increased, and the plaintext content is turned into unreadable ciphertext.

[0115] S402: Perform hashing on the metadata of the shards, where the metadata includes a location and a timestamp.

[0116] Chaotic encryption of fragmented video frames embedded with watermark information only protects the content itself and does not involve the structural information of the video. The structural information can be the position of the fragments in the video, the time sequence, the logical relationship between the fragments, etc. If the attacker has the structural information of the fragments, he can combine prior knowledge to carry out side channel attacks or structural analysis attacks; to prevent the fragment logic from being reverse analyzed, the metadata of the fragments is hashed and obfuscated.

[0117] According to an exemplary embodiment, the SHA-256 algorithm is used to calculate a hash value for the metadata combination to generate a fixed-length digest, which is then split into two parts: an obfuscation field used to replace the first 128 bits of the original information and concatenate it with the original metadata, and a check field used to store the last 128 bits separately to verify the authenticity of the metadata during decryption.

[0118] Replacing the original metadata with a hash value makes it difficult for attackers to restore the logical structure and order of the shards by analyzing the metadata, effectively preventing the shard logic from being reverse analyzed; using hash obfuscation and dynamic salt values ​​to protect metadata integrity and privacy, prevent the leakage of spatiotemporal information, recalculate the metadata hash during decryption, and compare the checksum fields. If they match, it is confirmed that the metadata has not been tampered with; if they do not match, it is confirmed that the metadata has been tampered with, triggering an alarm and discarding the abnormal shard.

[0119] S403: Setting up a dynamic key management mechanism to store key fragments on multiple nodes through distributed key storage of the blockchain;

[0120] Dynamic key rotation is triggered based on the playback environment; the playback environment can be an IP address or device ID. The dynamic key management mechanism is as follows:

[0121] S403a: Obtain key information about the playback environment, obtain the current IP address of the device through the network protocol stack, and read the device ID from the device system;

[0122] S403b: When a change in the IP address or device ID is detected, the system triggers a key update;

[0123] S403c: Obtain a new encryption key, including recalculating the video content hash and device fingerprint at the current playback position.

[0124] The new key is generated in the same way as the initial key. When the authorized user switches the network environment, the IP address will change. When the device is changed for playback, the device ID will change. The corresponding video content hash value and device fingerprint will also be updated. The keys used in the encryption and decryption process are updated in a timely manner to effectively prevent replay attacks.

[0125] Taking advantage of the decentralized and distributed characteristics of blockchain, the keys required for encrypted videos are stored in shards; the original key is split into r shards through a secret sharing algorithm. At least p shards are required to restore the original key, and each shard is stored on a different blockchain node. The key is divided into multiple parts, and each part is stored on a different node in the blockchain network.

[0126] According to an exemplary embodiment, the secret sharing algorithm utilizes the Shamir algorithm (r, p) threshold scheme.

[0127] The watermark information includes the authorized user ID and timestamp. Unique identification information is generated for each authorized user based on the watermark information. When a user is found to have abused behavior, the specific information of the corresponding user can be obtained by parsing the watermark, thereby tracing the source and overcoming the problem of being unable to track the abuse of authorized users.

[0128] S5: Set up a watermark self-repair mechanism and automatically match the decryption algorithm according to the terminal type.

[0129] Set the terminal-algorithm mapping table to be inserted into the encrypted video header. The terminal-algorithm mapping table includes information such as terminal type, encryption algorithm and key length; match the corresponding decryption algorithm according to the obtained terminal fingerprint information; the watermark self-repair mechanism is as follows: the terminal detects the integrity of the watermark after receiving it. After the watermark is embedded in the video, its watermark information content must be completely consistent with the originally generated watermark information, indicating that it has not been tampered with or replaced. If it has been tampered with, the hash value is used to reversely restore part of the watermark information to achieve watermark self-repair.

[0130] According to an exemplary embodiment, the timeline can be set to one block every 50ms. During decryption, the blockchain consensus mechanism is used to dynamically reorganize the key from multiple nodes to ensure that the video can be quickly and securely decrypted and played on different terminals, while reducing device resource consumption and ensuring content security. The terminal types can be H5 and Android. The PBFT (Practical Byzantine Fault Tolerance) consensus algorithm is used to ensure that even if some nodes fail or are attacked, the complete key can still be accurately and securely obtained, thereby improving the security and reliability of key storage. A GPU-accelerated parallel decryption algorithm is used to divide the encrypted video stream into blocks of equal size according to the timeline, which are transmitted to the GPU video memory via the PCIe bus. Parallel threads are started on the GPU side, and each thread is responsible for decrypting one data block. The decrypted data blocks are returned from the GPU video memory to the system memory, spliced ​​into the original video stream in chronological order, and sent to the decoder for rendering.

[0131] By setting up a sharding mechanism and utilizing dynamic sharding, the encryption calculation amount in non-sensitive areas can be reduced, processing speed and encryption efficiency can be improved, and watermarks can be embedded differently according to the characteristics of different regions. The watermark embedding strength is increased for the first-level sensitive shards to ensure that the copyright mark of the key content cannot be deleted; the embedding strength is reduced for the second-level sensitive shards and ordinary shards to reduce the amount of calculation and balance robustness and processing efficiency; chaotic encryption is superimposed on the video frame shards to increase the cost of video cracking and enhance the ability to resist attacks. The video frame shard metadata is also encrypted through hash obfuscation technology to effectively resist shard reassembly attacks. The risk of centralized server attacks is reduced through distributed key storage, and the probability of key leakage is reduced; the parallel encryption algorithm is used to support real-time video processing, and authorized users do not need to manually enter keys, reducing key management costs.

[0132] An exemplary embodiment of the present disclosure further provides a video encryption system based on dynamic segmentation and multi-layer obfuscation, comprising: an identification and rating module, a dynamic segmentation module, a watermark embedding unit, a multi-layer encryption unit, and a decryption adaptation module.

[0133] Identification and rating module: used for identifying all sensitive areas in each frame of the original image and calculating the frame-level sensitivity score for the original image;

[0134] Dynamic segmentation module: sets a segmentation mechanism based on frame-level sensitivity scores, and is used to dynamically divide the target video into segments according to the segmentation mechanism. The segmented segments include first-level sensitive segments, second-level sensitive segments, and ordinary segments.

[0135] Watermark embedding unit: used to embed the acquired watermark information into the frequency domain of the video frame after segmentation, and then convert the frequency domain of the video frame with embedded watermark information into the spatial domain, and realize the hiding of watermark by setting the visual perception threshold ∈;

[0136] Multi-layer encryption unit: used to encrypt video frames embedded with hidden watermark information using multi-layer encryption technology including slice content encryption, slice structure encryption and distributed key storage, and then slice and store the key required for encrypting the video; the multi-layer encryption unit includes a slice content encryption module, a slice structure encryption module and a key slice storage module;

[0137] Decryption adaptation module: used to set the watermark self-repair mechanism during decryption, and automatically matches the decryption algorithm according to the terminal type. The watermark self-repair mechanism is set to implement watermark self-repair after detecting incomplete watermark information during decryption.

[0138] The sharding mechanism in the dynamic sharding module is as follows:

[0139] when When , the window starts the secondary candidate state and resets the time counter;

[0140] when When , the window starts the first-level candidate state and resets the time counter;

[0141] If the time in the secondary candidate state is greater than the preset minimum fragmentation duration, the secondary candidate state is cancelled. When a sensitivity jump occurs or the time is equal to the preset maximum fragmentation duration, a secondary sensitive fragment is generated;

[0142] If the time in the first-level candidate state is greater than the preset minimum fragmentation duration, the first-level candidate state is cancelled. When a sensitivity jump occurs or the time is equal to the preset maximum fragmentation duration, a first-level sensitive fragment is generated;

[0143] The adjacent windows of the first-level candidate state without generated fragments, the second-level candidate state without generated fragments, and the normal state are merged into a normal fragment with a total duration less than or equal to the preset maximum fragment duration.

[0144] The watermark embedding unit includes a frequency domain conversion module, a watermark embedding module and a space domain conversion module;

[0145] Frequency domain conversion module: used to convert the original image from the spatial domain to the frequency domain through discrete cosine transform to obtain the frequency domain coefficient matrix F(u,v);

[0146] Watermark embedding module: adjusts the watermark strength based on the fragment sensitivity and is used to embed the watermark information into the frequency domain of the fragmented video frame;

[0147] Spatial domain conversion module: converts the video frame embedded with watermark from frequency domain to spatial domain, and hides the watermark by setting the visual perception threshold ∈.

[0148] The DCT coefficient position and frequency domain embedding position are selected according to the slice sensitivity: when the video frame slice is a first-level sensitive slice, the watermark signal is embedded in the low-frequency area; when the video frame slice is a second-level sensitive slice and a normal slice, the watermark signal is embedded in the high-frequency area.

[0149] The embedding method is adjusted according to the sensitivity of the slice: when the video frame slice is a first-level sensitive slice, the modified frequency domain coefficient matrix is ​​obtained after the watermark signal is superimposed on F(u,v); when the video frame slice is a second-level sensitive slice and a common slice, the modified frequency domain coefficient matrix is ​​obtained after the watermark signal is quantized on F(u,v).

[0150] The multi-layer encryption unit includes a slice content encryption module, a slice structure encryption module and a key slice storage module; wherein,

[0151] Sliced ​​content encryption module: Chaotic encryption is the first layer of multi-layer obfuscation encryption technology. It generates a random key stream based on Logistic mapping and performs chaotic encryption on the sliced ​​video frames embedded with hidden watermark information.

[0152] Segment structure encryption module: used to hash and obfuscate the metadata of the segments. The metadata includes the location and timestamp representing the structural information. The location refers to the position of the segment in the video, and the timestamp refers to the time when the segment was generated or created.

[0153] Key sharding storage module: Set up a dynamic key management mechanism and store key shards in multiple nodes through the distributed key storage of the blockchain.

[0154] Through the identification rating module, the dynamic segmentation module uses dynamic segmentation to reduce the encryption calculation amount of non-sensitive areas, and uses the watermark embedding unit to embed watermarks differentially according to the characteristics of different regions to ensure that the copyright identification of key content cannot be deleted, while balancing robustness and processing efficiency; through multi-layer encryption units, the video frame segmentation is superimposed with chaotic encryption to increase the cost of video cracking and enhance the anti-attack capability. The video frame segmentation metadata is also encrypted through hash obfuscation technology to effectively resist segmentation reorganization attacks. The risk of centralized server attacks is reduced through distributed key storage, and the probability of key leakage is reduced; through the decryption adaptation module, the parallel encryption algorithm is used to support real-time video processing, and authorized users do not need to manually enter the key, reducing the key management cost.

Claims

1. A video encryption method based on dynamic segmentation and multi-layer obfuscation, characterized in that: The steps include: S1: Identify all sensitive areas in each frame of the original image in the original video and calculate the frame-level sensitivity score for the original image; S2: Calculate the average sensitivity score of the frame images in the unit window based on the frame-level sensitivity score, set a fragmentation mechanism based on the average sensitivity score, and dynamically divide the target video into fragments according to the fragmentation mechanism; S3: Embed the acquired watermark information into the frequency domain of the fragmented video frame, then convert the frequency domain of the video frame with embedded watermark information into the spatial domain and perform watermark hiding processing; S4: Encrypt the video frames embedded with hidden watermark information using a multi-layer encryption technology including slice content encryption, slice structure encryption, and distributed key storage; S5: Automatically match the decryption algorithm according to the terminal type and set a watermark self-repair mechanism.

2. The video encryption method based on dynamic segmentation and multi-layer obfuscation according to claim 1, characterized in that: Calculate the average sensitivity score of the s-frame images in the unit window. The first-level sensitivity condition is that the average sensitivity score is greater than the second scoring threshold; the second-level sensitivity condition is that the average sensitivity score is less than or equal to the second scoring threshold and greater than or equal to the first scoring threshold; if the window status of the current window is different from that of the next window, it is considered that a sensitivity jump has occurred; if the average sensitivity score does not reach the first scoring threshold, it is a normal state. If the first-level sensitivity condition is met, the first-level candidate state is started, and if the second-level sensitivity condition is met, the second-level candidate state is started.

3. The video encryption method based on dynamic segmentation and multi-layer obfuscation according to claim 2, characterized in that: The sharding mechanism is as follows: The window starts in normal state; When the secondary candidate state is started, the time counter is reset; When the first-level candidate state is started, the time counter is reset; If the time in the secondary candidate state is greater than the preset minimum fragmentation duration, the secondary candidate state is cancelled. When a sensitivity jump occurs or the time is equal to the preset maximum fragmentation duration, a secondary sensitive fragment is generated; If the time in the first-level candidate state is greater than the preset minimum fragmentation duration, the first-level candidate state is cancelled. When a sensitivity jump occurs or the time is equal to the preset maximum fragmentation duration, a first-level sensitive fragment is generated; The adjacent windows of the first-level candidate state without generated fragments, the second-level candidate state without generated fragments, and the normal state are merged into a normal fragment with a total duration less than or equal to the preset maximum fragment duration.

4. The video encryption method based on dynamic segmentation and multi-layer obfuscation according to claim 1, characterized in that: Step S3 includes: S301: Obtain frequency domain coefficient matrix through discrete cosine transform ; S302: Adjust the watermark strength based on the fragment sensitivity and embed the watermark information into the fragmented video frame frequency domain; the watermark information includes the authorized user ID and timestamp; S303: Convert the watermarked video frame from frequency domain to spatial domain and set a visual perception threshold Implement hidden watermark.

5. The video encryption method based on dynamic segmentation and multi-layer obfuscation according to claim 3, characterized in that: Step S302: Adjust the embedding method according to the fragment sensitivity: When the video frame segment is a first-level sensitive segment, the first intensity gain, the weight matrix related to the frequency domain position and the watermark signal are obtained. The modified frequency domain coefficient matrix after superimposing the watermark signal on the basis; When the video frame is segmented into secondary sensitive segments and common segments, the watermark signal is obtained according to the quantization step size, the second intensity gain, the weight matrix set by the frequency domain energy distribution and the watermark signal. The modified frequency domain coefficient matrix after quantizing the watermark signal.

6. The video encryption method based on dynamic segmentation and multi-layer obfuscation according to claim 1, characterized in that: Step S4 includes: S401: Generate a random key stream based on Logistic mapping to perform chaotic encryption on the content of the sliced ​​video frames embedded with watermark information; S402: Hash the metadata of the shards, including the location and timestamp. S403: Set up a dynamic key management mechanism and store the key fragments in multiple nodes through the distributed key storage of the blockchain.

7. The video encryption method based on dynamic segmentation and multi-layer obfuscation according to claim 5, characterized in that: The S403 triggers dynamic key rotation according to the playback environment; the playback environment can be an IP address or device ID; the dynamic key management mechanism is as follows: S403a: Obtain key information about the playback environment, obtain the device's current IP address through the network protocol stack, and read the device ID from the device system; S403b: When a change in the IP address or device ID is detected, the system triggers a key update; S403c: Obtain a new encryption key, including recalculating the video content hash and device fingerprint at the current playback position.

8. The video encryption method based on dynamic segmentation and multi-layer obfuscation according to claim 1, characterized in that: The watermark self-repair mechanism is as follows: the terminal detects the integrity of the watermark after receiving it, and after the watermark is embedded in the video, When the watermark information content is completely consistent with the originally generated watermark information, it indicates that it has not been tampered with or replaced; When the watermark information content is not completely consistent with the originally generated watermark information, indicating that it has been tampered with, the watermark self-repair is achieved by reversely restoring part of the watermark information through the hash value.

9. An encryption system based on the video encryption method based on dynamic slicing and multi-layer obfuscation according to any one of claims 1 to 8, characterized in that: include: Identification and rating module, dynamic sharding module, watermark embedding unit, multi-layer encryption unit and decryption adaptation module; The recognition and rating module is used to identify all sensitive areas in each frame of the original image and perform an overall frame-level sensitivity rating. Dynamic fragmentation module: used to dynamically divide the target video into fragments according to the fragmentation mechanism. The fragments after pollen include primary sensitive fragments, secondary sensitive fragments and common fragments; Watermark embedding unit: used to embed the acquired watermark information into the video frame frequency domain after segmentation, and then convert the video frame frequency domain embedded with watermark information into the spatial domain, by setting the visual perception threshold Implement hidden watermark; Multi-layer encryption unit: used to encrypt video frames embedded with hidden watermark information using multi-layer encryption technology including slice content encryption, slice structure encryption and distributed key storage, and then slice and store the key required for encrypting the video; the multi-layer encryption unit includes a slice content encryption module, a slice structure encryption module and a key slice storage module; Decryption adaptation module: used to set the watermark self-repair mechanism during decryption, and automatically matches the decryption algorithm according to the terminal type. The watermark self-repair mechanism is set to implement watermark self-repair after detecting incomplete watermark information during decryption.

10. The video encryption method based on dynamic segmentation and multi-layer obfuscation according to claim 1, characterized in that: The watermark embedding unit includes a frequency domain conversion module, a watermark embedding module and a space domain conversion module; the multi-layer encryption unit includes a slice content encryption module, a slice structure encryption module and a key slice storage module; wherein, Frequency domain conversion module: used to convert the original image from the spatial domain to the frequency domain through discrete cosine transform to obtain the frequency domain coefficient matrix ; Watermark embedding module: adjusts the watermark strength based on the fragment sensitivity and is used to embed the watermark information into the frequency domain of the fragmented video frame; Spatial domain conversion module: converts the video frame embedded with watermark from frequency domain to spatial domain, and sets the visual perception threshold Implement hidden watermark; Sliced ​​content encryption module: Chaotic encryption is the first layer of multi-layer obfuscation encryption technology. It generates a random key stream based on Logistic mapping and performs chaotic encryption on the sliced ​​video frames embedded with hidden watermark information. Segment structure encryption module: used to hash and obfuscate the metadata of the segments. The metadata includes the location and timestamp representing the structural information. The location refers to the position of the segment in the video, and the timestamp refers to the time when the segment was generated or created. Key sharding storage module: Set up a dynamic key management mechanism and store key shards in multiple nodes through the distributed key storage of the blockchain.

Citation Information

Patent Citations

  • A method and apparatus for generating an encrypted video file library

    CN112188308B

Cited By

  • Sensitive picture adaptive shielding and authorization restoration method and device based on national cryptographic algorithm, equipment and storage medium

    CN121585866A