An information steganography method and system based on a video carrier

By processing sequence frames in the video carrier, obtaining attention weights, and generating dense videos in combination with information sequences and original videos, the problem of difficulty in processing time domain information in the video carrier is solved in the prior art, and the security and effectiveness of information steganography are improved.

CN114979667BActive Publication Date: 2025-05-30ENG UNIV OF THE CHINESE PEOPLES ARMED POLICE FORCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210514913.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-11
Publication Date
2025-05-30
Estimated Expiration
2042-05-11

AI Technical Summary

Technical Problem

The existing digital steganography mainly uses digital images as the carrier, making it difficult to effectively process the time domain information in the video carrier, resulting in poor robustness.

Method used

By processing the sequence frames in the video carrier, attention weights are obtained, and dense videos are generated in combination with the information sequence and the original video, thus realizing the processing of time domain information in the video carrier and information steganography.

Benefits of technology

It improves the security and effectiveness of embedded secret information in videos and enhances the processing capability of time domain information of video carriers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114979667B_ABST
    Figure CN114979667B_ABST
Patent Text Reader

Abstract

The present invention discloses an information steganography method based on a video carrier, comprising: processing an original video according to a preset number of sequential frames to obtain sequential frames; processing secret information to be sent according to the embedding capacity of the sequential frames to obtain an information sequence; obtaining attention weights according to the sequential frames; generating a stego video according to the attention weights, the information sequence and the original video; obtaining the information sequence according to the stego video; and restoring the secret information according to the information sequence. By analyzing and processing the time domain and spatial domain in the sequential frames to obtain attention weights, the present invention can process the time domain information in the video carrier, thereby effectively performing information steganography with the video as the carrier, and improving the security and effectiveness of embedding secret information in the video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security technology, and in particular, to an information steganography method and system based on a video carrier. Background Art

[0002] Steganography is a secret communication technology that hides secret information in carrier information so that the attacker cannot know whether the carrier contains secret information, thus achieving the purpose of covert transmission. It has a very wide range of applications in public services, network transmissions, military communications, etc. Compared with encryption technology, due to its imperceptibility, it is not easily analyzed and detected by malicious attackers, and is currently a hot topic in the field of information security, having important application value in military, intelligence and other departments.

[0003] However, most current digital steganography uses digital images as carriers, and relatively few studies have been conducted on digital video steganography technology. Considering that digital video has a larger absolute data volume than images, its embedding capacity and security perform better. With the development of 5G high-speed networks, a large amount of video media information has spread rapidly on the Internet, and digital video has become a commonly used media form and an ideal data hiding carrier. Therefore, the advantages of video steganography in terms of the number of carriers and transmission security are becoming increasingly prominent.

[0004] Due to the complex structure of video data, the video information flow mainly includes spatial domain information and temporal domain information. The existing image steganography convolutional networks cannot process the temporal domain information in the video carrier, resulting in poor robustness.

[0005] Therefore, how to process the temporal domain information in the video carrier to effectively perform information steganography with the video as the carrier is an urgent problem to be solved. Summary of the Invention

[0006] In view of this, the present invention provides an information steganography method and system based on a video carrier, which can process the temporal domain information in the video carrier to effectively perform information steganography with the video as the carrier.

[0007] The present invention provides an information steganography method based on a video carrier, including the following steps:

[0008] Process the original video according to a preset number of sequence frames to obtain sequence frames;

[0009] Process the secret information to be sent according to the embedding capacity of the sequence frames to obtain an information sequence;

[0010] Obtain attention weights according to the sequence frames;

[0011] Generate a stego video based on the attention weight, the information sequence, and the original video;

[0012] Obtain the information sequence according to the stego video;

[0013] Restore the secret information according to the information sequence.

[0014] Preferably, the processing of the secret information to be sent in advance according to the embedding capacity of the sequence frames to obtain an information sequence includes:

[0015] Convert the secret information into a binary sequence string;

[0016] Encode the binary sequence string to obtain an encoded binary sequence string;

[0017] Segment the encoded binary sequence string to obtain a plurality of the information sequences with the same embedding capacity as the sequence frames, wherein the insufficient bits of the last information sequence are filled with 0.

[0018] Preferably, the generating of the stego video according to the attention weight, the information sequence, and the original video includes:

[0019] Obtain an information mask according to the information sequence and the attention weight;

[0020] Generate a stego video according to the information mask and the original video.

[0021] Preferably, the obtaining of the information mask according to the information sequence and the attention weight includes:

[0022] Transform the information sequence to obtain an information tensor;

[0023] Obtain the information mask according to the information tensor and the attention weight.

[0024] Preferably, the generating of the stego video according to the information mask and the original video includes:

[0025] Obtain residual data according to the information mask and the original video;

[0026] Generate a stego video according to the residual data and the original video.

[0027] Preferably, the obtaining of the information sequence according to the stego video includes:

[0028] Obtain the attention weight according to the stego video;

[0029] Extract the residual data according to the stego video;

[0030] Obtain the information sequence according to the attention weight and the residual data.

[0031] Preferably, the obtaining the information sequence according to the attention weight and the residual data includes:

[0032] Obtain the average value of the product of the attention weight and the residual data in the first to third dimensions;

[0033] Obtain the information sequence based on the average value.

[0034] Preferably, the restoring the secret information according to the information sequence includes:

[0035] Decode the information sequence to restore the secret information.

[0036] The present invention also provides an information steganography system based on a video carrier, including an embedding end and an extraction end; wherein, the embedding end includes a first processing module, a second processing module, an attention module, and a generation module; the extraction end includes an acquisition module and a restoration module; wherein:

[0037] The first processing module is configured to process the original video according to a preset number of sequence frames to obtain sequence frames;

[0038] The second processing module is configured to process the pre-transmitted secret information according to the embedding capacity of the sequence frames to obtain an information sequence;

[0039] The attention module is configured to obtain an attention weight according to the sequence frames;

[0040] The generation module is configured to generate a stego video according to the attention weight, the information sequence, and the original video;

[0041] The acquisition module is configured to obtain the information sequence according to the stego video;

[0042] The restoration module is configured to restore the secret information according to the information sequence.

[0043] Preferably, it further includes:

[0044] An adversarial end, configured to simulate the noise attack existing in the channel transmission of the stego video.

[0045] In summary, the present invention provides a method for information steganography based on a video carrier, including: processing an original video according to a preset number of sequential frames to obtain sequential frames; processing pre-transmitted secret information according to the embedding capacity of the sequential frames to obtain an information sequence; obtaining an attention weight according to the sequential frames; generating a stego video according to the attention weight, the information sequence, and the original video; obtaining the information sequence according to the stego video; and restoring the secret information according to the information sequence. By analyzing and processing the time domain and spatial domain in the sequential frames to obtain the attention weight, the present invention can process the time domain information in the video carrier, thereby effectively performing information steganography with the video as the carrier, and improving the security and effectiveness of embedding secret information in the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1 It is a flowchart of a method for information steganography based on a video carrier provided by the present invention;

[0048] Figure 2 It is a schematic diagram of generating an information mask according to an information tensor and an attention weight in a method for information steganography based on a video carrier provided by the present invention;

[0049] Figure 3 It is a first schematic diagram of a composition of an information steganography system based on a video carrier provided by the present invention;

[0050] Figure 4 It is a second schematic diagram of a composition of an information steganography system based on a video carrier provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] In order to enable those skilled in the art to better understand the technical solutions in the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0052] As Figure 1 shown, an embodiment of the present invention provides a method for information steganography based on a video carrier, including:

[0053] S1. Process the original video according to the preset number of sequential frames to obtain sequential frames;

[0054] S2. Process the secret information to be sent according to the embedding capacity of the sequential frames to obtain an information sequence;

[0055] S3. Obtain attention weights according to the sequential frames;

[0056] S4. Generate a stego video according to the attention weights, the information sequence, and the original video;

[0057] In this embodiment, the embedding end obtains attention weights according to the sequential frames. Specifically, the embedding end inputs the sequential frames into the attention module to obtain the attention weights output by the attention module. Among them, the attention module is a module shared by the embedding end for embedding secret information and the extraction end for extracting secret information, and the attention module is obtained by training the initial attention module based on the sequential frames. In this embodiment, the attention module used is a four-layer three-dimensional convolutional network with a convolutional kernel of 3×3×3, a stride of 1, and a number of 64. The structure of the attention module is shown in Table 1:

[0058] Table 1 Structure of the attention module

[0059]

[0060] As shown in Table 1, the input layer first transforms the pixel values of the input sequential frames from the [0:255] interval to the [0:1] interval. (The pixel values are generally distributed in the interval from 0 to 255. For convenient data processing, they are normalized and transformed to between 0 and 1). 3DConv-BN-ReLU(64) in the intermediate layer 1 to the intermediate layer 3 represents a three-dimensional convolutional layer with a convolutional kernel size of 3×3×3, a stride of 1, and a number of 64, followed by a batch normalization BN layer and a ReLU activation function in sequence. 3DConv-BN-ReLU(d)-Softmax in the intermediate layer 4 represents a three-dimensional convolutional layer with a convolutional kernel size of 3×3×3, a stride of 1, and a number of 64, followed by a batch normalization BN layer, a ReLU activation function, and a Softmax layer in sequence. t, h, and w respectively represent the time dimension, height dimension, and width dimension of the video.

[0061] S5. Obtain the information sequence according to the stego video;

[0062] S6. Restore the secret information according to the information sequence.

[0063] Compared with the prior art, since the video sequence frames mainly contain two types of information, in addition to the spatial domain information, they also include the temporal domain information. In the existing image convolution networks, generally only the spatial domain information of the image is analyzed, ignoring the importance of the temporal domain information. The attention module proposed in this embodiment obtains the attention weights by analyzing and processing the spatial domain information and the temporal domain information in the sequence frames of the video, which can effectively solve the problem that the existing convolution networks cannot process the time information in the video carrier, and can process the temporal domain information in the video carrier, thereby effectively performing information steganography with the video as the carrier, improving the security and effectiveness of embedding secret information in the video.

[0064] In addition, compared with the prior art of performing information steganography based on picture carriers, since the video has many sequence frames, a large amount of information can be embedded in them, and even if a certain frame in the sequence frames is damaged, the embedding of information can still be guaranteed.

[0065] As a preferred embodiment, in the embodiment of the present invention, step S2 includes:

[0066] A1. Convert the secret information into a binary sequence string;

[0067] A2. Encode the binary sequence string to obtain an encoded binary sequence string;

[0068] In this embodiment, the binary sequence string is encoded using BCH codes, which can meet the requirements of lossless information transmission. Among them, the full name of BCH code is Bose Chaudhuri Hocquenghem code. BCH code is a linear block code in a finite field and has the ability to correct multiple random errors. It is usually used for error correction coding in the fields of communication and storage.

[0069] A3. Divide the encoded binary sequence string to obtain several information sequences with the same embedding capacity as the sequence frames. Among them, the insufficient bits of the last information sequence are filled with 0.

[0070] As a preferred embodiment, in the embodiment of the present invention, step S4 includes:

[0071] B1. Obtain an information mask according to the information sequence and the attention weights;

[0072] B2. Generate a stego video according to the information mask and the original video.

[0073] As a preferred embodiment, in the embodiment of the present invention, step B1 includes:

[0074] C1. Transform the information sequence to obtain an information tensor;

[0075] In this embodiment, specifically, an information sequence of size (1, d) is spatially replicated according to the first to third dimensions of the attention weight to be expanded into an information tensor of size (t, h, w, d).

[0076] C2. Obtain an information mask according to the information tensor and the attention weight.

[0077] In this embodiment, specifically, the attention weight and the information tensor are subjected to Hadamard product to obtain a data tensor of size (t, h, w, d), and the data tensor is averaged in the fourth dimension to obtain an information mask of size (t, h, w, 1).

[0078] In this embodiment, a schematic diagram of generating an information mask according to the information tensor and the attention weight is as Figure 2 shown.

[0079] As a preferred implementation manner, in the embodiment of the present invention, step B2 includes:

[0080] D1. Obtain residual data according to the information mask and the original video;

[0081] In this embodiment, specifically, the information mask and the original video are input into an embedding module. The input layer in the embedding module first connects the information mask and the original video in the fourth dimension to obtain a tensor of size (t, h, w, 4), and then inputs it into embedding layer 1 and embedding layer 2 in sequence to obtain a residual data of size (t, h, w, 3).

[0082] D2. Generate a encrypted video according to the residual data and the original video.

[0083] In this embodiment, specifically, the residual data and the original video are subjected to matrix summation in the fourth dimension to generate an encrypted video. The specific embedding module structure is shown in Table 2:

[0084] Table 2 Embedding module structure

[0085]

[0086] As shown in Table 2, 3DConv-BN-ReLU(64) in embedding layer 1 represents a three-dimensional convolutional layer with a convolutional kernel size of 3×3×3, a stride of 1, and a number of 64, followed by a batch normalization BN layer and a ReLU activation function in sequence; 3DConv-ReLU(3) in embedding layer 2 represents a three-dimensional convolutional layer with a convolutional kernel size of 3×3×3, a stride of 1, and a number of 3, followed by a ReLU activation function.

[0087] As a preferred implementation manner, in the embodiment of the present invention, step S5 includes:

[0088] E1. Obtain the attention weight according to the encrypted video;

[0089] In this embodiment, specifically, the extraction end inputs the received encrypted video into the attention module to obtain the attention weight output by the attention module. The attention module used here is the same as the one used by the embedding end, with the same structure and parameters as those in the module used by the embedding end. Therefore, the module does not need to be trained separately.

[0090] E2. Extract the residual data according to the encrypted video;

[0091] In this embodiment, specifically, the extraction end inputs the received encrypted video into the extraction module to obtain the residual data output by the extraction module. First, the encrypted video is used as the input of the extraction module, and the information residuals hidden in the encrypted video are extracted using three convolutional layers. At the same time, the encrypted video is input into the attention module, and then the Hadamard product is performed on the obtained attention weight of the encrypted video and the information residuals. The average value is calculated in the first to third dimensions of the product result, and the size of the result is changed from (t, h, w, d) to (1, d), aggregating the secret information in the scattered spatio-temporal dimensions. In addition, the specific structure of the extraction module is shown in Table 3:

[0092] Table 3 Structure of the extraction module

[0093]

[0094] As shown in Table 3, 3DConv-BN-ReLU(64) in extraction layer 1 represents a three-dimensional convolutional layer with 64 convolutional kernels, a size of 3×3×3, and a stride of 1, followed by a batch normalization BN layer and a ReLU layer in sequence; 3DConv-BN-ReLU(128) in extraction layer 2 represents a three-dimensional convolutional layer with 128 convolutional kernels, a size of 3×3×3, and a stride of 1, followed by a batch normalization BN layer and a ReLU layer in sequence; 3DConv-ReLU(d) in extraction layer 3 represents a three-dimensional convolutional layer with a convolutional kernel size of 3×3×3, a stride of 1, and a number of 64, followed by a ReLU layer.

[0095] E3. Obtain the information sequence according to the attention weight and the residual data.

[0096] As a preferred implementation manner, in the embodiment of the present invention, step E3 includes:

[0097] F1. Obtain the average value of the product of the attention weight and the residual data in the first to third dimensions;

[0098] F2. Obtain the information sequence based on the average value.

[0099] As a preferred embodiment, in the embodiment of the present invention, step S6 includes:

[0100] Decode the information sequence and recover the secret information.

[0101] In this embodiment, since the embedding end uses BCH code for encoding, the extraction end uses BCH code to decode the information sequence and recover the secret information.

[0102] The present invention also provides an information steganography system based on a video carrier, including an embedding end 301 and an extraction end 302; wherein, the embedding end 301 includes a first processing module, a second processing module, an attention module, and a generation module; the extraction end 302 includes an acquisition module and a restoration module; wherein:

[0103] The first processing module is used to process the original video according to the preset number of sequence frames to obtain sequence frames;

[0104] The second processing module is used to process the pre-transmitted secret information according to the embedding capacity of the sequence frames to obtain an information sequence;

[0105] The attention module is used to obtain attention weights according to the sequence frames;

[0106] The generation module is used to generate a stego video according to the attention weights, the information sequence, and the original video;

[0107] The acquisition module is used to obtain the information sequence according to the stego video;

[0108] The restoration module is used to restore the secret information according to the information sequence.

[0109] The working principle of the information steganography system based on a video carrier provided in this embodiment is the same as that of the method embodiment of the information steganography system based on a video carrier above, and will not be elaborated here.

[0110] As a preferred embodiment, in the embodiment of the present invention, it further includes:

[0111] An adversarial end 303, which is used to simulate the noise attack existing in the channel transmission of the stego video.

[0112] In this embodiment, specifically, since the embedding end transmits the stego video to the extraction end through the channel, but due to the general existence of noise attacks in the public channel, which causes information distortion of the stego video, the adversarial end is used to simulate the noise in the channel. Based on the sequence frames, a convolutional network is used as a tool for generating adversarial attack samples to train the robustness of the initial attention module to distortion and obtain the final attention module. The convolutional network structure is designed as:

[0113]

[0114] G adv (I) is an adversarial distortion sample, Conv3d 16 and Conv3d 3 are 3D convolutional networks with convolutional kernel sizes of 3×3×3 and numbers of channels of 16 and 3 respectively. LeakyReLU is used as the activation function to process the output of the first convolutional network and then input it into the next convolutional network, and finally a distortion sample with 3 channels is output. To train the attack network G adv , we will minimize the following adversarial training loss:

[0115]

[0116] and are the loss weights for encoding and decoding respectively, ||I adv -I en || 2 is the L2 loss between the adversarial distortion sample I adv and the encoded sample I en . X is the embedded secret information, F dec (I adv ) is the distorted information decoded from the adversarial distortion sample I adv , and L M is set to the L2 loss of the information.

[0117] In this specification, the embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other. In this specification, the embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0118] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0119] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in a software module executed by a processor, or in a combination thereof. The software module may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0120] The foregoing description of the disclosed embodiments enables those skilled in the art to make or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An information steganography method based on a video carrier, characterized in that, it includes: Processing the original video according to a preset number of sequence frames to obtain sequence frames; Processing the secret information to be sent according to the embedding capacity of the sequence frames to obtain an information sequence; Obtaining attention weights according to the sequence frames; Generating a stego video according to the attention weights, the information sequence and the original video; Obtaining the information sequence according to the stego video; Restoring the secret information according to the information sequence; The obtaining of the attention weights according to the sequence frames includes: The input layer of the attention module receives the input sequence frames, transforms the pixel values of the input sequence frames from the [0:255] interval to the [0:1] interval, and outputs data with a structure of (t, h, w, 3). The three-dimensional convolutional layer, batch normalization BN layer and ReLU activation function in the middle layer 1 to middle layer 3 of the attention module with a convolutional kernel size of 3×3×3, a stride of 1 and a number of 64 process the data output by the input layer and output data with a structure of (t, h, w, 64). The three-dimensional convolutional layer, batch normalization BN layer, ReLU activation function and Softmax layer in the middle layer 4 of the attention module with a convolutional kernel size of 3×3×3, a stride of 1 and a number of 64 process the data output by the middle layer 3 and output data with a structure of (t, h, w, d). The output layer of the attention module outputs attention weights with a structure of (t, h, w, d), where t, h, and w respectively represent the time dimension, height dimension, and width dimension of the video; The generating of the stego video according to the attention weights, the information sequence and the original video includes: Obtaining an information mask according to the information sequence and the attention weights; Generating a stego video according to the information mask and the original video.

2. The method according to claim 1, characterized in that, The processing of the secret information to be sent according to the embedding capacity of the sequence frames to obtain an information sequence includes: Converting the secret information into a binary sequence string; Encoding the binary sequence string to obtain an encoded binary sequence string; Segmenting the encoded binary sequence string to obtain a plurality of the information sequences with the same embedding capacity as the sequence frames, where the insufficient bits of the last information sequence are filled with 0.

3. The method according to claim 2, characterized in that, The obtaining of the information mask according to the information sequence and the attention weights includes: Transforming the information sequence to obtain an information tensor; Obtaining the information mask according to the information tensor and the attention weights.

4. The method according to claim 3, characterized in that, The generating of the stego video according to the information mask and the original video includes: Obtaining residual data according to the information mask and the original video; Generating a stego video according to the residual data and the original video.

5. The method according to claim 4, characterized in that, Obtaining the information sequence according to the encrypted video includes: Obtaining the attention weight according to the encrypted video; Extracting the residual data according to the encrypted video; Obtaining the information sequence according to the attention weight and the residual data.

6. The method according to claim 5, wherein, obtaining the information sequence according to the attention weight and the residual data includes: Obtaining the average value of the product of the attention weight and the residual data in the first to third dimensions; Obtaining the information sequence based on the average value.

7. The method according to claim 6, wherein, restoring the secret information according to the information sequence includes: Decoding the information sequence to restore the secret information.

8. An information steganography system based on a video carrier, wherein, it includes an embedding end and an extraction end; among them, the embedding end includes a first processing module, a second processing module, an attention module, and a generation module; the extraction end includes an acquisition module and a restoration module; wherein: The first processing module is used to process the original video according to the preset number of sequence frames to obtain sequence frames; The second processing module is used to process the pre-transmitted secret information according to the embedding capacity of the sequence frames to obtain an information sequence; The attention module is used to obtain the attention weight according to the sequence frames; The generation module is used to generate an encrypted video according to the attention weight, the information sequence, and the original video; The acquisition module is used to obtain the information sequence according to the encrypted video; The restoration module is used to restore the secret information according to the information sequence; When the attention module executes obtaining the attention weight according to the sequence frames, it specifically is used for: The input layer of the attention module receives the input sequence frames, transforms the pixel values of the input sequence frames from the [0:255] interval to the [0:1] interval, and outputs data with a structure of (t, h, w, 3). The three-dimensional convolutional layer, batch normalization BN layer, and ReLU activation function in the middle layer 1 to middle layer 3 of the attention module with a convolutional kernel size of 3×3×3, a stride of 1, and a number of 64 process the data output by the input layer, and output data with a structure of (t, h, w, 64). The three-dimensional convolutional layer, batch normalization BN layer, ReLU activation function, and Softmax layer in the middle layer 4 of the attention module with a convolutional kernel size of 3×3×3, a stride of 1, and a number of 64 process the data output by the middle layer 3, and output data with a structure of (t, h, w, d). The output layer of the attention module outputs the attention weight with a structure of (t, h, w, d), where t, h, and w respectively represent the time dimension, height dimension, and width dimension of the video; When the generation module executes generating the encrypted video according to the attention weight, the information sequence, and the original video, it specifically is used for: Obtaining an information mask according to the information sequence and the attention weight; Generate a cipher-containing video based on the information mask and the original video.

9. The system according to claim 8, wherein, it further comprises: an adversarial end for simulating noise attacks existing in the channel transmission of the cipher-containing video.

Citation Information

Patent Citations

  • Video processing method and device, electronic equipment and storage medium

    CN111988672A

  • Digital media protection text steganography method based on variational automatic encoder

    CN113987129A