Method and system for robust protection of the authorship of generated videos
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PUBLICHNOE AKTSIONERNOE OBSHCHESTVO SBERBANK ROSSII (PAO SBERBANK)
- Filing Date
- 2025-12-17
- Publication Date
- 2026-07-30
Smart Images

Figure RU2025000407_30072026_PF_FP_ABST
Abstract
Description
A METHOD AND SYSTEM FOR ROBUST PROTECTION OF AUTHORITY OF VIDEOS GENERATED BY A MACHINE LEARNING MODEL. FIELD OF TECHNOLOGY
[0001] The claimed solution relates to the field of information technology, namely to means of protecting digital data for the purpose of confirming their authenticity and authorship. LEVEL OF TECHNOLOGY
[0002] With the widespread use of artificial intelligence (AI) and various machine learning models today, which can generate video files (and any other ordered sequence of images) based on a user's text query, there is a need to apply mechanisms to protect the generated videos for the purpose of subsequent confirmation of their authenticity.
[0003] One of the imperatives of this type of mechanism is to prevent attribution of AI-generated videos. However, attribution may be illegally attributed to: 1. All video information in its entirety. 2. Separate sequences of frames or individual frames. 3. The contents of the frames themselves or the contents of groups of frames. An attacker may know or suspect that a video has a digital signature to protect authorship. They can use video transformation techniques to destroy it, such as compression / stretching / frame warping, video re-encoding, video color correction, video resolution changes, video compression, and others. Another issue concerns protecting videos from third-party additional changes or modifications that distort the original AI-generated video, which could subsequently be used as false information discrediting the creators of such models for their ability to generate videos with prohibited or obscene content.
[0004] One example of video protection in the prior art is the use of a kind of digital signatures embedded in video that are invisible to the human eye and can only be obtained by subsequent decryption of the protected videos using their algorithmic extraction (see the article "Using Steganographic and Cryptographic Tools to Protect Video Files" / / B. S. Markin, A. V. Chernyshova, 2014, URL: http: / / ea.donntu.ru:8080 / bitstream / l 23456789 / 26611 / 1 / %D 1%81 %D 1 %82%D0%B0%D 1 %8 2%Dl%8C%Dl%8F%201.pdf). According to the known solution, the signature is created in the spatial domain by slightly changing the intensity level of the image pixels. The dissertation work "Model and Method for Protecting Video Information from Integrity Threats Using Steganography" by I. R. Yu. Martimov is well known.The drawback of these approaches is that the embedded signature cannot be uniquely identified with a given video—there is no video hash or equivalent. Under certain conditions, this signature can be extracted from one video file and transferred to another.
[0005] A method is known for crypto-processing any digital data, including existing digital video, using computer technology, as a result of which, for example, according to GOST R 34.10-2012 [1] and GOST R 34.11-2012 [2], video data is signed with a so-called digital electronic signature (DS).
[0006] A method for embedding and transmitting information in a video image is disclosed in patent RU 2608150 C2. The technical result is to minimize distortion of the embedded video image while ensuring the stego-resistance of the information transmission system. A method for covertly transmitting data in a video image according to the MPEG-2 (H.262) standard is proposed. This method is based on modifying the least significant bits of a video frame with the values of a two-dimensional nonlinear code combination carrying the secretly transmitted information. The formation of a steganographic channel begins with processing the embedded data, including encryption and modulation with a pseudo-random signal, which is selected as two-dimensional nonlinear Frank-Walsh and / or Frank-Chrestenson signals. Simultaneously with the formation of the stego-signal, frames for its embedding are selected, assuming all I-frames, as well as B- and P-frames, are suitable.Stegochannel data is embedded only into those DCT coefficients located in the vicinity of the right diagonal of the DCT coefficient matrix written to the JPEG file and supplemented with system information and universal Huffman tables by modulo-two addition of the DCT coefficient bits. The disadvantage of this approach is that the information embedding performed lacks authorship verification capabilities, and the embedded signature cannot be uniquely identified with the given video.
[0007] A method for confirming photograph authorship is described in patent RU2633185C2 20171011. A drawback of this approach is that electronic signature information is stored in service fields of various image file formats. Information stored in this way is always detectable in the file and can be extracted or removed from the file without changing the image itself. A method for hiding data in an image is described in the publications of W. Bender, D. Gruhl, N. Morimoto, and A. Lu, "Techniques for Data Hiding," IBM Systems Journal, Vol. 35, Nos. 3&4, 1996, and Daniel Gruhl and Walter Bender, "Information Hiding to Foil the Casual Counterfeiter," 2nd Information Hiding Workshop, 1998.The drawback of these approaches is that they are poorly resilient to image compression / stretching / warping, lack the ability to control the stability and size of the embedding via a robustness coefficient, and lack the ability to extract and embed a unique value (identifier) into the image. Therefore, even applying such approaches to video files will not allow for the creation of non-extractable or irreplaceable labeling of video information or its parts that is resistant to video transformation. ESSENCE OF THE INVENTION
[0008] The present invention is aimed at solving the technical problem of creating a new, effective method for protecting digital videos with the possibility of subsequent confirmation of their authenticity and possible changes.
[0009] The technical result is to increase the efficiency of protecting AI-generated videos for the purpose of determining their authenticity.
[0010] An additional technical result is the ability to determine the immutability of video based on embedded security information. [OOP] In a preferred embodiment of the invention, a method is claimed for protecting the authenticity of videos generated based on a text query by a machine learning model, comprising the steps of: a) receive user input data to generate a video; b) generate a video file using a machine learning model; c) recording at least one type of user input data selected from the group: input time, user ID, or text query; d) extracting key 1-frames from the generated video file; ze) extract pixels of at least one color component from the received 1-frames; f) each color component is divided into blocks of pixels of fixed size Wb x Hb; g) using a given key, determine the mixing of 1-frames and their pixel blocks obtained in step e), form the first sequence of blocks from the first N blocks of the mixed sequence and the second sequence from all the remaining blocks; h) calculating a unique video identifier based on pairs of pixel blocks in 1-frames from the first sequence of blocks formed in step g), wherein for each pair of pixel blocks in 1-frames, the average value of the pixels in each block is calculated; i) forming an attachment for protecting the generated video, consisting of a sequence including at least one of: a unique video identifier formed in step h), the input time, the user ID, a text query and additional information; j) generate an encrypted sequence by encrypting the data obtained in step i) with a private key; k) embedding the encrypted sequence obtained in step [) in the form of a sequence of bits into pairs of pixel blocks from 1-frames of the generated video included in the second sequence of blocks determined in step g); l) provide protected video to the user.
[0012] In one particular embodiment, at step d), P-frames and / or B-frames are additionally extracted.
[0013] In another particular example of implementation, the color component is selected from the group: color space, color model, color channel, subtractive model of frames of the generated video.
[0014] In another particular example of implementation, the color component is at least one of: RGB, RGBA, HSL, HSLA, CMY, CMYK, XYZ, LMS, Hunt, ICtCp, LAB, RLAB, YCbCr, in grayscale gradations of the components of the frames of the generated video.
[0015] In another particular implementation example, the frame sizes of the generated video are adjusted to a given resolution.
[0016] In another particular example of implementation, the robustness coefficient R is determined for blocks of pixels in the first sequence.
[0017] In another particular example of implementation, a value is determined based on R for determining blocks for embedding bits of the sequence.
[0018] In another particular example of implementation, error-correcting coding is applied to the generated embedding at step j).
[0019] In another particular example of implementation, at step k), the value of the unique video identifier obtained at step h) is embedded.
[0020] In another particular example of implementation, error-correcting coding is applied to the generated encrypted sequence at step k).
[0021] In another particular implementation example, error-correcting coding is applied to the unique video identifier calculated in step h).
[0022] In another particular example of implementation, at step c), compressed images of one or more frames from the generated video are additionally formed.
[0023] In another particular example of implementation, a sequence of bytes is formed based on the compressed images of frames, which is added to the data at step i) when forming the embedding.
[0024] In another particular example of implementation, the video is generated according to a standard selected from the group: H.261, H.262, H.263, H.264, H.264+, H.265, H.265+, AVI, AVI, AVC1, VC-1, VP5, VP6, VP7, VP8, VP9, WebM, ASF, MCF or another standard.
[0025] In another preferred embodiment, a system for protecting the authenticity of videos generated based on a text query by a machine learning model is claimed, comprising at least one processor and at least one memory associated with the processor, storing machine-readable instructions that, when executed by the processor, implement the above-mentioned method. BRIEF DESCRIPTION OF DRAWINGS
[0026] Fig. 1 illustrates a general view of the video protection method.
[0027] Fig. 2 illustrates a block diagram of the claimed method.
[0028] Fig. 3A illustrates an example of the formation of blocks of pixels associated with the pixels of a certain frame.
[0029] Fig. ЗБ illustrates an example of obtaining sequences of pairs of pixel blocks from frames for calculating a unique identifier of the generated video and for embedding security information.
[0030] Fig. 4 illustrates an example of determining bits for obtaining a unique identifier of the generated video.
[0031] Fig. 5 illustrates an example of the change in average pixel values within frame blocks after embedding security information.
[0032] Fig. 6 illustrates an example of a generated attachment and its parts in the form of a byte sequence.
[0033] Fig. 7 illustrates a view of the computing device. IMPLEMENTATION OF THE INVENTION
[0034] Fig. 1 shows a general diagram of the claimed solution, which is implemented by processing incoming user requests (110), which represent a textual description of the generated video. In the field of AI, such textual requests, describing the generated object for the machine learning model, are also called "Prompts." The user's request (110) can be transmitted via an API (120) implemented in a web browser or messenger (e.g., Telegram), which subsequently enables data exchange directly with the video generation model (130) (e.g., Kandinsky) between the user (110) and the video protection module (140).
[0035] The video protection module (140) can be implemented using an external hardware and software system, such as a server or remote workstation, providing functionality for processing incoming videos (121) generated by the model (130). The module (140) can also be located in the same execution loop as the model (130), such as a server or cloud computing node.
[0036] Upon completion of the video protection module (140), security information containing an encrypted attachment is embedded into the initially generated videos (121). This information allows for at least the authorship of the video to be established as being created using the model (130), as well as the user ID of the user who generated the request and the request text itself. The protected video (141) is subsequently transmitted to the user (110) via the API (120), for example, by displaying / playing the video (141) in a web browser or instant messenger.
[0037] Data exchange within the framework of the system implementation is based on standard principles of data exchange, in particular with the help of the Internet computing network, which can be implemented using any known principles known from the state of the art.
[0038] Fig. 2 shows an example of implementing the claimed method (200) for protecting generated videos. In the first step (201), user input is received for video generation by the model (130). The user input includes a user text query, a user ID (e.g., IP address, MAC address, registration ID, etc.), and the time of the text query. This data is recorded by the video protection module (140) for subsequent generation of security information for embedding in the video.
[0039] In one particular embodiment of the invention, a machine learning model generates video using the H.262 protocol (video codec). The H.262 video codec [6] is part of the MPEG-2 Part 2 family of digital video and audio coding standards. Initially, this codec could encode video with a resolution of 720*480 pixels at 30 frames per second; with further improvements, compression of a resolution of 1920*1080 pixels at 30 frames per second was achieved. Compression is based on the elimination of spatial and temporal redundancies.
[0040] To address the issue of spatial and temporal redundancy in a video stream, the codec introduces the concept of a group of pictures (GOP), and frames in a GOP are divided into three types: basic I- (intrapictures), predicted P- (predicted), and bidirectional B- (bidirectional) frames. I-frames are the input to the decoder and are coded independently in accordance with the method used in the JPEG standard [3-5]. P-frames are coded based on prediction, by referencing blocks of previous I- or P-frames. B-frames use references to two frames located before and after them. A GOP must always begin with an I-frame, from which P- and B-frames are predicted.
[0041] The number of P- and B-frames can vary depending on the desired degree of compression: the more P- and B-frames, the greater the compression, but if the 1-frame is encoded incorrectly, or information is lost in the I-frame, the error will extend to the entire group of frames.
[0042] Video image compression occurs as follows. The frame sequence is divided into 16x16 sample macroblocks, as in the JPEG algorithm, and divided into three frame types: I-, P-, and B-frames. I-frames are compressed using the JPEG algorithm: in each 16x16 sample macroblock, a transition to the YCrCb color system is performed, the Ct and Cb components are decimated to 8x8 matrices (every second row and every second column are removed), a discrete cosine transform (DCT) of the matrices is performed, their quantization, and finally zig-zag scanning and compression of the resulting sample sequence using the Huffman method. P-frames no longer transmit the fully encoded image, but a prediction error taking into account motion compensation—the difference between the previous and subsequent frames with motion vectors.
[0043] Motion vectors are calculated as follows: each block from the previous frame is compared with the block from the next. If they are identical, no motion has occurred, and there is no motion vector. If not, the block is moved within a search region of the main image to find its position at which the root-mean-square difference between the block and the main image fragment is minimal. This vertical and horizontal displacement of the block relative to its original position is taken as the motion vector. When encoding, B-frames reference two frames: one in front of the frame and one behind it. This means they convey information about how to "assemble" the frame from the two surrounding frames. This is because each frame is transmitted as a reference to the main frames, rather than individually. The MPEG2 standard is well suited for storing movies, but due to its group-of-frame structure, it is completely unsuitable for video editing.Problems also arise when compressing videos with fast-moving objects or frequently changing shots. The object will be blurred, and if the 1-frame is not captured at the moment of the change in shot, all the information will be written to the P-frame, resulting in an increased file size.
[0044] In one particular embodiment of the invention, a machine learning model generates video using the H.264 protocol (video codec). The H.264 video codec [7] is part of the MPEG4 Part 10 standard group. This recommendation is currently the video compression standard. It is used to record high-definition video on Blu-ray and HD DVD, and is the standard for online video hosting services such as YouTube, as well as a standard in digital television broadcasting systems. Compared to H.262, this codec provides twice the compression with the same video image quality. This is achieved due to the significant complexity of the codec. This codec implements frame division not only into 16x16 macroblocks, but also 16x8, 8x16, 8x8, 8x4, 4x8, 4x4, depending on the presence of small details. This increases the clarity of the transmission of small objects and the quality of motion compensation: greater accuracy in the representation of motion vectors is ensured.The motion vector search accuracy is 1 / 4 or 1 / 8 of a macroblock, which was not the case in the H.262 codec. The intra-frame coding has been modified. The coding is also based on the JPEG algorithm, but with some additions. Before the DCT operation on a macroblock, its spatial prediction (intra-prediction) is performed based on the already encoded adjacent macroblocks. Then, the luminance and color difference samples of the predicted macroblock are subtracted from the corresponding samples of the macroblock being encoded, and the DCT is performed on the matrix with the difference components.
[0045] The encoder features nine directions of intra-prediction. Temporal redundancy elimination methods have also been modified. P- and B-frames, when generated, can reference multiple frames (more than two). Huffman entropy coding has been replaced by a more complex and resource-intensive Context Adaptive Binary Arithmetic Coder (CABAC). Each new codec and each innovation is based on increased computer power.
[0046] Next, at step (202), the model (130) generates a video (121) according to the user request and transmits it to the video protection module (140). The video is generated in H.262 format or another format specified in paragraph
[0024] . These video formats encode frame images in JPG format [3-5] or another format that allows obtaining color pixels. When the generated video (121) is received by the module (140), 1-frames are extracted from it at step (203). Next, at step (203), blocks of pixels for the color component of the image are extracted from the video frames. The color component is at least one of: RGB, RGBA, HSL, HSLA, CMY, CMYK, XYZ, LMS, Hunt, ICtCp, LAB, RLAB, YCbCr, an image of frames in grayscale.
[0047] Fig. 3A shows an example of a video frame and obtaining pixel blocks associated with the pixels of a certain frame. A color channel is decoded from the frame. When receiving video at step (121), the frames of which are stored in JPEG format, decoding the discrete cosine transform (DCT) coefficients is performed to obtain the luminance (Y) and color (CLb (relative blueness) and Cr (relative redness) components of the image. RGB color channels are obtained from the YCbCr components. When receiving video at step (121), the frames of which are stored in BMP format, RGB color channels can be immediately extracted from the frames. At least one color channel is used. The channel pixels are divided into pixel blocks of a fixed size Wb x Hb and a sequence of pixel blocks (300) is formed.
[0048] In one embodiment of the invention, the original video is pre-converted to a fixed frame size (resolution) of W frame x H frame. This function is used to ensure that each generated video, regardless of its size (resolution), can be verified for authorship. In one signature destruction scenario, an attacker can change the video frame resolution. While the signature may be preserved in the frame pixels, extracting it requires converting the video frames to a standard resolution—the resolution at which the signature was embedded.
[0049] Next, at step (204), a new sequence of frame pixel blocks is formed by shuffling it using a given key (seed). The resulting shuffled sequence may coincide with the original. The shuffled sequence is divided into pairs of blocks. The sequence of block pairs is divided into two parts. The first sequence (301) has a fixed size. The second sequence (302) includes all the remaining pairs of blocks.
[0050] Fig. 3B shows an example of generating two sequences of block pairs from a video frame. The original block sequence (300) is shuffled using a given key (seed). The new sequence is divided into two parts. In this example, the first sequence (301) consists of four pairs of blocks (eight pixel blocks). The remaining pairs of blocks form sequence (302).
[0051] At step (205), the unique identifier of the generated video (unique value, identifier) is calculated based on the first sequence (301) of blocks. The generated video identifier allows the generated embedding to be linked to this specific video. Subsequent embedding cannot be transferred to another video, since it contains the unique identifier of this specific video. This identifier is a kind of hashing analog and can be used as a cryptographic hash of the video.
[0052] The video identifier is formed based on the first sequence (301) of blocks. The next pair of blocks is selected. The average pixel value of the first block S1 and the average pixel value of the second block S2 in the pair are calculated. If the difference between their average values is greater than the specified robustness coefficient R: S1-S2 > R, then the unique value bit is fixed at 1. If the difference between their average values is less than the specified robustness coefficient R with a negative sign: S1-S2 < -R, then the unique value bit is fixed at 0. If the difference between their average values lies in the range from -R to R, then a certain value is added / subtracted from all pixels of the above-mentioned blocks S1, S2 such that the absolute difference between their averages exceeds R, after which the corresponding unique value bit is fixed.
[0053] In one embodiment of the invention, a certain value is added / subtracted so that the absolute difference between the average values of pixels S1 and S2 exceeds R not for all block pixels, but only for some pixels. Using this feature, by excluding pixels located on the block boundary from the change, it is possible to reduce signature embedding artifacts visible to the human eye when examining the boundaries of two blocks. Thus, pixels in blocks located on the block boundary are not changed, and the transition from one block to another in the video frame will be smoother and less noticeable.
[0054] The use of the robustness coefficient R in the invention allows for the creation of high stability (indestructibility) in the signature of the generated video during video color correction, video resolution changes, video compression, and other transformations.
[0055] Fig. 4 shows an example of video identifier generation for the selected boundary value R=10. Each pair of blocks generates its own bit of information. The generated bits define a unique byte sequence (number), which is the video identifier. Within each block is a number that denotes the average pixel value within the block. For the first pair of blocks, we have Sl=55 and S2=49. Their absolute difference is |55-49| <R, поэтому необходимо изменить пиксели внутри блоков. Прибавление ко всем пикселям первого блока значения 2 позволяет получить новое среднее значение пикселей Sl=57. Вычитанием от всех пикселей второго блока значения 3 позволяет получить новое среднее значение пикселей S2=46. Теперь их абсолютная разница |57-46|> R. The difference value 57-46=11>R itself defines a bit equal to 1.
[0056] At step (206), a security attachment is generated. The attachment consists of a public (unencrypted) portion and an encrypted portion. The encrypted portion is an encrypted data sequence that may include the following data: the input time, the user ID, the text query ii (prompt request), the value of the video identifier obtained at step (205), and additional text information (e.g., the designation "I am a video generated by the Sber model"). The public portion may include information such as a verification digit, the value of the video identifier, and the integrity check digit of the encrypted attachment. An example of such an attachment is shown in Fig. 6.
[0057] In one particular embodiment of the invention, arbitrary data of arbitrary size, varying depending on the request or its conditions, is embedded. In this case, when extracting the data, its quantity and size are initially unknown. To enable extraction, the following data is inserted at the beginning of the encrypted portion: the number of different data items in a fixed number of first bits, and the size of each embedding in a fixed number of subsequent bits. Figure 6 shows an example of such an embedding. The amount of data is located at the beginning of the encrypted portion of the embedding. In this example, the embedding contains three types of data: - unique video identifier; - text information, including a fixed text phrase, user ID, request time, request prompt; - other data about the video. Examples of other video data include data such as: a thumbnail of a video frame as a compressed original frame; a history of prompt requests if there are multiple versions of the current video; the personal signature of the video creator; or other data that the video creator wishes to preserve in the signature, including in the form of graphic / audio or other content. The first 32 bits of the encrypted portion of the attachment are allocated to the amount of data. The data size follows: 64 bytes for the video identifier; 163 bytes for text information; and 376 bytes for other video data. Each size is allocated the next 32 bits of the encrypted portion of the attachment.
[0058] Additionally, the secure attachment may include other data, such as the first compressed frame of the generated video (121) or a portion thereof, which is converted into a byte sequence appended to the remaining data. The resulting attachment is signed with the private key at step (206) and formed part of the secure attachment at step (207). The resulting encrypted byte sequence is represented as a bit sequence.
[0059] Next, at step (208), the information obtained at step (207) is embedded into blocks of pixels of video frames.
[0060] First, the plaintext (unencrypted) information is embedded. This information may include a sequence of data such as a verification number, a video identifier (obtained in step (205)), or a security embedding integrity checksum, such as one calculated using CRC16, CRC32, CRC64, or any other checksum or hashing algorithm. This information is represented as a sequence of bits.
[0061] Next, the protective embedding obtained in step (207) is embedded. The embedding in step (208) is embedded into the pixels of the pairs of blocks obtained in step (204), which are included in the shuffled sequence (302).
[0062] An example of encoding the embedding of video frames into pixel blocks is shown in Fig. 5. Before this step, a shuffled sequence of pixel blocks (302) is formed (example in Fig. 3B). Fixed embedding parameters are selected: R is the robustness coefficient (stability of the embedding), and D is the distortion coefficient of video frames. In Fig. 5, the values R = 10 and D = 40 are given as an example.
[0063] As shown in the example in Fig. 5, the first bit of the embedded information is 1, the second is 1, the third is 0, and the fourth is 1. The first pair of blocks from the sequence (302) is taken. In Fig. 5, a number is written inside each block, which denotes the average value of the pixels within the block, and the lines inside the blocks depict the pattern of the original frame. The average value of all pixels within the first block S1 and the average value of all pixels within the second block S2 are calculated. In the example, the pixel blocks are the same and the average values Sl=94 and S2=94. First, the conditions |S1-S2| <D: - If |S1-S2| <D, то данная пара блоков пикселей используется для встраивания информации; - If |S1-S2|>=D, then this pair of pixel blocks is not used for embedding information, and a transition to the next pair of blocks is made.
[0064] Further, if the condition |S1-S2| is met <D, то выбирается очередной бит вкладываемой информации. Для его внедрения считается среднее значение пикселейпервого блока Si n среднее значение пикселей второго блока S2 в паре. Если требуется встроить бит 1, то ко всем пикселям блоков 1 и 2 прибавляется / отнимается некоторое значение так, чтобы разница их средних была больше заданного значения коэффициента робастности R (меньшего D): S1-S2 > R. If it is necessary to embed bit 0, then a certain value is added / subtracted from all pixels of blocks 1 and 2 so that the difference between their averages is less than the specified value of the robustness coefficient R with a negative sign: S1-S2 < -R.
[0065] Step by step, this process looks like this: - If it is required to embed bit 1 and Sl-S2<-R and the condition |S1-S2| is met <D, то ко всем пикселям блоков 1 и 2 прибавляется / отнимается некоторое значение так, чтобы разница их средних была больше заданного значения коэффициента робастности R: S1-S2 > R; - If it is necessary to embed bit 0 and Sl-S2<-R and the condition |S1-S2| is met <D, то блоки пикселей остаются без изменения; - If it is required to embed bit 1 and S1-S2>R and the condition |S1-S2| is met <D, то блоки пикселей остаются без изменения; - If it is required to embed bit 0 and S1-S2>R and the condition |S1-S2| is met <D, то ко всем пикселям блоков 1 и 2 прибавляется / отнимается некоторое значение так, чтобы разница их средних была больше заданного значения коэффициента робастности R: S1-S2 < -R. - Если условие |S1-S2|<D не выполнено, осуществляется переход к следующей паре блоков пикселей. This process continues until all required bits have been embedded or until there are no more pairs of pixel blocks in the sequence (302).
[0066] The video frame distortion coefficient D is also one of the distinctive features of the invention, enabling new levels of quality while improving information security. It eliminates blocks with large differences in average pixel values from the set of block pairs into which embedding occurs. If such block pairs are retained, embedding a bit of information into them may require a significant change in the block's pixels. These changes will be clearly visible to the human eye. Such artifacts will reduce the quality of the resulting video generation and also allow an attacker to identify which blocks the embedding occurred in and target them for destruction.
[0067] The robustness coefficient (embedding stability) R also allows for improved information security. It allows one to control the balance between the embedding's resistance to frame distortion and the appearance of embedding artifacts visible to the human eye. Increasing the robustness coefficient R results in a more reliable authorship embedding that will not be corrupted by significant video distortion. However, this may lead to the appearance of embedding artifacts. By decreasing the robustness coefficient R, one can achieve a virtually complete absence of artifacts visible to the human eye.
[0068] In the example in Fig. 5, the first bit of the embedded information is equal to 1. Bit 1 means that the difference S 1 - S2 must be greater than R, i.e. S 1 - S2>R. In the given example, S 1 =94 and S2=94, which means 94-94=0. To embed a bit of information equal to 1, the value 5 is added to all pixels of the first block and the value 6 is subtracted from all pixels of the second block. It turns out S 1 =99, S2=88, then 99-88=11>R. The bit of information equal to 1 is embedded. Then, a transition to the embedding of the next bit of information and a transition to the next pair of blocks in the sequence (302) is performed.
[0069] The second bit of embedded information is equal to 1. Bit 1 means that the difference S 1 - S2 must be greater than R, i.e. S 1 - S2>R. In this case, S 1 = 45 and S2 = 52, which means 45 - 52 = -7. To embed a bit of information equal to 1, the value 9 is added to all pixels of the first block and the value 9 is subtracted from all pixels of the second block. The result is Sl = 54, S2 = 43, then 54 - 43 = 11>R. The bit of information equal to 1 is embedded. Then the transition to the embedding of the next bit of information is performed.
[0070] The third bit of the embedded information is equal to 0. Bit 1 means that the difference S1-S2 must be less than -R, i.e. Sl-S2<-R. In this case, 155-110=45. This contradicts the condition |S1-S2| <D. Данная пара блоков пропускается, осуществляется переход к следующей паре блоков. Для следующей пары блоков разница пикселей 85-91 = -6. Для внедрения бита информации равный 0 происходит вычитание из всех пикселей первого блока значения 2 и прибавление ко всем пикселям второго блока значения 3. Получается S 1 =83, S2=94, тогда 83-94 = -11 < -R. Бит информации равный 0 внедрен.
[0071] The next message bit is set to 1. Bit 1 means that the difference between S1 and S2 must be greater than R, i.e., S1-S2>R. For the next pair of blocks, 87-76=11. The condition S1-S2>R is already satisfied. This pair of blocks remains unchanged. An information bit equal to 1 is considered embedded.
[0072] The final embedded security data sequence may look like this: [check digit (e.g. 537), video identifier, encrypted attachment integrity check digit, encrypted attachment]. The data in the encrypted attachment may be the following sequence: [number of attachments (e.g. 3), size of 1st attachment, size of 2nd attachment, size of 3rd attachment, video ID, text ("I am a video generated by Sber's model..."), additional data].
[0073] In byte form it will look like this: 00000219ac3eb891bc32652a83b4ff62cc2e49de3293d99e24ae8907e90c96b55cb5578a78d78 e887f5f436f6b9b989c6896b689e986s64a3509ee90fca..., where 00000219 is the byte format of the number 537, ac3eb891bc32652a83b4ff62cc2e49de32 is the byte format of the video identifier, 93d99e24 is the byte format of the CRC32 check digit value equal to 2480512548, and then the encrypted data. Fig. 6 shows an example of a byte sequence to be embedded. This sequence is represented as a bit sequence.
[0074] The resulting pixel blocks are arranged into the original sequence (300). The protected video (141) with the information embedded in step (208) is displayed on the user's device at step (209).
[0075] Next, we will consider the process of extracting embedded information when checking protected video (141).
[0076] Similar to steps (203) - (205), the value of the video identifier and the sequence of pixel blocks into which the information was potentially embedded are obtained.
[0077] The following algorithm is then executed.
[0078] Step 1. The first 4 bytes are extracted to obtain a check digit. This check digit is verified to be the original check digit. If the check digit matches, extraction continues. If not, a message is generated indicating that the digital signature was not found or was corrupted.
[0079] Step 2. The next 64 bytes are extracted. This is the unique video identifier. The resulting value is checked against the actual video identifier. If the value matches, extraction continues. If not, a message is displayed indicating that the digital signature was not found or was corrupted.
[0080] Step 3: The next 4 bytes are extracted, which allows us to obtain the checksum of the encrypted attachment.
[0081] Step 4. The encrypted attachment is extracted. All subsequent bits in the attachment refer to the encrypted attachment. At this stage, the size of the attachment is unknown. However, in the first block of the encrypted attachment—256 bytes for RSA—the lengths of the attachments are found. The first 256 bytes are extracted and decrypted using the public key. The first 4 bytes in this block are the number of attachments—in the example in Fig. 6, there are three attachments. Each subsequent 4 bytes is the size of each of these attachments. They contain the numbers 64, 163, and 376. If the size of any attachment is greater than 30,000 (the maximum attachment limit), the ciphertext block has been modified, causing a message to be generated indicating that the digital signature was not detected or was corrupted. Otherwise, the potential sizes of the attachments are obtained and extraction continues.
[0082] Next, the size of the encrypted attachment is determined. According to the example in Fig. 6, the obtained 64 bytes are the size of the first attachment, which is the value of the video identifier, then 163 bytes are the size of the second attachment - this is the text size. The size of the third attachment is 3756 bytes - this is the thumbnail of the first frame (the compressed first frame of the video). Then the total size of the encrypted attachment is: 64 + 163 + 3756 = 3983 (the attachment itself); 4 bytes are added for the number of attachments and 3 more times 4 bytes for their lengths. The result is as follows: 64 + 163 + 376 + 4 + 3 * 4 = 619 bytes. Since the encryption block has a size of 256, then in the given example 619 / / 256 = 3 entire encryption blocks, which is equal to 768 bytes of encrypted attachment.
[0083] 768 bytes are extracted and used to calculate the CRC32 checksum. If the value matches the checksum obtained in step 3, extraction continues. If not, a message is generated indicating that the digital signature was not found or was corrupted.
[0084] Decryption of 768 bytes. From the first block, as noted above, the number of attachments, their lengths, and the attachments themselves are obtained. The decrypted byte sequence is broken down by the number of bytes in each attachment. The original attachments are obtained.
[0085] The video ID value from the decrypted data is compared with the video ID value from the attachment (which was already compared with the actual value in step 2). If the values match, a message is displayed confirming the digital signature has been verified, and the extracted information—the video ID value, text, and additional data—is displayed. If not, a message is displayed indicating the digital signature was not detected or was corrupted.
[0086] Fig. 7 shows a general view of a computing device (500) suitable for performing the method (200). The device (500) may be, for example, a server or another type of computing device that can be used to implement the claimed technical solution, including: a smartphone, tablet, laptop, computer, etc. The device (500) may also be part of a cloud computing platform.
[0087] In general, the computing device (500) comprises one or more processors (501), memory means such as RAM (502) and ROM (503), input / output interfaces (504), input / output devices (505), and a device for network interaction (506), connected by a common information exchange bus.
[0088] The processor (501) (or several processors, multi-core processor) can be selected from a range of devices that are widely used at the present time, for example, from Intel™, AMD™, Apple™, Samsung Exynos™, MediaTEK™, Qualcomm Snapdragon™, etc. A graphics processor can also be used as a processor (401), for example, from Nvidia, AMD, Graphcore, etc.
[0089] RAM (502) is random access memory (RAM) and is designed to store machine-readable instructions executed by the processor (501) to perform the necessary logical data processing operations. RAM (502) typically contains executable instructions from the operating system and corresponding software components (applications, software modules, etc.).
[0090] ROM (503) represents one or more permanent data storage devices, such as a hard disk drive (HDD), a solid-state drive (SSD), a flash memory (EEPROM, NAND, etc.), optical storage media (CD-R / RW, DVD-R / RW, BlueRay Disc, MD), etc.
[0091] To organize the operation of the device components (500) and to organize the operation of external connected devices, various types of I / O interfaces (504) are used. The choice of the appropriate interfaces depends on the specific design of the computing device, which may include, but are not limited to: PCI, AGP, PS / 2, IrDa, FireWire, LPT, COM, SATA, IDE, Lightning, USB (2.0, 3.0, 3.1, micro, mini, type C), TRS / Audio jack (2.5, 3.5, 6.35), HDMI, DVI, VGA, Display Port, RJ45, RS232, etc.
[0092] To ensure user interaction with the computing device (500), various I / O information means (505) are used, for example, a keyboard, a display (monitor), a touch display, a touchpad, a joystick, a mouse, a light pen, a stylus, a touch panel, a trackball, speakers, a microphone, augmented reality means, optical sensors, a tablet, light indicators, a projector, a camera, biometric identification means (a retinal scanner, a fingerprint scanner, a voice recognition module), etc.
[0093] The network interaction means (506) ensures the transmission of data by the device (500) via an internal or external computer network, for example, an Intranet, the Internet, a LAN, etc. One or more means (506) may be, but are not limited to: an Ethernet card, a GSM modem, a GPRS modem, an LTE modem, a 5G modem, a satellite communication module, an NFC module, a Bluetooth and / or BLE module, a Wi-Fi module, etc.
[0094] Additionally, satellite navigation tools included in the device (500) can also be used, for example, GPS, GLONASS, BeiDou, Galileo.
[0095] The submitted application materials disclose preferred examples of the implementation of the technical solution and should not be interpreted as limiting other, particular examples of its implementation that do not go beyond the scope of the requested legal protection, which are obvious to specialists in the relevant field of technology. Sources of information: 1. GOST R 34.10-2012. Cryptographic information protection. Processes for the formation and verification of electronic digital signatures. https: / / rst. gov.ru: 8443 / file-service / file / load / 1699366979620 2. GOST P 34.11-2012. Cryptographic protection of information. Hash function. https: / / rst.gov.ru: 8443 / file-service / file / lo ad / 1699367414693 3. Hamilton, Eric: .JPEG File Interchange Format, Version 1.02. 1 September 1992; 4. Recommendation ITU-T T.871 : Information technology - Digital compression and coding of continuous-tone still images: .JPEG File Interchange Format (JFIF). Approved 14 May 2011; posted 11 September 2012; 5. Recommendation ITU-T T.81 : Information technology - Digital compression and coding of continuous-tone still images - Requirements and guidelines. Approved 18 September 1992; posted 14 April 2004.H.262 / MPEG-2 Part 2: «ISO / IEC 13818-2:2013 - Information technology - Generic coding of moving pictures and associated audio information: Video». ISO. https: / / www.iso.org / standard / 61152.html. Доступность: 24 July 2024. MPEG-4, Advanced Video Coding (Part 10) (H.264) (Full draft). Sustainability of Digital Formats. Washington, D.C.: Library of Congress. December 5, 2011. https: / / www.loc.gOv / preservation / digital / formats / fdd / fdd000081.shtml. Доступность: 24 July 2024.
Claims
FORMULA 1. A method and system for robust authorship protection of videos generated by a machine learning model, comprising the following steps: a) receive user input data to generate a video; b) generate a video file using a machine learning model; c) recording at least one type of user input data selected from the group: input time, user ID, or text query; d) extracting key 1-frames from the generated video file; e) extracting pixels of at least one color component from the received 1-frames; f) each color component is divided into blocks of pixels of fixed size Wb x Hb; g) using a given key, determine the mixing of 1-frames and their pixel blocks obtained in step e), form the first sequence of blocks from the first N blocks of the mixed sequence and the second sequence from all the remaining blocks; h) calculating a unique video identifier based on pairs of pixel blocks in 1-frames from the first sequence of blocks formed in step g), wherein for each pair of pixel blocks in 1-frames, the average value of the pixels in each block is calculated; i) forming an attachment for protecting the generated video, consisting of a sequence including at least one of: a unique video identifier formed in step h), the input time, the user ID, a text query and additional information; j) generate an encrypted sequence by encrypting the data obtained in step i) with a private key; k) embedding the encrypted sequence obtained in step j) in the form of a sequence of bits into pairs of pixel blocks from 1-frames of the generated video included in the second sequence of blocks determined in step g); l) provide protected video to the user.
2. The method according to claim 1, wherein at step d) P-frames and / or B-frames are additionally extracted.
3. The method according to claim 1, wherein the color component is selected from the group: color space, color model, color channel, subtractive model of frames of the generated video.
4. The method according to paragraph 3, in which the color component is at least one of: RGB, RGBA, HSL, HSLA, CMY, CMYK, XYZ, LMS, Hunt, ICtCp, LAB, RLAB, YCbCr, in grayscale gradations of the components of the frames of the generated video.
5. The method according to claim 1, in which the frame sizes of the generated video are adjusted to a given resolution.
6. The method according to claim 1, wherein a robustness coefficient R is determined for blocks of pixels in the first sequence.
7. The method according to claim 6, wherein based on R a value is determined for determining blocks for embedding bits of the sequence.
8. The method according to claim 1, in which error-correcting coding is applied to the generated embedding at step]).
9. The method according to claim 1, wherein step k) comprises embedding the value of the unique video identifier obtained in step h).
10. The method according to paragraph 1, in which error-correcting coding is applied to the generated encrypted sequence at step k).
11. The method according to claim 1, wherein error-correcting coding is applied to the unique video identifier calculated in step h).
12. The method according to claim 1, wherein step c) further comprises forming compressed images of one or more frames from the generated video.
13. The method according to claim 12, wherein a sequence of bytes is formed based on the compressed images of frames, which is added to the data in step i) when forming the attachment.
14. The method according to claim 1, wherein the video is generated according to a standard selected from the group: H.261, H.262, H.263, H.264, H.264+, H.265, H.265+, AVI, AVI, AVC1, VC-1, VP5, VP6, VP7, VP8, VP9, WebM, ASF, MCF.
15. A system for protecting the authenticity of videos generated based on a text query by a machine learning model, comprising at least one processor and at least one memory associated with the processor, storing machine-readable instructions that, when executed by the processor, implement the method according to any one of paragraphs 1-14.