Method and system for protecting the authenticity of machine learning model-generated video

By embedding a secure attachment into DCT coefficients of AI-generated videos, the method addresses the challenge of authenticating and protecting AI-generated videos from unauthorized modifications, ensuring their integrity and authorship verification.

WO2026071899A1PCT designated stage Publication Date: 2026-04-02PUBLICHNOE AKTSIONERNOE OBSHCHESTVO SBERBANK ROSSII (PAO SBERBANK)
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing methods for protecting AI-generated videos lack the ability to uniquely identify and verify the authenticity of the video, making them susceptible to unauthorized modifications and attribution issues.

Method used

A method involving extracting discrete cosine transform (DCT) coefficients from key frames of AI-generated videos, calculating a hash function, encrypting the data, and embedding it into the DCT coefficients of the frames to create a secure attachment that confirms the video's authenticity and immutability.

Benefits of technology

Ensures the authenticity and integrity of AI-generated videos by allowing verification of the authorship and detecting any modifications, maintaining the video's visual indistinguishability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure RU2024000369_02042026_PF_FP_ABST
    Figure RU2024000369_02042026_PF_FP_ABST
Patent Text Reader

Abstract

A method for protecting the authenticity of video generated by a machine learning model in response to a text query includes receiving user input data for the generation of a video, generating a video file with the aid of a machine learning model, registering the time of input, the user ID or the text query, extracting I-frames in JPG format from the generated video file, and generating an insertion for protecting the generated video, said insertion consisting of the aforesaid data. An encrypted sequence is generated by encryption of the insertion data, then hash function values and the encrypted sequence are embedded in discrete cosine transform coefficients of JPG frames from among previously determined I-frames of the generated video, and the protected video is presented to the user. The technical result is that of providing more efficient protection of generated videos.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR PROTECTING THE AUTHENTICITY OF VIDEOS GENERATED BY A MACHINE LEARNING MODEL AREA OF TECHNOLOGY

[0001] The claimed solution relates to the field of information technology, namely to means of protecting digital data for the purpose of confirming their authenticity. LEVEL OF TECHNOLOGY

[0002] With the widespread use of artificial intelligence (AI) and various machine learning models today, which can generate video files (and any other ordered sequence of images) based on a user's text query, there is a need to apply mechanisms to protect the generated videos for the purpose of subsequent confirmation of their authenticity.

[0003] One of the needs for this type of mechanism is to prevent attribution of AI-generated videos. Another concern is protecting videos from third-party modifications that could distort the original AI-generated video, which could later be used as false information discrediting the creators of such models for generating videos with prohibited or obscene content. Furthermore, the following may be susceptible to malicious modification: 1. Entirely video information 2. Separate sequences of frames or individual frames 3. The content of the frames themselves or the content of groups of frames

[0004] One example of video protection in the prior art is the use of a kind of digital signatures embedded in videos that are invisible to the human eye and can only be obtained by subsequent decryption of the protected videos using their algorithmic extraction (see the article "Using Steganographic and Cryptographic Tools to Protect Video Files" / / B. S. Markin, A. V. Chernyshova, 2014, URL: http.7 / ea.donntu.ru:8080 / bitstream / l 23456789 / 26611 / 1 / %D 1 %81 %D 1 %82%D0%B0%D 1 %82% Dl%8C%Dl%8F%201.pdf). According to the known solution, the signature is created in the spatial domain by slightly changing the intensity level of the image pixels. A well-known dissertation is "Model and Method for Protecting Video Information from Integrity Threats Using Steganography" / / R. Yu. Martimov. The drawback of these approaches is that the embedded signature cannot be uniquely identified with a given video—there is no video hash or equivalent. Under certain conditions, this signature can be extracted from one video file and transferred to another.

[0005] A method is known for crypto-processing any digital data, including existing digital video, using computer technology, as a result of which, for example, according to GOST R 34.10-2012 [1] and GOST R 34.11-2012 [2], video data is signed with a so-called digital electronic signature (DS).

[0006] A method for embedding and transmitting information in a video image is disclosed in patent RU 2608150 C2. The technical result is to minimize distortion of the embedded video image while ensuring the stego-resistance of the information transmission system. A method for covertly transmitting data in a video image according to the MPEG-2 (H.262) standard is proposed. This method is based on modifying the least significant bits of a video frame with the values ​​of a two-dimensional nonlinear code combination carrying the secretly transmitted information. The formation of a steganographic channel begins with processing the embedded data, including encryption and modulation with a pseudo-random signal, which is selected as two-dimensional nonlinear Frank-Walsh and / or Frank-Chrestenson signals. Simultaneously with the formation of the stego-signal, frames for its embedding are selected, assuming all I-frames, as well as B- and P-frames, are suitable.The stegochannel data is embedded only into those DCT coefficients that are located in the vicinity of the right diagonal of the matrix of DCT coefficients written to the JPEG file and supplemented with system information and universal Huffman tables by adding the bits of the DCT coefficients modulo two.

[0007] The disadvantage of this approach is that the information being embedded does not have the ability to verify authorship and that the embedded signature cannot be uniquely identified with the video. ESSENCE OF THE INVENTION

[0008] The present invention is aimed at solving the technical problem of creating a new, effective method for protecting digital videos with the possibility of subsequent confirmation of their authenticity and possible changes.

[0009] The technical result is to increase the efficiency of protecting AI-generated videos for the purpose of determining their authenticity.

[0010] An additional technical result is the ability to determine the immutability of video based on embedded security information. [OOP] In a preferred embodiment of the invention, a method is claimed for protecting the authenticity of videos generated based on a text query by a machine learning model, comprising the steps of: a) receiving user input data for generating a video; b) generating a video file using a machine learning model; c) recording at least one type of user input data selected from the group: input time, user ID, or a text query; d) extracting key 1-frames in JPG format from the generated video file; e) extracting from the obtained 1-frames blocks of values ​​of discrete cosine transform (DCT) coefficients for at least one image component selected from the group: luminance (Y), color (CL- and Cr) components of the image;f) obtain a sequence of low-frequency DCT coefficients (DC coefficients) for all 1-frames for all blocks of the luminance component, and calculate a hash function from the obtained sequence of DC coefficients; g) determine the 1-frames, blocks of DCT coefficients, and positions of the DCT coefficients within the blocks using a given key, excluding the DC DCT coefficients used in step f in the blocks; h) form an embedding for protecting the generated video, consisting of a sequence including: input time, user ID, text query, and the value of the hash function obtained in step f); i) generate an encrypted sequence by encrypting the data obtained in step h) with a private key; j) embed the values ​​of the hash function obtained in step f) and the encrypted sequence obtained in step i into the DCT coefficients of the frames in JPG format from the 1-frames of the generated video, determined in step g);k) provide protected video to the user.;

[0012] In one particular embodiment, at step d), P-frames and / or B-frames are additionally extracted.

[0013] In one particular example of implementation, at step d), blocks are extracted for embedding information of at least one of the Y-, Ь- and Сг-component frames of the generated video.

[0014] In another particular example of implementation, at step f), the sequence of DC coefficients consists of the DCT coefficients of one of the Y-, Cb- and Cr-component frames of the generated video.

[0015] In another particular example of implementation, at step f), the hash function is calculated based on at least one high-frequency DCT coefficient (AC coefficient) and / or DC coefficient of the DCT of at least one of the Y-, Cb- and Cr-component frames of the generated video.

[0016] In another particular example of implementation, at step g), blocks of DCT coefficients and positions of DCT coefficients within the blocks are determined using a given key, including any AC coefficients and DC coefficients of the DCT of at least one of the Y-, CL- and CR-component frames of the generated video, excluding those used at step e).

[0017] In another particular example of implementation, at step c), compressed images of one or more frames from the generated video are additionally formed.

[0018] In another particular example of implementation, a sequence of bytes is formed based on the compressed images of frames, which is added to the data at step g) when forming the embedding.

[0019] In another particular example of implementation, video is generated according to the H.262 or H.264 standard.

[0020] In another preferred embodiment, a system for protecting the authenticity of videos generated based on a text query by a machine learning model is claimed, comprising at least one processor and at least one memory associated with the processor, storing machine-readable instructions that, when executed by the processor, implement the above-mentioned method. BRIEF DESCRIPTION OF DRAWINGS

[0021] Fig. 1 illustrates a general view of the video protection method.

[0022] Fig. 2 illustrates a block diagram of the claimed method.

[0023] Fig. 3A illustrates an example of blocks with DCT coefficients associated with blocks of pixels of a certain frame from a video.

[0024] Fig. 3B illustrates the positions of the DCT coefficients in a block of a certain frame from a video.

[0025] Fig. ЗВ illustrates an example of the values ​​of the DCT coefficients at different positions in a block of a certain frame from a video.

[0026] Fig. 4 illustrates an example of determining the DCT coefficients in a certain frame from a video for embedding security information.

[0027] Fig. 5 illustrates an example of changing the DCT coefficients of video frames after embedding security information.

[0028] Fig. 6 illustrates an example of an attachment and its parts as a byte sequence.

[0029] Fig. 7 illustrates a view of the computing device. IMPLEMENTATION OF THE INVENTION

[0030] Fig. 1 shows a general diagram of the claimed solution, which is implemented by processing incoming user requests (110), which represent a textual description of the generated video. In the field of AI, such textual requests, describing the generated object for the machine learning model, are also called "prompts." The user's request (110) can be transmitted via an API (120) implemented in a web browser or messenger (e.g., Telegram), which subsequently enables data exchange directly with the video generation model (130) (e.g., Kandinsky) between the user (110) and the video protection module (140).

[0031] The video protection module (140) can be implemented using an external hardware and software system, such as a server or remote workstation, providing functionality for processing incoming videos (121) generated by the model (130). The module (140) can also be located in the same execution loop as the model (130), such as a server or cloud computing node.

[0032] Upon completion of the video protection module (140), security information containing an encrypted attachment is embedded into the initially generated videos (121). This information allows for at least the video's authorship to be established as being created using the model (130), as well as the user ID of the user who generated the request and the request text itself. The protected video (141) is subsequently transmitted to the user (software) via the API (120), for example, by displaying / playing the video (141) in a web browser or instant messenger.

[0033] Data exchange within the framework of the system implementation is based on standard principles of data exchange, in particular with the help of the Internet computing network, which can be implemented using any known principles known from the prior art.

[0034] Fig. 2 shows an example of implementing the claimed method (200) for protecting generated videos. In the first step (201), user input is received for video generation by the model (130). The user input includes a user text query, a user ID (e.g., IP address, MAC address, registration ID, etc.), and the time of the text query. This data is recorded by the video protection module (140) for subsequent generation of security information for embedding in the video.

[0035] The H.262 video codec [6] is part of the MPEG-2 Part 2 family of digital video and audio coding standards. Initially, this codec could encode video with a resolution of 720*480 pixels at 30 frames per second; with further improvements, compression of a resolution of 1920*1080 pixels at 30 frames per second was achieved. Compression is based on the elimination of spatial and temporal redundancies.

[0036] To address spatial and temporal redundancy in a video stream, the codec introduces the concept of a group of pictures (GOP). Frames in a GOP are divided into three types: basic I- (intrapictures), predicted P- (predicted), and bidirectional B- (bidirectional) frames. I-frames are the input to the decoder and are encoded independently according to the method used in the JPEG standard. P-frames are encoded based on prediction, by referencing blocks of previous I- or P-frames. B-frames use references to two frames located before and after them. A GOP must always begin with an I-frame; P- and B-frames are predicted from it.

[0037] The number of P- and B-frames can vary depending on the desired degree of compression: the more P- and B-frames, the greater the compression, but if the 1-frame is encoded incorrectly, or information is lost in the I-frame, the error will extend to the entire group of frames.

[0038] Video image compression occurs as follows. The frame sequence is divided into 16x16 sample macroblocks as in the JPEG algorithm and divided into three frame types: I-, P-, and B-frames. I-frames are compressed using the JPEG algorithm: in each 16x16 sample macroblock, a transition to the YCrCb color system is performed, the Cr and Cb components are thinned to 8x8 matrices (every second row and every second column are removed), a discrete cosine transform (DCT) of the matrices is performed and they are quantized, and finally, zig-zag scanning and compression of the resulting sample sequence using the Huffman method. P-frames no longer transmit a fully encoded image, but a prediction error taking into account motion compensation - the difference between the previous and next frame with motion vectors.

[0039] Motion vectors are calculated as follows: each block from the previous frame is compared with the block from the next. If they are identical, no motion has occurred, and there is no motion vector. If not, the block is moved within a search region of the main image to find its position at which the root-mean-square difference between the block and the main image fragment is minimal. This vertical and horizontal displacement of the block relative to its original position is taken as the motion vector. When encoding, B-frames reference two frames: one in front of the frame and one behind it. This means they convey information about how to "assemble" the frame from the two surrounding frames. This is because each frame is transmitted as a reference to the main frames, rather than individually. The MPEG2 standard is well suited for storing movies, but due to its group-of-frame structure, it is completely unsuitable for video editing.Problems also arise when compressing videos with fast-moving objects or frequently changing shots. The object will be blurred, and if the 1-frame is not captured at the moment of the change in shot, all the information will be written to the P-frame, resulting in an increased file size.

[0040] The H.264 video codec [7] is part of the MPEG4 Part 10 group of standards. This recommendation is currently the standard for video compression. It is used to record high-definition video on Blu-ray and HD DVD, and is the standard for online video hosting services such as YouTube, as well as a standard in digital television broadcasting systems. Compared to H.262, this codec provides twice the compression with the same video image quality. This is achieved due to the significant complexity of the codec. This codec implements division of the frame not only into 16x16 macroblocks, but also 16x8, 8x16, 8x8, 8x4, 4x8, 4x4, depending on the presence of small details. This increases the clarity of the transmission of small objects and the quality of motion compensation: greater accuracy of the representation of motion vectors is ensured. Moreover, the accuracy of the motion vector search is 1 / 4 or 1 / 8 of a macroblock, which was not available in the H.262 codec. Changed internal coding of 1-frame.The encoding is also based on the JPEG algorithm, but with some additions. Before the DCT operation, a macroblock undergoes spatial prediction (intraprediction) based on the adjacent macroblocks that have already been encoded. The luminance and color difference samples of the predicted macroblock are then subtracted from the corresponding samples of the macroblock being encoded, and the DCT is performed on the matrix containing the difference components.

[0041] The encoder features nine directions of intra-prediction. Temporal redundancy elimination methods have also been modified. P- and B-frames, when generated, can reference multiple frames (more than two). Huffman entropy coding has been replaced by a more complex and resource-intensive Context Adaptive Binary Arithmetic Coder (CABAC). Each new codec and each innovation is based on increased computer power.

[0042] Next, at step (202), the model (130) generates a video (121) according to the user request and transmits it to the video protection module (140). The video is generated in the H.262 or H.264 format. These video formats encode frame images in the JPG (or JPEG) format [3-5]. When the generated video (121) in the H.262 or H.264 format is received by the module (140), I-frames are extracted from it. These frames represent an image in the JPG format. At step (203), the values ​​of the discrete cosine transform (DCT) coefficients for the luminance (Y) and color (CL (relative blueness) and CR (relative redness) components of the image are extracted from the JPG image of the video block frames.

[0043] Fig. 3A shows an example of a video frame and the derivation of blocks with discrete cosine transform (DCT) coefficients. When encoding the original frame (an image represented as RGB pixels) into JPEG, these coefficients are obtained as follows. The resulting image is divided into color channels. The resulting image channels (121) are divided into blocks (3001 - 300n) with dimensions of 8x8 pixels. The number of blocks (3001 - 300 п ) depends on the resolution of the generated image (121) or the size of the raster image. For example, for an image measuring 1920x1080 pixels, the number of blocks will be 240 horizontally and 135 vertically.

[0044] Each block (3001 - 300 п ) is subjected to a discrete cosine transform (DCT), which is a type of discrete Fourier transform. In this case, each of the blocks (3001 - 300 п) contains one DC coefficient (3011) at the position (0,0) within the block and 63 AC coefficients (30164) at other positions within the block. The DC coefficient is the average value of all values ​​within the block, as roughly shown in Fig. ЗВ.

[0045] Fig. 3B shows an example of positioning DCT coefficients within a block for video frames. DC coefficients are also called low-pass coefficients. In the current description, the position within a block (0,0) is also denoted as 1- th position. Positions (Y1, Y2), Y1 = 0.7, Y2 = 0.7, (Y1, Y2) (0.0) are designated from the 2nd to the 64th position.

[0046] The JPEG frame (image) itself is a compressed storage of the aforementioned DCT coefficients. Thus, at step (203), these blocks of DCT coefficients are extracted directly from the JPEG-encoded image data.

[0047] At step (204), the DC coefficient sequence is hashed. This is done by taking all DC coefficient values ​​in blocks across all I-frames and applying the selected hashing algorithm, such as SHA-256, SHA-512, etc.

[0048] As shown in Fig. 4, at step (205), using the selected key (seed), 1-frames are determined and within them the AC coefficients in blocks (3001 - 300 п ) for subsequent embedding of protective information. The initial ordered sequence (401) of frames (F) from the generated video is taken, within the frames there are blocks of DCT coefficients (X1, X2) and the positions of the coefficients themselves within the blocks (Y1, Y2): {E - (X1, X2) - (Y1, Y2)}. Where F is the frame number in the video, (X1, X2) is the coordinate of the block of DCT coefficients within the image: X1 = 0, H изо b ра zheniya — 1 , X2 = images; (Y1, Y2) - coordinates of the positions of the DCT coefficients inside the blocks: Y1 = 0.7, Y2 = 0.7, (Y1, Y2) Φ (0.0). Fig. 3B shows an example of such positions. Next, the selected sequence (401) of AC coefficients is mixed using a given key, subsequently forming a new sequence (402) of AC coefficients.

[0049] At step (206), a security attachment is generated. The attachment consists of a public (unencrypted) part and an encrypted part. The encrypted part is an encrypted data sequence, including the following data: the entry time, the user ID, the text query (prompt request), the hash value obtained at step (204), and additional text information (e.g., the designation "I am a Sber model"). The public part may include information such as a verification digit, a hash sequence of the selected DCT coefficients, and a check digit of the integrity of the encrypted attachment. An example of such an attachment is shown in Fig. 6.

[0050] In one of the particular variants of the invention, arbitrary data of arbitrary size, changing depending on the request or its conditions. In this case, when extracting data, their quantity and size are initially unknown. To enable their extraction, the following data is inserted at the beginning of the encrypted part of the data: the quantity of different data in a fixed number of the first bits, the size of each of the attachments in a fixed number of subsequent bits. An example of such an attachment is shown in Fig. 6. At the beginning of the encrypted part of the attachment, there is the quantity of data - 3 data: the video hash; text information, including a fixed text phrase, user ID, request time, request "Prompt"; compressed video frame. The first 32 bits of the encrypted part of the attachment are allocated for the quantity of data. Next come the sizes of the data - 64 bytes for the hash; 163 bytes for the text information, 3756 bytes for the compressed video frame. The next 32 bits of the encrypted part of the attachment are allocated for each size.

[0051] Additionally, the secure embedding may include a compressed image of a certain frame (e.g., the first) from the generated video (121), which is converted into a byte sequence appended to the byte sequence of the remaining data. The resulting embedding is signed with a private key at step (207) in step (206), thereby forming part of the secure embedding. The resulting encrypted byte sequence is represented as a bit sequence.

[0052] Next, at step (208), information is embedded into the sequence of DCT coefficients from 1-frames obtained at step (205).

[0053] First, the plaintext (unencrypted) information is embedded. This information may include a data sequence such as a check digit, a hash of the DC coefficient sequence (obtained in step (204)), or a security embedding integrity check digit, such as one calculated using CRC16, CRC32, CRC64, or any other checksum or hashing algorithm. This information is represented as a bit sequence.

[0054] Next, the secure attachment obtained at step (207) is embedded. The protected video (141) with the information embedded at step (208) is displayed / played at step (209) on the user's device.

[0055] The embedding at step (208) is built directly into the AC block coefficients (3001 - 300) selected at step (205) п ), included in the shuffled sequence (402).

[0056] An example of encoding an embedding inside a video file is shown in Fig. 5. When this step is performed, a mixed sequence of video frames, blocks of DCT coefficients, and the positions of the coefficients themselves within the blocks is formed. In Fig. 5 io The first bit is embedded in frame #14, block (31, 70), into the coefficient at position (1, 0). The second bit is embedded in frame #21, block (8, 69), into the coefficient at position (0, 2), and so on. Embedding is performed as the remainder of dividing the DCT coefficient value by 2. In other words, if the remainder of dividing by 2 is 0, this means that a bit with the value 0 is embedded in the coefficient. If the remainder of dividing by 2 is 1, this means that a bit with the value 1 is embedded in the coefficient. If the current value of the DCT coefficient does not produce the required remainder, this value is increased or decreased by 1.

[0057] As shown in the example in Fig. 5, the first bit of the embedded information is 1. The value of the coefficient in frame #14, in block (31, 70), at position (1, 0) is 17. The remainder of dividing 17 by 2 is 1. The remainder coincides with the value of the embedded bit. The second bit of the embedded information is 0. The value of the coefficient in frame #21, in block (8, 69), at position (0, 2) is 23. The remainder of dividing 23 by 2 is 1. The remainder does not coincide with the value of the embedded bit. We change the value of 23 by subtracting 1 from it. The value becomes equal to 22. The remainder of dividing 22 by 2 is now 0, which coincides with the value of the embedded bit.

[0058] Information embedding can be achieved by encoding information into the least significant bit (LSB) in blocks of DCT coefficients. The LSB is the last bit in the bitmap whose change least significantly alters the value itself.

[0059] An example of coding in the DCT block of coefficients in Fig. ЗВ: 15 - 00010101. DC coefficient excluded from embedding. 4 - 00000100. NZB=0. Requires embedding 0 -> no change required 6 - 00000110. NZB=0. Requires embedding 1 - change required 3- 00000011 (change). NZB=1. Requires embedding 0 — ► change required

[0060] Let's look at an example of embedding text information. The text to be embedded is: SBER - 01010011 01000010 01000101 01010010. Let's consider embedding the letter "S" - 01010011. The embedding proceeds from left to right: 0-1 -0-1, etc. - Step 1. The first bit from "S" = 0 and is embedded in the DCT coefficient at position (0; 1) with the value 4. The LSB for 4 is 0 (00000100). It is necessary to embed the 0 from "S" and for 4 the LSB is also 0. Therefore, the value does not change. - Step 2. The second bit from "S" = 1 and it is embedded in the DCT coefficient with the value 6. The LSB for 6 is 0 (00000110). It is necessary to embed 1 from "S", and for 6 the LSB = 0. li Therefore, the value changes to 1 and takes the bit form 00000111. In other words, 6 is replaced by 7. - Step 3. The third bit of "S" is 0 and is embedded in the DCT coefficient with the value 3. The LSB for 3 is 1 (00000011). We need to embed the 0 from "S", and for 3 the LSB = 1. Therefore, this value is changed to 0 and takes the bit form 00000010. In other words, 3 is replaced by 2. This process continues until all required bits are embedded. Ultimately, a DCT matrix of coefficients is formed with encoded information regarding the change in the selected AS coefficients.

[0061] The final embedded sequence of security data may have the following form: [check digit (e.g., 537), DC coefficient hash, encrypted attachment integrity check digit, encrypted attachment]. The data in the encrypted attachment may be the following sequence: [number of attachments (e.g., 3), size of 1st attachment, size of 2nd attachment, size of 3rd attachment, DC coefficient hash, text ("I am a Sber model..."), video frame thumbnail (highly compressed image of a frame from the video)].

[0062] In byte form it will look like this: 00000219ac3eb891bc32652a83b4ff62cc2e49de3293d99e24ae8907e90c96b55cb5578a78d78e88 7f5f436f6b9b989c6896b689e986s64a3509ee90fca..., where 00000219 is the byte form of the number 537, ac3eb891bc32652a83b4ff62cc2e49de32 is the byte form of the hash, 93d99e24 is the byte form of the CRC32 check digit value equal to 2480512548, then comes the encrypted data. Figure 6 shows an example of a byte sequence to be embedded. This sequence is represented as a bit sequence.

[0063] The obtained new values ​​of the DCT coefficients (Fig. 5) are compressed (encoded) according to the JPEG standard, form a new 1-frame video, and are saved as a new video file.

[0064] When embedding this sequence into selected frames and DCT coefficients, the resulting video (141) remains visually indistinguishable from the originally generated video (121). Even small changes to the pixels of any frame lead to changes in the DCT coefficients of that frame, which destroys the embedded embedding. Any rearrangement of the original frames also destroys the embedded embedding. This fact allows for subsequent integrity verification of the embedded security information to confirm that the video (141) is authentic, i.e., created using the model (130), and that it has not been modified in any part of its content.

[0065] Next, we will consider the process of extracting embedded information when checking protected video (141).

[0066] Similar to steps (203) - (205), the hash value of the DC coefficients of video frames (video hash) and the sequence of frames and DCT coefficients into which information embedding was potentially performed are obtained.

[0067] The following algorithm is then executed.

[0068] Step 1. The first 4 bytes are extracted to obtain a check digit. This check digit is verified to be the original check digit. If the check digit matches, extraction continues. If not, a message is generated indicating that the digital signature was not found or was corrupted.

[0069] Step 2. The next 64 bytes are extracted. This is the hash of the video frame DC coefficients. The resulting value is checked against the actual DC coefficient hash. If the hash matches, extraction continues. If not, a message is displayed indicating that the digital signature was not found or was corrupted.

[0070] Step 3: The next 4 bytes are extracted, which allows us to obtain the checksum of the encrypted attachment.

[0071] Step 4. The encrypted attachment is extracted. All subsequent bits in the attachment refer to the encrypted attachment. At this stage, the size of the attachment is unknown. However, in the first block of the encrypted attachment—256 bytes for RSA—the lengths of the attachments are found. The first 256 bytes are extracted and decrypted using the public key. The first 4 bytes in this block are the number of attachments—in the example in Fig. 6, there are 3 attachments. Each subsequent 4 bytes is the size of each of these attachments. They contain the numbers 64, 163, and 3756. If the size of any attachment is greater than 30,000 (the maximum attachment limit), the ciphertext block has been modified, and a message is generated indicating that the digital signature was not detected or was corrupted. Otherwise, the potential sizes of the attachments are obtained and extraction continues.

[0072] Next, the size of the encrypted attachment is determined. In accordance with the example in Fig. 6, the obtained 64 bytes are the size of the first attachment, which is the hash of the DC coefficients of the video frames, then 163 bytes are the size of the second attachment - this is the text size. The size of the third attachment is 3756 bytes - this is a thumbnail of the first video frame (compressed frame). Then the total size of the encrypted attachment is: 64 + 163 + 3756 = 3983 (the attachment itself); 4 bytes are added for the number of attachments and 3 more times 4 bytes for their lengths. The result is as follows: 64 + 163 + 3756 + 4 + 3 * 4 = 3999 bytes. Since the encryption block has a size of 256, then in the given example 3999 / / 256 = 16 whole encryption blocks, which is equal to 4096 bytes of encrypted attachment.

[0073] 4096 bytes are extracted and used to calculate the CRC32 checksum. If the value matches the checksum obtained in step 3, extraction continues. If not, a message is generated indicating that the digital signature was not found or was corrupted.

[0074] Decryption of 4096 bytes. From the first block, as noted above, the number of attachments, their lengths, and the attachments themselves are obtained. The decrypted byte sequence is broken down by the number of bytes in each attachment. The original attachments are obtained.

[0075] The hash from the decrypted data is verified against the hash from the attachment (which was already compared with the actual hash in step 2). If the hashes match, a message is displayed confirming the digital signature has been verified, along with the extracted information—the hash of the video's DC coefficients, the text, and a thumbnail of the first frame of the generated video. If not, a message is displayed indicating the digital signature was not detected or was corrupted.

[0076] Fig. 7 shows a general view of a computing device (500) suitable for performing the method (200). The device (500) may be, for example, a server or another type of computing device that can be used to implement the claimed technical solution, including: a smartphone, tablet, laptop, computer, etc. The device (500) may also be part of a cloud computing platform.

[0077] In general, the computing device (500) comprises one or more processors (501), memory means such as RAM (502) and ROM (503), input / output interfaces (504), input / output devices (505), and a device for network interaction (506), connected by a common information exchange bus.

[0078] The processor (501) (or several processors, multi-core processor) can be selected from a range of devices that are widely used at the present time, for example, from Intel™, AMD™, Apple™, Samsung Exynos™, MediaTEK™, Qualcomm Snapdragon™, etc. A graphics processor can also be used as a processor (401), for example, from Nvidia, AMD, Graphcore, etc.

[0079] RAM (502) is random access memory (RAM) and is designed to store machine-readable instructions executed by the processor (501) to perform the necessary logical data processing operations. RAM (502) typically contains executable instructions from the operating system and corresponding software components (applications, software modules, etc.).

[0080] ROM (503) represents one or more permanent data storage devices, such as a hard disk drive (HDD), a solid-state drive (SSD), flash memory (EEPROM, NAND, etc.), optical storage media (CD-R / RW, DVD-R / RW, BlueRay Disc, MD), etc.

[0081] To organize the operation of the device components (500) and to organize the operation of external connected devices, various types of I / O interfaces (504) are used. The choice of the appropriate interfaces depends on the specific design of the computing device, which may include, but are not limited to: PCI, AGP, PS / 2, IrDa, FireWire, LPT, COM, SATA, IDE, Lightning, USB (2.0, 3.0, 3.1, micro, mini, type C), TRS / Audio jack (2.5, 3.5, 6.35), HDMI, DVI, VGA, Display Port, RJ45, RS232, etc.

[0082] To ensure user interaction with the computing device (500), various I / O information means (505) are used, for example, a keyboard, a display (monitor), a touch display, a touchpad, a joystick, a mouse, a light pen, a stylus, a touch panel, a trackball, speakers, a microphone, augmented reality means, optical sensors, a tablet, light indicators, a projector, a camera, biometric identification means (a retinal scanner, a fingerprint scanner, a voice recognition module), etc.

[0083] The network interaction means (506) ensures the transmission of data by the device (500) via an internal or external computer network, for example, an Intranet, the Internet, a LAN, etc. One or more means (506) may be, but are not limited to: an Ethernet card, a GSM modem, a GPRS modem, an LTE modem, a 5G modem, a satellite communication module, an NFC module, a Bluetooth and / or BLE module, a Wi-Fi module, etc.

[0084] Additionally, satellite navigation tools included in the device (500) can also be used, for example, GPS, GLONASS, BeiDou, Galileo.

[0085] The submitted application materials disclose preferred examples of the implementation of the technical solution and should not be interpreted as limiting other, particular examples of its implementation that do not go beyond the scope of the requested legal protection, which are obvious to specialists in the relevant field of technology. Sources of information: 1. GOST R 34.10-2012. Cryptographic information protection. Processes for the formation and verification of electronic digital signatures. https: / / rst.gov.ru:8443 / file-service / file / load / l 699366979620 2. GOST R 34.11-2012. Cryptographic information protection. Hash function. https: / / rst.gov.ru:8443 / file-service / file / load / l 699367414693 3. Hamilton, Eric: JPEG File Interchange Format, Version 1.02. 1 September 1992; 4. Recommendation ITU-T T.871 : Information technology - Digital compression and coding of continuous-tone still images: JPEG File Interchange Format (JFIF). Approved 14 May 2011; posted 11 September 2012; 5. Recommendation ITU-T T.81 : Information technology - Digital compression and coding of continuous-tone still images - Requirements and guidelines. Approved 18 September 1992; posted 14 April 2004. 6. H.262 / MPEG-2 Part 2: «ISO / IEC 13818-2:2013 - Information technology - Generic coding of moving pictures and associated audio information: Video». ISO. htt s: / / www.iso.org / standard / 61152.html. Доступность: 24 July 2024. 7. MPEG-4, Advanced Video Coding (Part 10) (H.264) (Full draft). Sustainability of Digital Formats. Washington, D.C.: Library of Congress. December 5, 2011. https: / / www.loc.gov / preservation / digital / formats / fdd / fdd000081.shtml. Доступность: 24 July 2024.

Claims

FORMULA 1. A method for protecting the authenticity of videos generated based on a text query by a machine learning model, comprising the steps of: a) receiving user input data for generating a video; b) generating a video file using a machine learning model; c) recording at least one type of user input data selected from the group: input time, user ID, or a text query; d) extracting key 1-frames in JPG format from the generated video file; e) extracting from the obtained 1-frames blocks of discrete cosine transform (DCT) coefficient values ​​for at least one image component selected from the group: luminance (Y), color (CL- and Cr) image components; f) obtaining a sequence of low-frequency DCT coefficients (DC- coefficients) for all 1-frames for all blocks of the luminance component, and calculating a hash function from the obtained sequence of DC- coefficients;g) using a given key, determine 1-frames, blocks of DCT coefficients, and positions of DCT coefficients within blocks, excluding in the blocks the DC DCT coefficients used in step i); h) form an embedding for protecting the generated video, consisting of a sequence including: input time, user ID, text query, value of the hash function obtained in step f); i) generate an encrypted sequence by encrypting the data obtained in step h) with a private key; j) embed the hash function values ​​obtained in step f) and the encrypted sequence obtained in step i) into the DCT coefficients of frames in JPG format from the 1-frames of the generated video, determined in step g); k) provide the protected video to the user.

2. The method according to claim 1, wherein the video is generated according to the H.262 or H.264 standard.

3. The method according to claim 1, wherein in step d) P-frames and / or B-frames are additionally extracted.

4. The method according to claim 1, wherein step d) comprises extracting blocks for embedding information of at least one of the Y-, Ь- and Сг- constituent frames of the generated video.

5. The method according to claim 1, wherein in step f) the sequence of DC coefficients consists of the DCT coefficients of one of the Y-, Cb- and Cr-component frames of the generated video.

6. The method according to claim 1, wherein in step f) the hash function is calculated based on at least one high-frequency AC coefficient of the DCT and / or DC coefficient of the DCT of at least one of the Y-, Cb- and Cr-component frames of the generated video.

7. The method according to claim 1, wherein at step g), using a given key, blocks of DCT coefficients and positions of DCT coefficients within the blocks are determined, including AC coefficients and DC coefficients of the DCT, of at least one of the Y-, Cb- and Cr-component frames of the generated video, excluding those used at step 1).

8. The method according to claim 1, wherein step c) further comprises forming compressed images of one or more frames from the generated video.

9. The method according to claim 8, wherein a sequence of bytes is formed based on the compressed images of the frames, which is added to the data at step b) when forming the attachment.

10. A system for protecting the authenticity of videos generated based on a text query by a machine learning model, comprising at least one processor and at least one memory associated with the processor, storing machine-readable instructions that, when executed by the processor, implement the method according to any one of paragraphs 1-9.

Citation Information

Patent Citations

  • Method of hidden data transfer in video image

    RU2608150C2

  • Method and device for generating video clip from text description and sequence of key points synthesized by diffusion model

    RU2823216C1

  • Immutable watermarking for authenticating and verifying ai-generated output

    US20240242128A1

  • Generating videos using sequences of generative neural networks

    US20240320965A1

  • Variable length video generation from textual descriptions

    WO2024072999A1