Method and system for protecting the authenticity of machine learning model-generated images

By employing pixel block shuffling and encryption, the method strengthens the protection of AI-generated images against modifications, ensuring robust authenticity confirmation through unique identifiers.

WO2026054670A1PCT designated stage Publication Date: 2026-03-12PUBLICHNOE AKTSIONERNOE OBSHCHESTVO SBERBANK ROSSII (PAO SBERBANK)
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing methods for protecting AI-generated images lack robustness against modifications and do not provide efficient means to confirm authenticity, as they are vulnerable to compression, distortion, and do not contain unique identifiers.

Method used

A method involving pixel block shuffling, robustness coefficient calculation, and encryption of image identifiers is used to embed security information into AI-generated images, ensuring authenticity and immutability.

Benefits of technology

The method enhances the protection of AI-generated images by providing robustness against image distortions and enabling efficient confirmation of authenticity through unique identifiers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure RU2024000371_12032026_PF_FP_ABST
    Figure RU2024000371_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to the field of information technology, and more particularly to digital data protection means for confirming the authenticity of digital data. A method for protecting the authenticity of images generated by a machine learning model in response to a text query comprises the steps of: receiving user input data for the generation of an image; generating an image with the aid of a machine learning model; registering the type of user input data; extracting from the image pixels of at least one colour component thereof; separating each colour component into blocks of pixels with a set size; forming a first sequence of blocks; calculating an identifier of the image; generating an insertion for protecting the generated image; generating an encrypted sequence; generating a sequence to be embedded; embedding the sequence; and presenting the protected image to the user. The technical result is that of providing more efficient protection of AI-generated images in order to determine the authenticity thereof.
Need to check novelty before this filing date? Find Prior Art

Description

A METHOD AND SYSTEM FOR ROBUST PROTECTION OF AUTHORITY OF IMAGES GENERATED BY A MACHINE LEARNING MODEL AREA OF TECHNOLOGY

[0001] The claimed solution relates to the field of information technology, namely to means of protecting digital data for the purpose of confirming their authenticity. LEVEL OF TECHNOLOGY

[0002] With the widespread use of artificial intelligence (AI) and various machine learning models that can generate images based on user text queries, there is a need to apply mechanisms to protect the generated images for the purpose of subsequent confirmation of their authenticity.

[0003] One of the needs for this type of mechanism is to prevent attribution of AI-generated images. Another concern is protecting images from third-party modifications that could distort the original AI-generated image, which could subsequently be used as false information discrediting the creators of such models for their ability to generate images of prohibited or obscene content.

[0004] One example of image protection in the prior art involves the use of digital signatures embedded in generated images. These signatures are invisible to the human eye and can only be obtained by subsequent decryption of the protected images using algorithmic extraction (see the article "Copyright Protection of Images Using Robust Digital Signatures" / / N. Nikolaidis et al., 1996). According to this prior art solution, the signature is generated in the spatial domain by slightly varying the intensity level of randomly selected image pixels. The signature is determined by comparing the average intensity of the marked pixels with that of the unmarked pixels. Statistical hypothesis testing is used for this purpose. The signature can be constructed in such a way that it is resistant to JPEG compression and low-pass filtering.

[0005] This approach was the basis for the declared solution in terms of the use of the implementation of protective information in the brightness coefficients of the JPEG format, as well as a number of improvements that allow not only more efficient labeling of generated images, but also detection of possible future changes.

[0006] A method is known for cryptographic processing of any digital data, including an existing digital image, using computer technology, as a result of which, for example, according to GOST R 34.10-2012 [4] and GOST R 34.11-2012 [5], image data is signed with a so-called digital electronic signature (ES).

[0007] A method for confirming photograph authorship is described in patent RU2633185C2 20171011. The drawback of this approach is that electronic signature information is stored in service fields of various image file formats. Information stored in this way is always detectable in the file and can be extracted or removed from the file without altering the image itself.

[0008] A method for hiding data in an image is known in the publications of W. Bender, D. Grahl, N. Morimoto, A. Lu, "Techniques for Data Hiding" IBM Systems Journal, Vol. 35 Nos. 3&4, 1996 and Daniel Gruhl and Walter Bender, "Information Hiding to Foil the Casual Counterfeiter", 2nd Information Hiding Workshop, 1998. The disadvantage of these approaches is that these methods have low resistance to compression / stretching / distortion of the image, do not contain the ability to control the stability and size of the embedding through some robustness coefficient, do not contain the receipt and implementation of some unique value (identifier) ​​of the image. ESSENCE OF THE INVENTION

[0009] The present invention is aimed at solving the technical problem of creating a new, effective method for protecting digital images with the possibility of subsequent confirmation of their authenticity and possible changes.

[0010] The technical result is to increase the efficiency of protecting AI-generated images for the purpose of determining their authenticity.

[0011] An additional technical result is the ability to determine the immutability of images based on embedded security information.

[0012] In a preferred embodiment of the invention, a method is claimed for protecting the authenticity of images generated based on a text query by a machine learning model, comprising the steps of: a) receiving user input data for generating an image containing a text query; b) generating an image using a machine learning model; PC17RU2024 / 000371 c) recording at least one type of user input data selected from the group: input time, user ID, text query; d) extracting pixels of at least one color component from the generated image; e) dividing each color component into blocks of pixels of a fixed size Wb x Hb; f) generating a first sequence of blocks by shuffling the pixel blocks obtained in step e according to a given key; g) calculating an image identifier based on a first part of the pairs of blocks in the first sequence of blocks, wherein for each pair of blocks an average value of the pixels in each block is calculated; h) generating an embedding for protecting the generated image, consisting of a sequence including at least one of: an identification message, an image identifier generated in step g), input time, user ID, a text query, and additional text information;i) generating an encrypted sequence by encrypting the embedding to protect the generated image with a private key; j) generating an embedding sequence consisting of the image identifier generated in step g) and the encrypted sequence obtained in step i); k) embedding the sequence obtained in step j) in the form of a bit sequence into pairs of pixel blocks included in the second part of the blocks of the first sequence of blocks; l) providing the protected image to the user.

[0013] In one particular example of implementation, the image is generated in at least one of the formats selected from the group: JPG, BMP, or PNG.

[0014] In another particular example of implementation, the color component is selected from the group: color space, color model, color channel, subtractive model.

[0015] In another particular example of implementation, the color component is at least one of: RGB, RGB A, HSL, HSLA, CMY, CMYK, XYZ, LMS, Hunt, ICtCp, LAB, RLAB, YCbCr, grayscale image.

[0016] In another particular example of implementation, the generated image is converted to a given resolution Wimg x Himg.

[0017] In another particular example of implementation, the robustness coefficient R is determined for blocks of pixels in the first sequence.

[0018] In another particular example of implementation, a value is determined based on R for determining blocks for embedding bits of the sequence.

[0019] In another particular example of implementation, error-correcting coding is applied to the generated embedding at step h).

[0020] In another particular example of implementation, error-correcting coding is applied to the generated encrypted sequence at step i).

[0021] In another particular example of implementation, error-correcting coding is applied to the sequence formed in step j).

[0022] The claimed solution is also implemented using a system for protecting the authenticity of images generated on the basis of a text query by a machine learning model, comprising at least one processor and at least one memory in which machine-readable instructions are stored, which, when executed by the processor, perform the above-mentioned method. BRIEF DESCRIPTION OF DRAWINGS

[0023] Fig. 1 illustrates a general view of the method for protecting images.

[0024] Fig. 2 illustrates a block diagram of the claimed method.

[0025] Fig. 3A illustrates an example of blocks associated with blocks of image pixels.

[0026] Fig. ЗБ illustrates an example of obtaining sequences of block pairs for calculating the image identifier and for embedding security information.

[0027] Fig. 4 illustrates an example of defining bits for obtaining an image identifier.

[0028] Fig. 5 illustrates an example of the change in average pixel values ​​within blocks after embedding security information.

[0029] Fig. 6 illustrates an example of an attachment and its parts as a byte sequence.

[0030] Fig. 7 illustrates a view of the computing device. IMPLEMENTATION OF THE INVENTION

[0031] Fig. 1 shows a general diagram of the claimed solution, which is implemented by processing incoming user requests (110), which are a text description of the generated image. In the field of AI, such text Requests describing the object to be generated for a machine learning model are also called "prompts." A request from a user (110) can be transmitted via an API (120) implemented in a web browser or messenger (e.g., Telegram), which then enables the subsequent exchange of data directly with the image generation model (130) (e.g., Kandinsky) between the user (software) and the image protection module (140).

[0032] The image protection module (140) can be implemented using an external hardware and software system, such as a server or remote workstation, providing functionality for processing incoming images (121) generated by the model (130). The module (140) can also be located in the same execution loop as the model (130), such as a server or cloud computing node.

[0033] Upon completion of the image protection module (140), security information containing an encrypted attachment is embedded into the initially generated images (121). This information allows for at least the image's authorship to be established as having been created using the model (130), as well as the identifier of the user who generated the request and the request text itself. The protected image (141) is subsequently transmitted to the user (110) via the API (120), for example, by displaying the image (141) in a web browser or instant messenger.

[0034] Data exchange within the framework of the system implementation is based on standard principles of data exchange, in particular with the help of the Internet computing network, which can be implemented using any known principles known from the state of the art.

[0035] Fig. 2 shows an example of implementing the claimed method (200) for protecting generated images. At the first stage (201), user input is received for image generation by the model (130). The user input includes a user text query, a user ID (e.g., IP address, MAC address, registration ID, etc.), and the time of the text query. This data is recorded by the image protection module (140) for subsequent generation of security information for embedding in the image.

[0036] Next, at step (202), the model (130) generates an image (121) according to the user request and transmits it to the image protection module (140). The image is generated in JPG (JPEG), BMP, PNG, or other format containing explicit or encoded pixel values ​​of the color channel.

[0037] Upon receipt of the generated image (121) by the module (140), at step (203) a pixel block file is extracted for the color component of the image. The color component is at least one of: RGB, RGBA, HSL, HSLA, CMY, CMYK, XYZ, LMS, Hunt, ICtCp, LAB, RLAB, YCbCr, a grayscale image.

[0038] Fig. 3A shows an example of obtaining pixel blocks. A color channel is decoded from an image file. When the original JPEG image is received at step (121) to obtain the luminance (Y) and color (CLb (relative blueness) and Cr (relative redness) components of the image, the discrete cosine transform (DCT) coefficients are decoded. RGB color channels are obtained from the YCbCr components. When the original BMP image is received at step (121) from the file, the RGB color channels can be immediately extracted. At least one color channel is used. The channel pixels are divided into pixel blocks of a fixed size Wb x Hb and a sequence of pixel blocks (300) is formed.

[0039] In one embodiment of the invention, the original image is pre-sized to a fixed size of Wimg x Himg. This function is used to ensure that each generated image, regardless of its size (resolution), can be verified for authorship. In one signature destruction scenario, an attacker can resize the image. While this may preserve the signature in the image data, retrieving it requires resizing the image to a standard size—the size at which the signature was embedded.

[0040] Next, at step (204), a new block sequence is formed by shuffling it using a given key (seed). The resulting shuffled sequence may be identical to the original. The shuffled sequence is divided into pairs of blocks. The sequence of block pairs is divided into two parts. The first sequence (301) has a fixed size. The second sequence (302) includes all remaining pairs of blocks.

[0041] Fig. 3B shows an example of forming two sequences of block pairs. The original block sequence (300) is shuffled using a given key (seed). The new sequence is divided into two parts. In this example, the first sequence (301) consists of four pairs of blocks (eight pixel blocks). The remaining pairs of blocks form sequence (302).

[0042] At step (205), the image identifier (unique value, unique identifier) ​​is calculated based on the first sequence (301) of blocks. The image identifier allows the generated embedding to be linked to a specific image. Subsequent embedding cannot be transferred to another image, as it contains a unique identifier for that image. This identifier is a kind of hashing and can be used as a cryptographic hash of the image.

[0043] The image identifier is formed based on the first sequence (301) of blocks. The next pair of blocks is selected. The average pixel value of the first block S1 and the average pixel value of the second block S2 in the pair are calculated. If the difference between their average values ​​is greater than the specified robustness coefficient R: S1-S2 > R, then the unique value bit is fixed at 1. If the difference between their average values ​​is less than the specified robustness coefficient R with a negative sign: S1-S2 < -R, then the unique value bit is fixed at 0. If the difference between their average values ​​lies in the range from -R to R, then a certain value is added / subtracted from all pixels of the above-mentioned blocks SI, S2 such that the absolute difference between their averages exceeds R, after which the corresponding unique value bit is fixed.

[0044] In one embodiment of the invention, a certain value is added / subtracted so that the absolute difference between the average values ​​of pixels S1 and S2 exceeds R not for all block pixels, but only for some pixels. Using this feature, by excluding pixels located on the block boundary from the change, it is possible to reduce signature embedding artifacts visible to the human eye when examining the boundaries of two blocks. Thus, pixels in blocks located on the block boundary are not changed, and the transition in the image from one block to another will be smoother and less noticeable.

[0045] The use of the robustness coefficient R in the invention allows for the formation of high stability (indestructibility) in the image signature during color correction of the image, resizing of the image, compression of the image and other transformations.

[0046] Fig. 4 shows an example of generating an image identifier for the selected boundary value R=10. Each pair of blocks generates its own bit of information. The generated bits define a unique sequence of bytes (a number), which is the image identifier. Within each block is a number that denotes the average pixel value within the block. For the first pair of blocks, we have S1=55 and S2=49. Their absolute difference is |55-49| <R, поэтому необходимо изменить пиксели внутри блоков. Прибавление ко всем пикселям первого блока значения 2 позволяет получить новое среднее значение пикселей S 1 =57. Вычитанием от всех пикселей The second block's value of 3 yields a new average pixel value of S2=46. Now their absolute difference is |57-46|>R. The difference value itself, 57-46=11>R, determines the bit to be equal to 1.

[0047] At step (206), a security attachment is generated. The attachment consists of a public (unencrypted) portion and an encrypted portion. The encrypted portion is an encrypted data sequence that may include the following: entry time, user ID, text query (prompt request), image identifier value obtained at step (205), and additional text information (e.g., the designation "I am a Sber model"). The public portion may include information such as a verification digit, image identifier value, and an integrity check digit for the encrypted attachment. An example of such an attachment is shown in Fig. 6.

[0048] In one particular embodiment of the invention, arbitrary data of arbitrary size, varying depending on the request or its conditions, is embedded. In this case, when extracting the data, its quantity and size are initially unknown. To enable extraction, the following data is inserted at the beginning of the encrypted portion: the number of different data items in a fixed number of first bits, and the size of each embedding in a fixed number of subsequent bits. Figure 6 shows an example of such an embedding. The amount of data is located at the beginning of the encrypted portion of the embedding. In this example, the embedding contains three types of data: - image identifier; - text information, including a fixed text phrase, user ID, request time, request prompt; - other data about the image. Examples of other image data include data such as: a thumbnail of the image as a compressed version of the original image; a history of prompt requests if there are multiple generations of the current image; the personal signature of the image generator; or other data that the image creator wishes to preserve in the signature, including in the form of graphic / audio or other content. The first 32 bits of the encrypted portion of the attachment are allocated to the amount of data. Next come the data sizes: 64 bytes for the image identifier; 163 bytes for text information; and 376 bytes for other image data. Each size is allocated the next 32 bits of the encrypted portion of the attachment.

[0049] Additionally, the secure attachment may include other data, such as a compressed generated image (121) or a portion thereof, which is converted into a byte sequence appended to the remaining data. The resulting attachment is signed with the private key at step (206) and formed part of the secure attachment at step (207). The resulting encrypted byte sequence is represented as a bit sequence.

[0050] Next, at step (208), the information obtained at step (207) is embedded into the pixels of the image.

[0051] First, the plaintext (unencrypted) information is embedded. This information may include a data sequence such as a verification digit, an image identifier (obtained in step (205)), or a security embedding integrity check digit, such as one calculated using CRC16, CRC32, CRC64, or any other checksum or hashing algorithm. This information is represented as a bit sequence.

[0052] Next, the protective embedding obtained in step (207) is embedded. The embedding in step (208) is embedded into the pixels of the pairs of blocks obtained in step (204), which are included in the shuffled sequence (302).

[0053] An example of embedding encoding within image pixel blocks is shown in Fig. 5. Before this step, a shuffled sequence of pixel blocks (302) is formed (example in Fig. 3B). Fixed embedding parameters are selected: R is the robustness coefficient (stability of the embedding), and D is the image distortion coefficient. In Fig. 5, the values ​​R=10 and D=40 are given as an example.

[0054] As shown in the example in Fig. 5, the first bit of the embedded information is 1, the second is 1, the third is 0, and the fourth is 1. The first pair of blocks from the sequence (302) is taken. In Fig. 5, a number is written inside each block, which denotes the average value of the pixels within the block, and the lines inside the blocks depict the pattern of the original image. The average value of all pixels within the first block S1 and the average value of all pixels within the second block S2 are calculated. In the example, the pixel blocks are the same and the average values ​​Sl=94 and S2=94. First, the conditions |S1-S2| <D: - If |S1-S2| <D, то данная пара блоков пикселей используется для встраивания информации; - If |S1-S2|>=D, then this pair of pixel blocks is not used for embedding information, and a transition to the next pair of blocks is made.

[0055] Further, if the condition |S1-S2| is met <D, то выбирается очередной бит вкладываемой информации. Для его внедрения считается среднее значение пикселей первого блока S1 и среднее значение пикселей второго блока S2 в паре. Если требуется встроить бит 1, то ко всем пикселям блоков 1 и 2 прибавляется / отнимается некоторое значение так, чтобы разница их средних была больше заданного значения коэффициента робастности R (меньшего D): S1-S2 > R. If it is necessary to embed bit 0, then a certain value is added / subtracted from all pixels of blocks 1 and 2 so that the difference between their averages is less than the specified value of the robustness coefficient R with a negative sign: Sl-S2 < -R.

[0056] Step by step, this process looks like this: - If it is required to embed bit 1 and Sl-S2<-R and the condition |S1-S2| is met <D, то ко всем пикселям блоков 1 и 2 прибавляется / отнимается некоторое значение так, чтобы разница их средних была больше заданного значения коэффициента робастности R: S1-S2 > R; - If it is required to embed bit 0 and Sl-S2<-R and the condition |S1-S2| is met <D, то блоки пикселей остаются без изменения; - If it is required to embed bit 1 and S1-S2>R and the condition |S1-S2| is met <D, то блоки пикселей остаются без изменения; - If it is required to embed bit 0 and S1-S2>R and the condition |S1-S2| is met <D, то ко всем пикселям блоков 1 и 2 прибавляется / отнимается некоторое значение так, чтобы разница их средних была больше заданного значения коэффициента робастности R: S1-S2 < -R. - If the condition |S1-S2| <D не выполнено, осуществляется переход к следующей паре блоков пикселей. This process continues until all required bits have been embedded or until there are no more pairs of pixel blocks in the sequence (302).

[0057] The image distortion coefficient D is also one of the distinctive features of the invention, enabling new quality improvements in information security. It eliminates blocks with large differences in average pixel values ​​from the set of block pairs into which embedding occurs. If such block pairs are retained, embedding a bit of information into them may require a very large change in the block's pixels. These changes will be Clearly visible to the human eye. Such artifacts will reduce the quality of the resulting image generation and also allow an attacker to identify which blocks the embedding was made in and target it for destruction.

[0058] The robustness coefficient (embedding stability) R also allows for improved information security. It allows one to control the balance between the embedding's resistance to image distortion and the appearance of embedding artifacts visible to the human eye. Increasing the robustness coefficient R results in a more reliable signature embedding that will not be destroyed by significant image distortion. However, this may lead to the appearance of embedding artifacts. By decreasing the robustness coefficient R, one can achieve a virtually complete absence of artifacts visible to the human eye.

[0059] In the example in Fig. 5, the first bit of the embedded information is equal to 1. Bit 1 means that the difference S1-S2 must be greater than R, i.e. S1-S2>R. In the given example, Sl=94 and S2=94, which means 94-94=0. To embed a bit of information equal to 1, the value 5 is added to all pixels of the first block and the value 6 is subtracted from all pixels of the second block. It turns out that S1=99, S2=88, then 99-88=11>R. The bit of information equal to 1 is embedded. Then, a transition to the embedding of the next bit of information and a transition to the next pair of blocks in the sequence (302) is performed.

[0060] The second bit of embedded information is equal to 1. Bit 1 means that the difference S1-S2 must be greater than R, i.e. S1-S2>R. In this case, S1=45 and S2=52, which means 45-52 = -7. To embed a bit of information equal to 1, the value 9 is added to all pixels of the first block and the value 9 is subtracted from all pixels of the second block. The result is S1=54, S2=43, then 54-43=11>R. The bit of information equal to 1 is embedded. Then the transition to the embedding of the next bit of information is performed.

[0061] The third bit of the embedded information is equal to 0. Bit 1 means that the difference S1- S2 must be less than -R, i.e. Sl-S2<-R. In this case, 155-110=45. This contradicts the condition |S1-S2| <D. Данная пара блоков пропускается, осуществляется переход к следующей паре блоков. Для следующей пары блоков разница пикселей 85-91 = -6. Для внедрения бита информации равный 0 происходит вычитание из всех пикселей первого блока значения 2 и прибавление ко всем пикселям второго блока значения 3. Получается S 1 =83, S2=94, тогда 83-94 = -11 < -R. Бит информации равный 0 внедрен.

[0062] The next bit of the message is 1. Bit 1 means that the difference S1-S2 must be greater than R, i.e. S1-S2>R. For the next pair of blocks, we have 87-76=11. Condition S1- li S2>R has already been executed. This block pair remains unchanged. An information bit equal to 1 is considered embedded.

[0063] The final embedded security data sequence may look like this: [check digit (e.g., 537), image identifier, encrypted attachment integrity check digit, encrypted attachment]. The data in the encrypted attachment may look like this: [number of attachments (e.g., 3), size of 1st attachment, size of 2nd attachment, size of 3rd attachment, image identifier, text ("I am a Sber model..."), additional data].

[0064] In byte form it will look like this: 00000219ac3eb891bc32652a83b4ff62cc2e49de3293d99e24ae8907e90c96b55cb5578a78d78e8 87f5f436f6b9b989c6896b689e986s64a3509ee90fca..., where 00000219 is the byte form of the number 537, ac3eb891bc32652a83b4ff62cc2e49de32 is the byte form of the image identifier, 93d99e24 is the byte form of the CRC32 check digit value equal to 2480512548, then comes the encrypted data. Figure 6 shows an example of a byte sequence to be embedded. This sequence is represented as a bit sequence.

[0065] The resulting pixel blocks are arranged in the original sequence (300). The protected image (141) with the information embedded in step (208) is displayed on the user's device at step (209).

[0066] Next, we will consider the process of extracting embedded information when checking a protected image (141).

[0067] Similar to steps (203) - (205), the image identifier value is obtained and the sequence of pixel blocks where information embedding has potentially been performed.

[0068] The following algorithm is then executed.

[0069] Step 1. The first 4 bytes are extracted to obtain a check digit. This check digit is verified to be the original check digit. If the check digit matches, extraction continues. If not, a message is generated indicating that the digital signature was not found or was corrupted.

[0070] Step 2. The next 64 bytes are extracted. This is the hash of the DC coefficients. The resulting value is checked against the actual image ID value. If the value matches, extraction continues. If not, a message is displayed indicating that the digital signature was not found or was corrupted.

[0071] Step 3: The next 4 bytes are extracted, which allows us to obtain the checksum of the encrypted attachment.

[0072] Step 4. The encrypted attachment is extracted. All subsequent bits in the attachment refer to the encrypted attachment. At this stage, the size of the attachment is unknown. However, in the first block of the encrypted attachment—256 bytes for RSA—the lengths of the attachments are found. The first 256 bytes are extracted and decrypted using the public key. The first 4 bytes in this block are the number of attachments—in the example in Fig. 6, there are three attachments. Each subsequent 4 bytes is the size of each of these attachments. They contain the numbers 64, 163, and 376. If the size of any attachment is greater than 30,000 (the maximum attachment limit), the ciphertext block has been modified, causing a message to be generated indicating that the digital signature was not detected or was corrupted. Otherwise, the potential sizes of the attachments are obtained and extraction continues.

[0073] Next, the size of the encrypted attachment is determined. According to the example in Fig. 6, the obtained 64 bytes are the size of the first attachment, which is the image identifier value, then 163 bytes are the size of the second attachment - this is the text size. The size of the third attachment is 3756 bytes - this is the image thumbnail (compressed image). Then the total size of the encrypted attachment is: 64 + 163 + 3756 = 3983 (the attachment itself); 4 bytes are added for the number of attachments and 3 more times 4 bytes for their lengths. The result is as follows: 64 + 163 + 376 + 4 + 3 * 4 = 619 bytes. Since the encryption block has a size of 256, then in the given example 619 / / 256 = 3 whole encryption blocks, which is equal to 768 bytes of encrypted attachment.

[0074] 768 bytes are extracted and used to calculate the CRC32 checksum. If the value matches the checksum obtained in step 3, extraction continues. If not, a message is generated indicating that the digital signature was not found or was corrupted.

[0075] Decryption of 768 bytes. From the first block, as noted above, the number of attachments, their lengths, and the attachments themselves are obtained. The decrypted byte sequence is broken down by the number of bytes in each attachment. The original attachments are obtained.

[0076] The image ID value from the decrypted data is compared with the image ID value from the attachment (which was already compared with the actual value in step 2). If the values ​​match, a message is displayed indicating the digital signature has been verified, and the extracted information—the value—is displayed. Image ID, text, and additional data. If not, a message is displayed indicating that the digital signature was not found or was corrupted.

[0077] Fig. 7 shows a general view of a computing device (500) suitable for performing the method (200). The device (500) may be, for example, a server or another type of computing device that can be used to implement the claimed technical solution, including: a smartphone, tablet, laptop, computer, etc. The device (500) may also be part of a cloud computing platform.

[0078] In general, the computing device (500) comprises one or more processors (501), memory means such as RAM (502) and ROM (503), input / output interfaces (504), input / output devices (505), and a device for network interaction (506), connected by a common information exchange bus.

[0079] The processor (501) (or several processors, multi-core processor) can be selected from a range of devices that are widely used at the present time, for example, from Intel™, AMD™, Apple™, Samsung Exynos™, MediaTEK™, Qualcomm Snapdragon™, etc. A graphics processor can also be used as a processor (401), for example, from Nvidia, AMD, Graphcore, etc.

[0080] RAM (502) is random access memory (RAM) and is designed to store machine-readable instructions executed by the processor (501) to perform the necessary logical data processing operations. RAM (502) typically contains executable instructions from the operating system and corresponding software components (applications, software modules, etc.).

[0081] ROM (503) represents one or more permanent data storage devices, such as a hard disk drive (HDD), a solid-state drive (SSD), flash memory (EEPROM, NAND, etc.), optical storage media (CD-R / RW, DVD-R / RW, BlueRay Disc, MD), etc.

[0082] To organize the operation of the device components (500) and to organize the operation of external connected devices, various types of I / O interfaces (504) are used. The choice of the appropriate interfaces depends on the specific design of the computing device, which may include, but are not limited to: PCI, AGP, PS / 2, IrDa, FireWire, LPT, COM, SATA, IDE, Lightning, USB (2.0, 3.0, 3.1, micro, mini, type C), TRS / Audio jack (2.5, 3.5, 6.35), HDMI, DVI, VGA, Display Port, RJ45, RS232, etc.

[0083] To ensure user interaction with the computing device (500), various means (505) of I / O information are used, for example, a keyboard, a display (monitor), a touch screen, a touchpad, a joystick, a mouse, a light pen, stylus, touchpad, trackball, speakers, microphone, augmented reality tools, optical sensors, tablet, indicator lights, projector, camera, biometric identification tools (retinal scanner, fingerprint scanner, voice recognition module), etc.

[0084] The network interaction means (506) ensures the transmission of data by the device (500) via an internal or external computer network, for example, an Intranet, the Internet, a LAN, etc. One or more means (506) may be, but are not limited to: an Ethernet card, a GSM modem, a GPRS modem, an LTE modem, a 5G modem, a satellite communication module, an NFC module, a Bluetooth and / or BLE module, a Wi-Fi module, etc.

[0085] Additionally, satellite navigation tools included in the device (500) can also be used, for example, GPS, GLONASS, BeiDou, Galileo.

[0086] The submitted application materials disclose preferred examples of the implementation of the technical solution and should not be interpreted as limiting other, particular examples of its implementation that do not go beyond the scope of the requested legal protection, which are obvious to specialists in the relevant field of technology. Sources of information: 1. Hamilton, Eric: JPEG File Interchange Format, Version 1.02. 1 September 1992; 2. Recommendation ITU-T T.871: Information technology - Digital compression and coding of continuous-tone still images: JPEG File Interchange Format (JFIF). Approved 14 May 2011; posted 11 September 2012; 3. Recommendation ITU-T T.81: Information technology - Digital compression and coding of continuous-tone still images - Requirements and guidelines. Approved 18 September 1992; posted April 14, 2004. 4. GOST P 34.10-2012. Cryptographic protection of information. Processes for the formation and verification of electronic digital signatures. https: / / rst.gov.ru:8443 / file-service / file / load / 1699366979620

Claims

FORMULA 1. A method for protecting the authenticity of images generated based on a text query by a machine learning model, comprising the steps of: a) receiving user input data for generating an image, containing a text query; b) generating an image using a machine learning model; c) recording at least one type of user input data selected from the group: input time, user ID, text query; d) extracting pixels of at least one color component from the generated image; e) dividing each color component into blocks of pixels of a fixed size Wb x Hb; I) forming a first sequence of blocks by mixing the pixel blocks obtained in step e according to a given key); g) calculating an image identifier based on a first part of the pairs of blocks in the first sequence of blocks, wherein for each pair of blocks an average value of pixels in each block is calculated; h) forming an embedding for protecting the generated image, consisting of a sequence including at least one of: an identification message, an image identifier formed in step g), the input time, a user ID, a text query and additional text information; i) generating an encrypted sequence by encrypting the embedding for protecting the generated image with a private key; j) generating an embedding sequence consisting of the image identifier formed in step g) and the encrypted sequence obtained in step i);k) embedding the sequence obtained in step j) in the form of a sequence of bits into pairs of pixel blocks included in the second part of the blocks of the first sequence of blocks; l) providing the protected image to the user.

2. The method according to claim 1, wherein the image is generated in at least one of the formats selected from the group: JPG, BMP, or PNG.

3. The method according to claim 1, wherein the color component is selected from the group: color space, color model, color channel, subtractive model.

4. The method according to paragraph 3, in which the color component is at least one of: RGB, RGBA, HSL, HSLA, CMY, CMYK, XYZ, LMS, Hunt, ICtCp, LAB, RLAB, YCbCr, a grayscale image.

5. The method according to claim 1, wherein the generated image is reduced to a given resolution Wimg x Himg.

6. The method according to claim 1, wherein a robustness coefficient R is determined for blocks of pixels in the first sequence.

7. The method according to claim 6, wherein based on R a value is determined for determining blocks for embedding bits of the sequence.

8. The method according to claim 1, in which error-correcting coding is applied to the generated embedding at step h).

9. The method according to claim 1, in which error-correcting coding is applied to the generated encrypted sequence at step i).

10. The method according to claim 1, in which error-correcting coding is applied to the sequence formed in step j).

11. A system for protecting the authenticity of images generated on the basis of a text query by a machine learning model, comprising at least one processor and at least one memory in which machine-readable instructions are stored, which, when executed by the processor, perform the method according to any one of paragraphs 1-10.

Citation Information

Patent Citations

  • Method of creating and checking electronic image certified by digital watermark

    RU2399953C1

  • Protecting deep learning models using watermarking

    US20190370440A1

  • Digital watermarking of machine learning models

    US20210019605A1

  • Immutable watermarking for authenticating and verifying ai-generated output

    US20210390447A1

  • Robust digital watermarking

    US6724911B1