Artificial Intelligence Generated Image Recognition Method and System Based on Generative Watermarking

By watermarking and semantic analysis of images, and noise reduction processing of noise images is solved, the problem of high-quality artificial intelligence-generated images is solved, and accurate image identification is achieved.

CN119919783BActive Publication Date: 2025-07-22STATE GRID DIGITAL TECHNOLOGY HOLDING CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510413329.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively identify and identify high-quality artificial intelligence-generated images, resulting in their overflow and spread and adverse effects.

Method used

By extracting the image watermark information and semantic information of the image latent vector, analyzing the noise image and denoising the noise, the image similarity is judged to determine whether an image is generated by artificial intelligence.

Benefits of technology

It improves the recognition accuracy of images generated by artificial intelligence and avoids misjudgment and missed detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919783B_ABST
    Figure CN119919783B_ABST
Patent Text Reader

Abstract

The present invention provides an artificial intelligence generated image recognition method and system based on generative watermarking. The implementation solution is as follows: semantic analysis is respectively performed on the first watermark information vector and the first image latent vector in the first image to obtain the first image semantic information, the first non-image semantic information, and the second image semantic information; when the similarity between the first image semantic information and the second image semantic information does not meet the requirements, the first image semantic information is encoded to obtain a prompt word vector, and based on the first image semantic information and the noise information related to the first non-image semantic information, multiple noise images are generated, and based on the prompt word vector, noise reduction processing is respectively performed on each noise image to obtain multiple second images; when the image similarity between the first image and each second image meets the requirements, it is determined that the first image is an artificial intelligence generated image. By adopting the present invention, the image identification accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image technology, and in particular, to an artificial intelligence generated image recognition method and system based on generative watermarking. Background Art

[0002] With the rapid development of Generative Artificial Intelligence (GAI), the quality of images generated by artificial intelligence is getting higher and higher, almost reaching the level of being indistinguishable from real ones. The proliferation of these images generated by artificial intelligence has brought a series of adverse effects after extensive dissemination. Therefore, there is an urgent need for a method for effectively identifying or authenticating images generated by artificial intelligence. Summary of the Invention

[0003] The present invention provides an artificial intelligence generated image recognition method and system based on generative watermarking, which can solve at least one of the above technical problems.

[0004] According to one aspect of the present invention, there is provided an artificial intelligence generated image recognition method based on generative watermarking, including:

[0005] Performing watermark extraction on a first image to obtain a first watermark information vector and a first image latent vector;

[0006] Performing semantic parsing on the first watermark information vector to obtain first image semantic information and first non-image semantic information;

[0007] Performing semantic parsing on the first image latent vector to obtain second image semantic information;

[0008] In the case where the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, encoding the first image semantic information to obtain a prompt word vector, and generating a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information, and performing noise reduction processing on each of the noise images based on the prompt word vector to obtain a plurality of second images;

[0009] In the case where the image similarity between the first image and each of the second images meets a preset image similarity condition, determining that the first image is an image generated by artificial intelligence.

[0010] According to another aspect of the present invention, there is provided an artificial intelligence generated image recognition device based on generative watermarking, including:

[0011] A watermark extraction module for performing watermark extraction on a first image to obtain a first watermark information vector and a first image latent vector;

[0012] A first semantic parsing module, configured to perform semantic parsing on the first watermark information vector to obtain first image semantic information and first non-image semantic information;

[0013] A second semantic parsing module, configured to perform semantic parsing on the first image latent vector to obtain second image semantic information;

[0014] An image generation module, configured to, when the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, encode the first image semantic information to obtain a prompt word vector, and generate a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information, and perform noise reduction processing on each of the noise images based on the prompt word vector to obtain a plurality of second images;

[0015] A first image recognition module, configured to determine that the first image is an AI-generated image when the image similarity between the first image and each of the second images meets a preset image similarity condition.

[0016] According to another aspect of the present invention, there is provided an AI-generated image recognition method based on generative watermarking, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any one of the AI-generated image recognition methods based on generative watermarking in the embodiments of the present invention.

[0017] According to another aspect of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute any one of the AI-generated image recognition methods based on generative watermarking in the embodiments of the present invention.

[0018] Adopting the technical solution of the present invention, watermark extraction is performed on the first image to be detected to obtain a first watermark information vector and a first image latent vector. Semantic parsing is performed on the first watermark information vector to obtain first image semantic information and first non-image semantic information. At the same time, semantic parsing is performed on the first image latent vector to obtain second image semantic information. If the semantic similarity between the first image semantic information and the second image semantic information does not meet the preset semantic similarity condition, it indicates that the semantics expressed by the watermark in the image do not match the semantics of the image itself. At this time, further judgment needs to be made on the first image to determine whether it is an AI-generated image. A prompt word vector is generated using the image semantic information in the watermark information, and multiple noise images are generated based on the image semantic information in the watermark information and the noise information related to the non-image semantic information in the watermark information. Denoising processing is performed on each noise image based on the prompt word vector, so that an AI image related to the watermark information, that is, a second image, can be generated. In this way, an AI image strongly related to the semantic information in the watermark information can be generated. Thus, it is determined whether the image similarity between the first image and each second image meets the preset similarity condition. If it meets, the first image is determined to be an AI-generated image. Therefore, the recognition accuracy of AI-generated images can be improved.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present invention. Among them:

[0021] Figure 1 is a flowchart of a method for identifying AI-generated images based on generative watermarks according to an embodiment of the present invention;

[0022] Figure 2 is a flowchart of a watermark image parsing process according to an embodiment of the present invention;

[0023] Figure 3 is a flowchart of an AI generation process of a watermark image according to an embodiment of the present invention

[0024] Figure 4 is a structural block diagram of a device for identifying AI-generated images based on generative watermarks according to an embodiment of the present invention;

[0025] Figure 5 is a block diagram of an electronic device for implementing the method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present invention. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0027] Figure 1 It is a flowchart of an artificial intelligence-generated image recognition method based on generative watermarking according to an embodiment of the present invention.

[0028] As Figure 1 shown, the artificial intelligence-generated image recognition method based on generative watermarking may include:

[0029] S110, extracting a watermark from a first image to obtain a first watermark information vector and a first image latent vector;

[0030] S120, semantically parsing the first watermark information vector to obtain first image semantic information and first non-image semantic information;

[0031] S130, semantically parsing the first image latent vector to obtain second image semantic information;

[0032] S140, in the case where the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, encoding the first image semantic information to obtain a prompt word vector, and generating a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information, and performing denoising processing on each noise image based on the prompt word vector to obtain a plurality of second images;

[0033] S150, in the case where the image similarity between the first image and each second image meets a preset image similarity condition, determining that the first image is an artificial intelligence-generated image.

[0034] Exemplarily, the first image is an image with a plaintext watermark or a hidden watermark. If a watermark information vector can be obtained by extracting the watermark from the first image, it indicates that the first image is an image with a plaintext watermark or a hidden watermark. The first image can also be referred to as a watermark image.

[0035] Exemplarily, a pre-trained watermark extraction network can be used to extract the watermark from the first image. The first image is input into the watermark extraction network, and the image encoder in the watermark extraction network converts the first image into a 1024-dimensional watermark image latent vector and upsamples it to a 1792-dimensional watermark image latent vector. Then, using the network parameters of the watermark extraction network, the 1792-dimensional watermark image latent vector is feature-split to obtain the first watermark information vector and the first image latent vector.

[0036] Exemplarily, the following formula can be used to represent the above feature-splitting process:

[0037] v_w_watermark = (v_w_combined − α·v_w_image) ÷ (1−α);

[0038] Where, v_w_watermark represents the first watermark information vector, v_w_combined represents the watermark image latent vector, α represents the network parameters of the watermark extraction network, and v_w_image represents the first image latent vector.

[0039] Exemplarily, as Figure 2 shown, the watermarked image to be detected is input into the watermark extraction network to obtain a watermark text vector, also known as a watermark information vector. If the watermark information vector is encrypted information, the watermark information can be decrypted. The watermark text decoder is used to decode the watermark text vector to obtain a watermark text sequence. The first image semantic information can be extracted from the watermark text sequence. Thus, based on the first image semantic information, the image recognition process of the above steps is performed until the identification result of whether the watermarked image is an AI-generated image is obtained.

[0040] Exemplarily, a pre-trained semantic recognition network can be used to perform semantic parsing on the first watermark information vector and the first image latent vector to obtain the first image semantic information and the second image semantic information. The first watermark information vector includes not only image semantic information but also non-image semantic information. The non-image semantic information can include the generation platform, generation time, and generation method of the first image, or some unrecognizable noise information, etc.

[0041] Exemplarily, the first image semantic information refers to the semantic information in the watermark information of the first image to be recognized, and the second image semantic information refers to the semantic information in the image information of the first image to be recognized.

[0042] Exemplarily, the image semantic information may include multiple types of information, such as object category, scene description, attribute information, spatial location relationship, and emotional intention, etc. For example, the object category may be objects such as people, animals, or objects included in the image. The scene description may be the overall scene of the image. The attribute information may include the color, shape, and actions of the object, etc. The spatial relationship may include the positional relationship between objects. The emotional intention may include the emotion or intention conveyed by the image.

[0043] Exemplarily, the first semantic similarity condition may be that the semantic similarity between the first image semantic information and the second image semantic information is less than a preset semantic similarity threshold. For example, the semantic similarity between the two is less than 80%, which means that the semantic information in the watermark information in the image to be detected has a low or no match with the semantic information in the image information in this image, indicating that there is a high probability that the image to be detected is not an AI-generated image. However, to avoid missing the detection of this image, it is necessary to further identify whether the image to be detected is an AI-generated image.

[0044] It can be understood that the first image is the image to be detected, and the second image is a reference image for detecting the image to be detected. The second image is an image with a plaintext watermark or an image with a hidden watermark. Therefore, the second image can also be referred to as a watermark reference image.

[0045] Exemplarily, a text encoder is used to encode the first image semantic information to obtain a prompt word vector.

[0046] Exemplarily, after synonym mutation and recombination of each sub-semantic information in the first image semantic information, multiple text-to-image prompt words are obtained. A text encoder is used to encode each text-to-image prompt word to obtain each text-to-image prompt word vector. Based on the platform information and time information in the first non-image semantic information, a noise vector is generated. Based on each text-to-image prompt word vector and the noise vector, multiple noise images are generated.

[0047] Exemplarily, the image similarity condition may be that the total number of second images with a similarity greater than a preset threshold to the first image is greater than a preset quantity threshold. For another example, the similarity between the first image and each second image is greater than the preset similarity threshold.

[0048] According to the above embodiments, a preliminary judgment is made on the similarity between the semantic information in the watermark information of the first image and the semantic information in the image information of the image. If it is determined that the similarity between the semantic information in the watermark information of the first image and the semantic information in the image information of the image does not meet the preset similarity condition, the first image semantic information is encoded to obtain a prompt word vector, and based on the first image semantic information and the noise information related to the first non-image semantic information, multiple noise images are generated. Based on the prompt word vector, each noise image is respectively denoised to obtain multiple second images. In this way, since the prompt word vector is generated from the semantic information in the watermark information of the image to be detected, by using this prompt word vector to denoise the noise images, the image content related to the semantic information in the watermark information can be retained as much as possible after denoising. In this way, the accuracy of the reference image used to detect the image to be detected can be improved. Furthermore, a re-judgment is made on the first image and each of the second images used as references. If the similarity between the first image and each of the second images used as references is greater than the preset image similarity condition, it is determined that the first image is an AI-generated image. Thus, by using this example, it can be accurately determined whether the watermark image to be detected is an AI-generated image, avoiding missed detections.

[0049] In one embodiment, generating multiple noise images based on the first image semantic information and the noise information related to the first non-image semantic information includes: performing semantic classification on the first image semantic information to obtain multiple sub-semantic information; determining a corresponding set of synonymous texts for each sub-semantic information based on each sub-semantic information; extracting one text from each set of synonymous texts respectively, and combining the extracted texts based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information to obtain a text-to-image prompt word; determining the generation time and generation platform of the first image based on the first non-image semantic information; determining the noise information based on the generation time and generation platform; generating multiple noise images based on each text-to-image prompt word and the noise information.

[0050] Exemplarily, the image semantic information may include multiple category information, such as object category, scene description, attribute information, spatial position relationship, and emotional intention, etc. For example, the object category may be objects such as people, animals, or objects included in the image. The scene description may be the overall scene of the image. The attribute information may include the color, shape, and action of the object, etc. The spatial relationship may include the position relationship between objects. The emotional intention may include the emotion or intention conveyed by the image.

[0051] For example, the semantic information semantic_info is exemplified as follows:

[0052] {

[0053] "object": ["kitten", "sofa"];

[0054] "scene": ["living room"];

[0055] "status": ["orange cat", "cat scratching", "blue sofa"];

[0056] "relation": ["cat on the sofa"];

[0057] "emotion": ["pleased"]

[0058] }

[0059] Exemplarily, the sub-semantic information can be a short text message and corresponds to one of the above categories.

[0060] Exemplarily, the set of synonymous texts can include multiple short texts with the same or similar meanings.

[0061] Exemplarily, one text is extracted from each set of synonymous texts, and based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information, the extracted texts are combined to obtain a text-to-image prompt. In this way, multiple text-to-image prompts can be generated. Moreover, the image semantic information expressed by these multiple text-to-image prompts is the same or highly similar.

[0062] Exemplarily, based on the generation time and generation platform, the noise information of the images generated by the platform during the time period where the time is located is obtained. For example, the historical images generated by the platform during the time period where the time is located can be used to determine their common noise from the historical images, and these common noises are used as the noise information.

[0063] Exemplarily, a text encoder is used to encode the text-to-image prompt and the noise information respectively to obtain a text-to-image prompt vector and a noise vector; an image generation network is used to process the text-to-image prompt vector and the noise vector to obtain a noise image.

[0064] According to the above embodiments, semantic classification is performed on the first image semantic information to obtain a plurality of sub-semantic information; based on each sub-semantic information, a corresponding set of synonymous texts for each sub-semantic information is determined; a text is extracted from each set of synonymous texts, and based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information, the extracted texts are combined to obtain a text-to-image prompt. In this way, multiple text-to-image prompts representing the same image semantics can be generated. Based on the first non-image semantic information, the generation time and generation platform of the first image are determined; based on the generation time and generation platform, noise information is determined. In this way, noise information related to the generation time and generation platform of the image to be detected can be obtained. Furthermore, based on multiple text-to-image prompts representing the same image semantics as the image to be detected, and noise information related to the generation time and generation platform of the image to be detected, multiple noise images highly related to the image to be detected can be generated.

[0065] In one embodiment, denoising processing is performed on each noise image based on the prompt vector to obtain a plurality of second images, including: encoding the noise image using an image encoder to obtain an initial image latent vector; inputting the initial image latent vector and the prompt vector into a denoising network to obtain a target image latent vector output by the denoising network; determining watermark text information based on the second non-image semantic information related to the noise information in the noise image and the first image semantic information, and encoding the watermark text information to obtain a second watermark information vector; using a feature fusion network to perform feature fusion on the target image latent vector and the second watermark information vector to generate a second image. The noise image is input into the image encoder. The convolutional layer in the image encoder extracts features from the noise image to obtain a feature map. The global average pooling layer performs average pooling on the feature map to obtain a feature vector. The fully connected layer maps the feature vector to a 1024-dimensional space to obtain an initial image latent vector;

[0066] Exemplarily, the image encoder is an encoder trained using the convolutional neural network ResNet50. The image encoder includes a convolutional layer, a global average pooling layer, and a fully connected layer.

[0067] Exemplarily, the noise image is input into the image encoder. The convolutional layer in the image encoder extracts image features from the noise image to obtain a feature map. The global average pooling layer performs pooling processing on the feature map to obtain a feature vector. The fully connected layer maps the feature vector to a 1024-dimensional space to obtain a 1024-dimensional initial image latent vector.

[0068] Exemplarily, in the denoising network, based on the image semantics expressed by the prompt vector, the initial image latent vector is gradually sampled for features related to the image semantics, so as to obtain a target image latent vector, which has removed noise information and retains features that are the same as or similar to the image semantics expressed by the prompt vector.

[0069] Exemplarily, the denoising network can be based on the UNet network, and a time step embedding and a conditional mechanism component are referenced for each sampling layer in the UNet network.

[0070] Among them, the UNet network can be composed of 4 downsampling layers, 1 intermediate layer, 4 upsampling layers, and 1 output layer. The downsampling layer consists of multiple convolutional blocks, and each convolutional block extracts a feature image. The upsampling layer consists of a transposed convolutional layer component, which is used to output the gradually restored feature map. The output layer is a fully connected layer, which is used to splice the feature maps output by the upsampling layer and the downsampling layer in the channel dimension to gradually generate the detailed information of the image.

[0071] The time step embedding component uses sine position encoding. After converting the time step t representing the current denoising stage into a vector, the time step embedding vector is mapped to the corresponding feature dimension through a fully connected layer, and the mapped time step embedding vector is added to the feature map of each convolutional block in the UNet network.

[0072] The conditional mechanism component is used to add the original prompt vector representation to the feature map of each convolutional block in the UNet network through the cross-attention mechanism.

[0073] The UNet network is used to predict the latent representation of the image during the image denoising generation process. The main role of the time step embedding is to embed the current denoising stage into the network, and the main role of the conditional mechanism is to input the prompt vector into the network to control the content generated by the UNet network.

[0074] The operation process of the denoising network is as follows: the input is the image latent vector v_image and the current time step t, that is, which stage of image denoising it is. The time step t is converted into a time step embedding vector, and the prompt word conditional information v_prompt is injected into the UNet through cross-attention. The UNet gradually denoises to generate the final image, and the output of the denoising network is the processed image information vector representation v_g_image.

[0075] It can be understood that since the target image latent vector does not include watermark information, in order to generate a reference watermark image of the same category as the watermark image to be detected, it is necessary to determine the watermark information, and then perform feature fusion on the watermark information vector and the target image latent vector. In this way, a reference watermark image can be generated.

[0076] Exemplarily, based on the noise information in the noise image, candidate noise information that matches the noise information is searched for, and the platform information and time information mapped by the candidate noise information are determined as the second non-image semantic information, and the second non-image semantic information is merged with the first image semantic information into watermark text information.

[0077] For example, the merged watermark text information is:

[0078] {

[0079] "platform":"p1",

[0080] "time":"2025-01-01 12:00:00",

[0081] "semantic_info":{

[0082] "object":["kitten","sofa"],

[0083] "scene":["living room"],

[0084] "status":["orange cat","cat scratching","blue sofa"],

[0085] "relation":["the cat is on the sofa"],

[0086] "emotion":["pleased"]}

[0087] }

[0088] Exemplarily, a feature fusion network is used to perform feature fusion on the target image latent vector and the second watermark information vector to generate a watermark image latent vector, and a watermark image decoder is used to decode and reconstruct the watermark image latent vector to obtain a second image.

[0089] Exemplarily, the feature fusion network is used to perform feature fusion operations on the 1024-dimensional target image latent vector v_g_image and the 768-dimensional second watermark information vector v_watermark, including feature weighted splicing and dimensionality reduction operations, so that the features of the generated image carry watermark information.

[0090] Among them, in terms of feature dimensions, the target image latent vector v_g_image and the second watermark information vector v_watermark are weighted and fused to form a high-dimensional vector. Among them, the target image latent vector v_g_image is the output of the noise reduction network, and its feature dimension is 1024 dimensions. The second watermark information vector v_watermark is one of the outputs of the text encoder, and the feature dimension is 768 dimensions. The generated watermark image latent vector after fusion is 1792 dimensions.

[0091] Exemplarily, the formula for weighted fusion of the feature fusion network can be:

[0092] v_combined = α·v_g_image + (1−α)·v_watermark

[0093] Among them, the initial value of α can be 0.98 to minimize the impact of the spliced watermark information latent vector on the image content and quality during image reconstruction as much as possible.

[0094] Finally, a fully connected layer is used to map the spliced watermark image latent vector v_combined into a space with the same dimension as the target image latent vector v_g_image, that is, 1024 dimensions, to obtain the final watermark image latent vector.

[0095] Exemplarily, watermark image decoding is used to decode and reconstruct the watermark image latent vector v_combined obtained after the fusion of the image and watermark features, and an AI-generated image with watermark information of the prompt word is obtained.

[0096] Exemplarily, the network structure of the watermark image decoder can include a fully connected layer and multiple transposed convolutional layers.

[0097] Among them, the fully connected layer is used to map the watermark image latent vector v_combined into an initial feature map suitable for transposed convolution. For multiple transposed convolutional layers, batch normalization and ReLU activation functions are added after each transposed convolutional layer, and the last transposed convolutional layer uses the Sigmoid activation function to limit the pixels within a reasonable range, thereby gradually upsampling the feature map, restoring the image resolution, and generating the final AI-generated image w_image.

[0098] As Figure 3 shown, the text processing network processes the input prompt word to obtain image semantic information. The text encoder encodes the image semantic information and the text information composed of the image semantic information, platform information, and time information respectively to obtain the prompt word text vector and the watermark text vector. The latent space processing network processes the prompt word text vector to generate an image latent vector. The feature fusion network fuses the image latent vector and the watermark text vector to obtain a watermark image latent vector. The watermark image decoder decodes and reconstructs the watermark image latent vector to obtain an AI-generated reference watermark image.

[0099] According to the above embodiments, the semantic information of the first image is encoded to obtain a prompt vector, and a plurality of noise images are generated based on the semantic information of the first image and noise information related to the first non-image semantic information. Based on the prompt vector, each noise image is denoised to obtain a plurality of second images. In this way, since the prompt vector is generated from the semantic information in the watermark information of the image to be detected, by using this prompt vector to denoise the noise image, the image after denoising can retain as much image content related to the semantic information in the watermark information as possible. In this way, the accuracy of the reference image used to detect the image to be detected can be improved.

[0100] In one embodiment, the denoising network includes N upsampling layers connected in sequence and N downsampling layers connected in sequence, and the N upsampling layers are connected to the N downsampling layers through an intermediate layer. When the upsampling layer samples the input image latent vector, based on the prompt vector and the sampling time step embedding vector corresponding to the upsampling layer, the input image latent vector is upsampled through a cross-attention mechanism to obtain the image latent vector output by the upsampling layer. When the downsampling layer samples the input image latent vector, based on the prompt vector and the sampling time step embedding vector corresponding to the upsampling layer, the input image latent vector is downsampled through a cross-attention mechanism to obtain the image latent vector output by the downsampling layer.

[0101] Exemplarily, the input image latent vector of the first downsampling layer is the above-mentioned initial image latent vector. The output image latent vector of the last downsampling layer serves as the input image latent vector of the first upsampling layer. The output image latent vector of the last upsampling layer is the above-mentioned target image latent vector.

[0102] Exemplarily, a sinusoidal positional encoding is used to convert the time step t representing the denoising stage where the current sampling layer is located into a vector to obtain a sampling time step embedding vector.

[0103] Exemplarily, based on the residual information between the i-th downsampling layer and the (N - i + 1)-th downsampling layer, the sampling time step embedding vectors of the i-th downsampling layer and the (N - i + 1)-th downsampling layer are adjusted, and the adjusted sampling time step embedding vectors are added to the corresponding i-th downsampling layer and (N - i + 1)-th downsampling layer. The residual information can be used to determine the synchronization situation of the time steps, and the synchronization difference information between these two corresponding up and down sampling layers is used to adjust the sampling time step embedding vectors of these two sampling layers so that the delay time between them can be eliminated to achieve relative synchronization.

[0104] According to the above embodiments, the prompt word vector and the sampled time step embedding vector are injected into the corresponding sampling layer through the cross-attention mechanism, and the sampling layer determines the sampling direction according to the information that can be injected. The image features sampled in this sampling direction match the prompt words and the time steps, so as to gradually sample the image latent vector that meets the target requirements and remove the noise information.

[0105] In one embodiment, the above method further includes: when the semantic similarity between the first image semantic information and the second image semantic information meets the first semantic similarity condition, calculating the similarity between the first image semantic information and the third image semantic information corresponding to each text-image prompt word in the preset historical text-to-image prompt word set; when the similarity between the first image semantic information and the third image semantic information meets the second semantic similarity condition, determining that the first image is an AI-generated image.

[0106] In this example, when the semantic similarity between the first image semantic information and the second image semantic information meets the first semantic similarity condition, it is also possible to directly determine that the first image is an AI-generated image. However, to avoid misjudgment, it is possible to further use the preset historical text-to-image prompt word set to determine whether the semantic information in the watermark information of the watermarked image to be detected matches the semantic information in the historical text-to-image prompt words used by major platforms. If they match, it can be determined that the first image is an AI-generated image. In this way, the recognition accuracy of AI-generated images can be improved.

[0107] In one embodiment, the above method further includes: obtaining the historical text-to-image prompt words in the AI generation platforms at a preset frequency from each AI generation platform; updating the historical text-to-image prompt word set based on the historical text-to-image prompt words.

[0108] Exemplarily, the preset frequency can be every 1 hour, every 12 hours, etc.

[0109] Exemplarily, if the historical text-to-image prompt word set does not include the historical text-to-image prompt word, add the historical text-to-image prompt word to the historical text-to-image prompt word set. If the historical text-to-image prompt word set includes the historical text-to-image prompt word.

[0110] According to the above embodiments, the historical text-to-image prompt word set can be updated in a timely manner, which is convenient for improving the recognition accuracy when identifying AI-generated images later.

[0111] Figure 4 It is the structural block diagram of an AI-generated image recognition device based on generative watermarking according to an embodiment of the present invention.

[0112] Such as Figure 4As shown, the artificial intelligence generated image recognition device based on generative watermarking includes:

[0113] A watermark extraction module 410, configured to extract a watermark from a first image to obtain a first watermark information vector and a first image latent vector;

[0114] A first semantic parsing module 420, configured to perform semantic parsing on the first watermark information vector to obtain first image semantic information and first non-image semantic information;

[0115] A second semantic parsing module 430, configured to perform semantic parsing on the first image latent vector to obtain second image semantic information;

[0116] An image generation module 440, configured to, when the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, encode the first image semantic information to obtain a prompt word vector, and generate a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information, and perform noise reduction processing on each of the noise images based on the prompt word vector to obtain a plurality of second images;

[0117] A first image recognition module 450, configured to determine that the first image is an artificial intelligence generated image when the image similarity between the first image and each of the second images meets a preset image similarity condition.

[0118] In one implementation, the image generation module 440 includes:

[0119] A semantic classification unit, configured to perform semantic classification on the first image semantic information to obtain a plurality of sub-semantic information;

[0120] A text set determination unit, configured to determine a corresponding synonymous text set for each sub-semantic information based on each sub-semantic information;

[0121] A text combination unit, configured to extract one text from each of the synonymous text sets, and combine the extracted texts based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information to obtain a text-to-image prompt word;

[0122] A first information determination unit, configured to determine the generation time and generation platform of the first image based on the first non-image semantic information;

[0123] A second information determination unit, configured to determine noise information based on the generation time and the generation platform;

[0124] A noise image generation unit, configured to generate a plurality of noise images based on each of the text-to-image prompt words and the noise information.

[0125] In one embodiment, the image generation module 340 includes:

[0126] An initial image vector determination unit, configured to encode the noise image by using an image encoder to obtain an initial image latent vector;

[0127] A target image vector determination unit, configured to input the initial image latent vector and the prompt word vector into a denoising network to obtain a target image latent vector output by the denoising network;

[0128] A watermark vector determination unit, configured to determine watermark text information based on second non-image semantic information related to the noise information in the noise image and the first image semantic information, and encode the watermark text information to obtain a second watermark information vector;

[0129] A feature fusion unit, configured to perform feature fusion on the target image latent vector and the second watermark information vector by using a feature fusion network to generate the second image.

[0130] In one embodiment, the denoising network includes N upsampling layers connected in sequence and N downsampling layers connected in sequence, and the N upsampling layers are connected to the N downsampling layers through an intermediate layer;

[0131] When the upsampling layer samples the input image latent vector, based on the prompt word vector and the sampling time step embedding vector corresponding to the upsampling layer, perform upsampling on the input image latent vector through a cross-attention mechanism to obtain the image latent vector output by the upsampling layer;

[0132] When the downsampling layer samples the input image latent vector, based on the prompt word vector and the sampling time step embedding vector corresponding to the upsampling layer, perform downsampling on the input image latent vector through a cross-attention mechanism to obtain the image latent vector output by the downsampling layer.

[0133] In one embodiment, the above device further includes:

[0134] A similarity calculation module, configured to calculate the similarity between the first image semantic information and third image semantic information corresponding to each text-to-image prompt word in a preset historical text-to-image prompt word set when the semantic similarity between the first image semantic information and the second image semantic information meets a first semantic similarity condition;

[0135] A second image recognition module, configured to determine that the first image is an AI-generated image when the similarity between the first image semantic information and the third image semantic information meets a second semantic similarity condition.

[0136] In one implementation, the above device further includes:

[0137] A historical prompt word acquisition module, configured to acquire historical text-to-image prompt words in each AI generation platform at a preset frequency;

[0138] A prompt word set generation module, configured to update the historical text-to-image prompt word set based on the historical text-to-image prompt words.

[0139] For the specific functions and examples of the modules and sub-modules of the system according to the embodiments of the present invention, reference may be made to the relevant descriptions of the corresponding steps in the above method embodiments, which will not be elaborated here.

[0140] In the technical solution of the present invention, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0141] According to the embodiments of the present invention, the present invention also provides a system and a readable storage medium.

[0142] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0143] As Figure 5 shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0144] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0145] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the artificial intelligence-generated image recognition method based on generative watermarks. For example, in some embodiments, the artificial intelligence-generated image recognition method based on generative watermarks can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the artificial intelligence-generated image recognition method based on generative watermarks described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the artificial intelligence-generated image recognition method based on generative watermarks in any other suitable manner (e.g., by means of firmware).

[0146] The various embodiments of the systems and technologies described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a special or general programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0147] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine, or entirely on a remote machine or server.

[0148] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0149] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer that has: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0150] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0151] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating blockchain.

[0152] It should be understood that various forms of the flow shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. There is no limitation herein.

[0153] The above specific implementation manners do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An artificial intelligence-generated image recognition method based on generative watermarking, characterized in that Including: Performing watermark extraction on a first image to obtain a first watermark information vector and a first image latent vector; Performing semantic parsing on the first watermark information vector to obtain first image semantic information and first non-image semantic information; Performing semantic parsing on the first image latent vector to obtain second image semantic information; In the case where the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, encoding the first image semantic information to obtain a prompt word vector, and generating a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information, and performing noise reduction processing on each of the noise images based on the prompt word vector to obtain a plurality of second images; In the case where the image similarity between the first image and each of the second images meets a preset image similarity condition, determining that the first image is an AI-generated image.

2. The method according to claim 1, wherein The generating a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information includes: Performing semantic classification on the first image semantic information to obtain a plurality of sub-semantic information; Based on each sub-semantic information, determining a corresponding set of synonymous texts for each sub-semantic information; Extracting one text from each of the sets of synonymous texts, and combining the extracted texts based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information to obtain a text-to-image prompt word; Based on the first non-image semantic information, determining the generation time and generation platform of the first image; Based on the generation time and the generation platform, determining noise information; Generating a plurality of noise images based on each of the text-to-image prompt words and the noise information.

3. The method according to claim 1, characterized in that, The performing noise reduction processing on each of the noise images based on the prompt word vector to obtain a plurality of second images includes: Encoding the noise image using an image encoder to obtain an initial image latent vector; Inputting the initial image latent vector and the prompt word vector into a noise reduction network to obtain a target image latent vector output by the noise reduction network; Based on second non-image semantic information related to the noise information in the noise image and the first image semantic information, determining watermark text information, and encoding the watermark text information to obtain a second watermark information vector; Using a feature fusion network to perform feature fusion on the target image latent vector and the second watermark information vector to generate the second image.

4. The method according to claim 3, wherein The noise reduction network includes N upsampling layers connected in sequence and N downsampling layers connected in sequence, and the N upsampling layers are connected to the N downsampling layers through an intermediate layer; When the upsampling layer samples the input image latent vector, based on the prompt word vector and the sampling time step embedding vector corresponding to the upsampling layer, performing upsampling on the input image latent vector through a cross-attention mechanism to obtain the image latent vector output by the upsampling layer; When sampling the input image latent vector in the downsampling layer, based on the prompt vector and the sampling time step embedding vector corresponding to the upsampling layer, the input image latent vector is downsampled through a cross-attention mechanism to obtain the image latent vector output by the downsampling layer.

5. The method according to claim 1, wherein It further includes: When the semantic similarity between the first image semantic information and the second image semantic information meets the first semantic similarity condition, calculate the similarity between the first image semantic information and the third image semantic information corresponding to each text-image prompt word in the preset historical text-to-image prompt word set. When the similarity between the first image semantic information and the third image semantic information meets the second semantic similarity condition, determine that the first image is an AI-generated image.

6. The method according to claim 5, wherein It further includes: At a preset frequency, obtain the historical text-to-image prompt words in each AI generation platform from each AI generation platform. Based on the historical text-to-image prompt words, update the historical text-to-image prompt word set.

7. An artificial intelligence generated image recognition device based on generative watermarking, characterized in that, It includes: A watermark extraction module for extracting a watermark from the first image to obtain a first watermark information vector and a first image latent vector. A first semantic parsing module for semantically parsing the first watermark information vector to obtain first image semantic information and first non-image semantic information. A second semantic parsing module for semantically parsing the first image latent vector to obtain second image semantic information. An image generation module for, when the semantic similarity between the first image semantic information and the second image semantic information does not meet the preset first semantic similarity condition, encoding the first image semantic information to obtain a prompt vector, and generating multiple noise images based on the first image semantic information and the noise information related to the first non-image semantic information, and performing noise reduction processing on each of the noise images based on the prompt vector to obtain multiple second images. A first image recognition module for determining that the first image is an AI-generated image when the image similarity between the first image and each of the second images meets the preset image similarity condition.

8. The device according to claim 7, characterized in that, The image generation module includes: A semantic classification unit for semantically classifying the first image semantic information to obtain multiple sub-semantic information. A text set determination unit for determining the synonymous text set corresponding to each sub-semantic information based on each sub-semantic information. A text combination unit for extracting one text from each of the synonymous text sets and combining the extracted texts based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information to obtain a text-to-image prompt word. A first information determination unit for determining the generation time and generation platform of the first image based on the first non-image semantic information. A second information determination unit for determining noise information based on the generation time and the generation platform. A noise image generation unit for generating multiple noise images based on each of the text-to-image prompt words and the noise information.

9. An artificial intelligence generated image recognition system based on generative watermarking, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-6.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Image digital watermark embedding method, image digital watermark extracting method and image digital watermark embedding system

    CN118632012A

  • Image category judgment method and device, equipment and storage medium

    CN119540651A