Artificial intelligence generated image recognition method and system based on generative watermark
By performing watermark extraction, semantic analysis and noise reduction processing on the artificial intelligence generated images, we can identify whether the image is generated by artificial intelligence, which solves the problem of difficult to identify high-quality artificial intelligence generated images in the prior art, and achieves higher recognition accuracy.
Patent Information
- Application Number
- CN202510413329.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The prior art is difficult to effectively identify and identify images generated by artificial intelligence, especially the quality of these images is high and can almost reach the level of being fake and real.
Using an artificial intelligence generated image recognition method based on generative watermarks, multiple noise images are generated by watermark extraction, semantic analysis, encoding and noise reduction processing on the image, and image similarity analysis is used to determine whether the image is an artificial intelligence generated image.
It improves the recognition accuracy of artificial intelligence-generated images, and can effectively identify artificial intelligence-generated images to avoid missed inspections.
Smart Images

Figure CN119919783A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image technology, and in particular to an artificial intelligence generated image recognition method and system based on generative watermark. Background Art
[0002] With the rapid development of Generative Artificial Intelligence (GAI), the quality of AI-generated images is getting higher and higher, almost to the point of being indistinguishable from the real thing. The proliferation of these AI-generated images has brought about a series of adverse effects after being widely disseminated. Therefore, there is an urgent need for methods to effectively identify or authenticate AI-generated images. Summary of the invention
[0003] The present invention provides an artificial intelligence generated image recognition method and system based on generative watermark, which can solve at least one of the above technical problems.
[0004] According to one aspect of the present invention, there is provided an artificial intelligence generated image recognition method based on generative watermark, comprising: Extracting a watermark from the first image to obtain a first watermark information vector and a first image latent vector; Performing semantic analysis on the first watermark information vector to obtain first image semantic information and first non-image semantic information; Performing semantic analysis on the first image latent vector to obtain semantic information of the second image; When the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, the first image semantic information is encoded to obtain a prompt word vector, and a plurality of noise images are generated based on the first image semantic information and noise information related to the first non-image semantic information, and each of the noise images is subjected to denoising processing based on the prompt word vector to obtain a plurality of second images; When the image similarity between the first image and each of the second images meets a preset image similarity condition, it is determined that the first image is an artificial intelligence generated image.
[0005] According to another aspect of the present invention, there is provided an artificial intelligence generated image recognition device based on generative watermark, comprising: A watermark extraction module, used to extract the watermark from the first image to obtain a first watermark information vector and a first image latent vector; A first semantic parsing module, configured to perform semantic parsing on the first watermark information vector to obtain first image semantic information and first non-image semantic information; A second semantic parsing module, configured to perform semantic parsing on the first image latent vector to obtain second image semantic information; an image generation module, configured to encode the first image semantic information to obtain a prompt word vector when the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, and to generate a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information, and to perform noise reduction processing on each of the noise images based on the prompt word vector to obtain a plurality of second images; The first image recognition module is used to determine that the first image is an artificial intelligence generated image when the image similarity between the first image and each of the second images meets a preset image similarity condition.
[0006] According to another aspect of the present invention, there is provided an artificial intelligence generated image recognition method based on generative watermark, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute any artificial intelligence generated image recognition method based on generative watermark as described in any embodiment of the present invention.
[0007] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the artificial intelligence generated image recognition method based on generative watermark as described in any one of the embodiments of the present invention.
[0008] According to the technical solution of the present invention, a watermark is extracted from the first image to be detected to obtain a first watermark information vector and a first image latent vector. The first watermark information vector is semantically parsed to obtain the first image semantic information and the first non-image semantic information. The first image latent vector is semantically parsed to obtain the second image semantic information. If the semantic similarity between the first image semantic information and the second image semantic information does not meet the preset semantic similarity condition, it means that the semantics expressed by the watermark in the image does not match the semantics of the image itself. At this time, the first image needs to be further judged to determine whether it is an artificial intelligence generated image. The image semantic information in the watermark information is used to generate a prompt word vector, and multiple noise images are generated based on the image semantic information in the watermark information and the noise information related to the non-image semantic information in the watermark information. The noise reduction process is performed on each noise image based on the prompt word vector, so that an artificial intelligence image related to the watermark information, that is, a second image, can be generated. In this way, an artificial intelligence image strongly related to the semantic information in the watermark information can be generated. Thus, it is determined whether the image similarity between the first image and each second image meets the preset similarity condition. If it does, the first image is determined to be an artificial intelligence image. Thereby, the recognition accuracy of AI-generated images can be improved.
[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention. Figure 1 is a flow chart of an artificial intelligence generated image recognition method based on generative watermark according to an embodiment of the present invention; Figure 2 is a flow chart of a watermark image parsing process according to an embodiment of the present invention; Figure 3 Flowchart of the artificial intelligence generation process of a watermark image according to an embodiment of the present invention Figure 4 It is a structural block diagram of an artificial intelligence generated image recognition device based on generative watermark according to an embodiment of the present invention; Figure 5 The block diagram is a block diagram of an electronic device for implementing the method according to the embodiment of the present invention. DETAILED DESCRIPTION
[0011] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0012] Figure 1 It is a flowchart of an artificial intelligence generated image recognition method based on generative watermark according to an embodiment of the present invention.
[0013] like Figure 1 As shown, the artificial intelligence generated image recognition method based on generative watermark may include: S110, extracting a watermark from the first image to obtain a first watermark information vector and a first image latent vector; S120, performing semantic analysis on the first watermark information vector to obtain first image semantic information and first non-image semantic information; S130, performing semantic analysis on the first image latent vector to obtain second image semantic information; S140, when the semantic similarity between the first image semantic information and the second image semantic information does not meet the preset first semantic similarity condition, encode the first image semantic information to obtain a prompt word vector, generate multiple noise images based on the first image semantic information and noise information related to the first non-image semantic information, and perform noise reduction processing on each noise image based on the prompt word vector to obtain multiple second images; S150: When the image similarity between the first image and each of the second images meets a preset image similarity condition, determine that the first image is an artificial intelligence generated image.
[0014] Exemplarily, the first image is an image with a plain text watermark or an image with a dark watermark. If a watermark information vector can be obtained by extracting the watermark from the first image, it means that the first image is an image with a plain text watermark or an image with a dark watermark. The first image can also be called a watermark image.
[0015] Exemplarily, a pre-trained watermark extraction network can be used to extract watermarks from the first image. The first image is input into the watermark extraction network, and the image encoder in the watermark extraction network converts the first image into a 1024-dimensional watermark image latent vector and upgrades the dimension to a 1792-dimensional watermark image latent vector. Then, the network parameters of the watermark extraction network are used to perform feature splitting on the 1792-dimensional watermark image latent vector to obtain a first watermark information vector and a first image latent vector.
[0016] Exemplarily, the above feature splitting process can be expressed by the following formula: v_w_watermark = (v_w_combined − α·v_w_image) ÷ (1−α); Among them, v_w_watermark represents the first watermark information vector, v_w_combined represents the watermark image latent vector, α represents the network parameters of the watermark extraction network, and v_w_image represents the first image latent vector.
[0017] For example, Figure 2 As shown, the watermarked image to be detected is input into the watermark extraction network to obtain a watermark text vector, also called a watermark information vector. If the watermark information vector is encrypted information, the watermark information can be decrypted. The watermark text vector is decoded using a watermark text decoder to obtain a watermark text sequence. The first image semantic information can be extracted from the watermark text sequence. Thus, based on the first image semantic information, the image recognition process of the above steps is performed until the identification result of whether the watermarked image is an artificial intelligence generated image is obtained.
[0018] Exemplarily, a pre-trained semantic recognition network can be used to perform semantic analysis on the first watermark information vector and the first image latent vector to obtain the first image semantic information and the second image semantic information. The first watermark information vector includes non-image semantic information in addition to the image semantic information. The non-image semantic information may include the generation platform, generation time and generation method of the first image or some unrecognizable noise information.
[0019] Exemplarily, the first image semantic information refers to the semantic information in the watermark information in the first image to be identified, and the second image semantic information refers to the semantic information in the image information in the first image to be identified.
[0020] Exemplarily, image semantic information may include multiple categories of information, such as object categories, scene descriptions, attribute information, spatial position relationships, and emotional intent. For example, object categories may be objects such as people, animals, or objects included in an image. Scene descriptions may be the overall scene of an image. Attribute information may include the color, shape, and motion of an object. Spatial relationships may include positional relationships between objects. Emotional intent may include the emotions or intent conveyed by an image.
[0021] Exemplarily, the first semantic similarity condition may be that the semantic similarity between the semantic information of the first image and the semantic information of the second image is less than a preset semantic similarity threshold. For example, the semantic similarity between the two is less than 80%, which means that the semantic information in the watermark information in the image to be detected has a low match or does not match the semantic information in the image information in the image, indicating that the image to be detected is likely not to be an AI-generated image. However, in order to avoid missing the image, it is necessary to further identify whether the image to be detected is an AI-generated image.
[0022] It can be understood that the first image is the image to be detected, and the second image is a reference image for detecting the image to be detected. The second image is an image with a plain text watermark or an image with a dark watermark. Therefore, the second image can also be called a watermark reference image.
[0023] Exemplarily, a text encoder is used to encode the semantic information of the first image to obtain a prompt word vector.
[0024] Exemplarily, each sub-semantic information in the first image semantic information is subjected to synonym mutation and then combined to obtain multiple text-image prompt words, each text-image prompt word is encoded using a text encoder to obtain each text-image prompt word vector, a noise vector is generated based on the platform information and time information in the first non-image semantic information, and multiple noise images are generated based on each text-image prompt word vector and the noise vector.
[0025] For example, the image similarity condition may be that the total number of second images having a similarity with the first image greater than a preset threshold is greater than a preset number threshold. For another example, the similarity between the first image and each second image is greater than a preset similarity threshold.
[0026] According to the above embodiment, the similarity between the semantic information in the watermark information in the first image and the semantic information in the image information in the image is preliminarily judged. If it is determined that the similarity between the semantic information in the watermark information in the first image and the semantic information in the image information in the image does not meet the preset similarity condition, the semantic information of the first image is encoded to obtain a prompt word vector, and multiple noise images are generated based on the first image semantic information and the noise information related to the first non-image semantic information. Based on the prompt word vector, each noise image is subjected to denoising respectively to obtain multiple second images. In this way, since the prompt word vector is generated by the semantic information in the watermark information in the image to be detected, the prompt word vector is used to perform denoising on the noise image, so that the image after denoising can retain the image content related to the semantic information in the watermark information as much as possible. In this way, the accuracy of the reference image used for detecting the image to be detected can be improved. Furthermore, the first image and each second image used as a reference are judged again. If the similarity between the first image and each second image used as a reference is greater than the preset image similarity condition, the first image is determined to be an artificial intelligence generated image. Therefore, this example can be used to accurately determine whether the watermark image to be detected is an artificial intelligence generated image, thereby avoiding missed detection.
[0027] In one embodiment, multiple noise images are generated based on first image semantic information and noise information related to first non-image semantic information, including: semantically classifying the first image semantic information to obtain multiple sub-semantic information; determining, based on each sub-semantic information, a set of synonymous texts corresponding to each sub-semantic information; extracting a text from each synonymous text set, and combining the extracted texts based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information to obtain a text image prompt word; determining the generation time and generation platform of the first image based on the first non-image semantic information; determining the noise information based on the generation time and generation platform; and generating multiple noise images based on each text image prompt word and the noise information.
[0028] Exemplarily, image semantic information may include multiple categories of information, such as object categories, scene descriptions, attribute information, spatial position relationships, and emotional intent. For example, object categories may be objects such as people, animals, or objects included in an image. Scene descriptions may be the overall scene of an image. Attribute information may include the color, shape, and motion of an object. Spatial relationships may include positional relationships between objects. Emotional intent may include the emotions or intent conveyed by an image.
[0029] For example, the semantic information semantic_info is as follows: { "object":["cat","sofa"]; "scene": ["living room"]; "status":["orange cat", "cat is scratching", "blue sofa"]; “relation”:[“the cat is on the sofa”]; “emotion”: [“joy”] } Exemplarily, the sub-semantic information may be a short text information and corresponds to one of the above categories.
[0030] For example, a synonymous text set may include multiple short texts with the same or similar meanings.
[0031] Exemplarily, a text is extracted from each synonymous text set, and the extracted texts are combined based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information to obtain a text image prompt word. In this way, multiple text image prompt words can be generated in this way. Moreover, the image semantic information expressed by these multiple text image prompt words is the same or highly similar.
[0032] Exemplarily, based on the generation time and the generation platform, the noise information of the image generated by the platform in the time period at the time is obtained. For example, the historical images generated by the platform in the time period at the time can be used to determine their common noise from the historical images, and the common noise is used as the noise information.
[0033] Exemplarily, a text encoder is used to encode the text image prompt word and noise information respectively to obtain a text image prompt word vector and a noise vector; an image generation network is used to process the text image prompt word vector and the noise vector to obtain a noise image.
[0034] According to the above-mentioned implementation, semantic classification is performed on the first image semantic information to obtain multiple sub-semantic information; based on each sub-semantic information, a synonymous text set corresponding to each sub-semantic information is determined; a text is extracted from each synonymous text set, and based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information, the extracted texts are combined to obtain a text image prompt word. In this way, multiple text image prompt words representing the same image semantics can be generated. Based on the first non-image semantic information, the generation time and generation platform of the first image are determined; based on the generation time and generation platform, noise information is determined. In this way, noise information related to the generation time and generation platform of the image to be detected can be obtained. Furthermore, based on multiple text image prompt words representing the same image semantics as the image to be detected, and noise information related to the generation time and generation platform of the image to be detected, multiple noise images highly correlated with the image to be detected can be generated.
[0035] In one embodiment, based on the prompt word vector, each noise image is subjected to denoising processing to obtain multiple second images, including: encoding the noise image using an image encoder to obtain an initial image latent vector; inputting the initial image latent vector and the prompt word vector into a denoising network to obtain a target image latent vector output by the denoising network; determining watermark text information based on the second non-image semantic information related to the noise information in the noise image and the first image semantic information, and encoding the watermark text information to obtain a second watermark information vector; using a feature fusion network to perform feature fusion on the target image latent vector and the second watermark information vector to generate a second image. The noise image is input into an image encoder, the convolution layer in the image encoder extracts features from the noise map to obtain a feature map, the global average pooling layer performs average pooling on the feature map to obtain a feature vector, and the fully connected layer maps the feature vector to a 1024-dimensional space to obtain an initial image latent vector; Exemplarily, the image encoder is an encoder generated by training the convolutional neural network ResNet50. The image encoder includes a convolutional layer, a global average pooling layer, and a fully connected layer.
[0036] Exemplarily, a noisy image is input into an image encoder, a convolutional layer in the image encoder extracts image features from the noisy image to obtain a feature map, a global average pooling layer performs pooling on the feature map to obtain a feature vector, and a fully connected layer maps the feature vector to a 1024-dimensional space to obtain a 1024-dimensional initial image latent vector.
[0037] Exemplarily, in the denoising network, based on the image semantics expressed by the prompt word vector, the features related to the image semantics are gradually sampled from the initial image latent vector to obtain a target image latent vector, which has removed noise information and retains features that are the same or similar to the image semantics expressed by the prompt word vector.
[0038] Exemplarily, the denoising network can be based on a UNet network, and a time step embedding and a conditional mechanism component is referenced for each sampling layer in the UNet network.
[0039] Among them, the UNet network can be composed of 4 downsampling layers, 1 intermediate layer, 4 upsampling layers and 1 output layer. The downsampling layer is composed of multiple convolution blocks, and outputs the feature image extracted by each convolution block. The upsampling layer is composed of a transposed convolution layer component, which is used to output the gradually restored feature map. The output layer is a fully connected layer, which is used to splice the feature maps output by the upsampling layer and the downsampling layer in the channel dimension to gradually generate the detailed information of the image.
[0040] The time step embedding component uses sinusoidal position encoding to convert the time step t representing the current denoising stage into a vector, maps the time step embedding vector to the corresponding feature dimension through a fully connected layer, and adds the mapped time step embedding vector to the feature map of each convolutional block of the UNet network.
[0041] The conditional mechanism component is used to add the original prompt word vector representation to the feature map of each convolutional block of the UNet network through the cross-attention mechanism.
[0042] The UNet network is used to predict the potential representation of the image in the image denoising generation process. The main function of the time step embedding is to embed the current denoising stage into the network. The main function of the conditional mechanism is to input the prompt word vector into the network to control the content generated by the UNet network.
[0043] The operation process of the denoising network is to input the image latent vector v_image and the current time step t, that is, the stage of image denoising, convert the time step t into a time step embedding vector, inject the prompt word conditional information v_prompt into UNet through cross attention, and UNet gradually denoises to generate the final image. The output of the denoising network is the processed image information vector representation v_g_image.
[0044] It can be understood that since the target image latent vector does not include watermark information, in order to generate a reference watermark image of the same category as the watermark image to be detected, it is necessary to determine the watermark information, and then perform feature fusion on the watermark information vector and the target image latent vector. In this way, a reference watermark image can be generated.
[0045] Exemplarily, based on the noise information in the noise image, candidate noise information matching the noise information is searched, the platform information and time information mapped by the candidate noise information are determined as the second non-image semantic information, and the second non-image semantic information is merged with the first image semantic information into watermark text information.
[0046] For example, the merged watermark text information is: { "platform":"p1", "time":"2025-01-01 12:00:00", "semantic_info":{ "object":["cat","sofa"], “scene”:[”living room”], "status":["orange cat","cat is scratching","blue sofa"], “relation”:[“The cat is on the sofa”], "emotion": ["joy"]} } Exemplarily, a feature fusion network is used to perform feature fusion on the target image latent vector and the second watermark information vector to generate a watermark image latent vector, and a watermark image decoder is used to decode and reconstruct the watermark image latent vector to obtain a second image.
[0047] Exemplarily, the feature fusion network is used to perform feature fusion operations on the 1024-dimensional target image v_g_image and the 768-dimensional second watermark information vector v_watermark, including feature weighted concatenation and dimensionality reduction operations, so that the features of the generated image carry the watermark information.
[0048] In the feature dimension, the target image latent vector v_g_image and the second watermark information vector v_watermark are weightedly fused to form a high-dimensional vector, where the target image latent vector v_g_image is the output of the denoising network, and its feature dimension is 1024 dimensions, and the second watermark information vector v_watermark is one of the outputs of the text encoder, and its feature dimension is 768 dimensions. The watermark image latent vector obtained after fusion is 1792 dimensions.
[0049] Exemplarily, the formula for weighted fusion of the feature fusion network may be: v_combined =α·v_g_image + (1−α)·v_watermark The initial value of α may be 0.98, so as to minimize the influence of the spliced watermark information latent vector on the image content and quality during image reconstruction.
[0050] Finally, a fully connected layer is used to map the concatenated watermark image latent vector v_combined to a space with the same dimension as the target image latent vector v_g_image, that is, 1024 dimensions, to obtain the final watermark image latent vector.
[0051] Exemplarily, the watermark image decoding is used to decode and reconstruct the watermark image latent vector v_combined obtained after the fusion of the image and the watermark feature, so as to obtain an artificial intelligence generated image with a prompt word information watermark.
[0052] Exemplarily, the network structure of the watermark image decoder may include a fully connected layer and a plurality of transposed convolutional layers.
[0053] Among them, the fully connected layer is used to map the watermark image latent vector v_combined into an initial feature map suitable for transposed convolution. For multiple transposed convolution layers, batch normalization and ReLU activation functions are added after each transposed convolution layer. The last transposed convolution layer uses the Sigmoid activation function to limit the pixels within a reasonable range, thereby gradually upsampling the feature map, restoring the image resolution, and generating the final AI-generated image w_image.
[0054] like Figure 3 As shown in the figure, the text processing network processes the input prompt word to obtain image semantic information. The text encoder encodes the image semantic information and the text information composed of image semantic information, platform information and time information respectively to obtain the prompt word text vector and watermark text vector. The prompt word text vector is processed by the latent space processing network to generate the image latent vector. The feature fusion network fuses the image latent vector with the watermark text vector to obtain the watermark image latent vector. The watermark image decoder is used to decode and reconstruct the watermark image latent vector to obtain the reference watermark image generated by artificial intelligence.
[0055] According to the above-mentioned implementation, the semantic information of the first image is encoded to obtain a prompt word vector, and based on the semantic information of the first image and the noise information related to the first non-image semantic information, multiple noise images are generated, and based on the prompt word vector, each noise image is subjected to denoising respectively to obtain multiple second images. In this way, since the prompt word vector is generated by the semantic information in the watermark information in the image to be detected, the prompt word vector is used to perform denoising on the noise image, so that the image after denoising can retain the image content related to the semantic information in the watermark information as much as possible. In this way, the accuracy of the reference image used to detect the image to be detected can be improved.
[0056] In one embodiment, the denoising network includes N upsampling layers connected in sequence and N downsampling layers connected in sequence, and the N upsampling layers are connected to the N downsampling layers through an intermediate layer; when the upsampling layer samples the input image latent vector, the input image latent vector is upsampled through a cross-attention mechanism based on the prompt word vector and the sampling time step embedding vector corresponding to the upsampling layer to obtain the image latent vector output by the upsampling layer; when the downsampling layer samples the input image latent vector, the input image latent vector is downsampled through a cross-attention mechanism based on the prompt word vector and the sampling time step embedding vector corresponding to the upsampling layer to obtain the image latent vector output by the downsampling layer.
[0057] Exemplarily, the input image latent vector of the first downsampling layer is the initial image latent vector. The output image latent vector of the last downsampling layer is used as the input image latent vector of the first upsampling layer. The output image latent vector of the last upsampling layer is the target image latent vector.
[0058] Exemplarily, sinusoidal position coding is used to convert the time step t representing the denoising stage of the current sampling layer into a vector to obtain a sampling time step embedding vector.
[0059] Exemplarily, based on the residual information between the i-th downsampling layer and the N-i+1-th downsampling layer, the sampling time step embedding vector of the i-th downsampling layer and the sampling time step embedding vector of the N-i+1-th downsampling layer are adjusted, and the adjusted sampling time step embedding vector is added to the corresponding i-th downsampling layer and the N-i+1-th downsampling layer. The synchronization of the time step can be determined using the residual information, and the sampling time step embedding vectors of the two sampling layers are adjusted using the synchronization difference information between the two corresponding upper and lower sampling layers so that the delay time between them can be eliminated and relative synchronization can be achieved.
[0060] According to the above implementation, the prompt word vector and the sampling time step embedding vector are injected into the corresponding sampling layer through the cross attention mechanism. The sampling layer can determine the sampling direction based on the injected information. The image features sampled according to such sampling direction match the prompt word and the time step, thereby gradually sampling the image latent vector that meets the target requirements and removing the noise information.
[0061] In one embodiment, the above method also includes: when the semantic similarity between the first image semantic information and the second image semantic information meets the first semantic similarity condition, calculating the similarity between the first image semantic information and the third image semantic information corresponding to each text image prompt word in a preset historical text image prompt word set; when the similarity between the first image semantic information and the third image semantic information meets the second semantic similarity condition, determining that the first image is an artificial intelligence generated image.
[0062] In this example, when the semantic similarity between the semantic information of the first image and the semantic information of the second image meets the first semantic similarity condition, the first image can also be directly determined to be an image generated by artificial intelligence. However, in order to avoid misjudgment, a preset set of historical cultural image prompt words can be further used to determine whether the semantic information in the watermark information in the watermark image to be detected matches the semantic information in the cultural image prompt words used historically on major platforms. If they match, the first image can be determined to be an image generated by artificial intelligence. In this way, the recognition accuracy of artificial intelligence-generated images can be improved.
[0063] In one embodiment, the above method also includes: obtaining historical text image prompt words in the artificial intelligence generation platform from each artificial intelligence generation platform according to a preset frequency; and updating the historical text image prompt word set based on the historical text image prompt words.
[0064] For example, the preset frequency may be every 1 hour, every 12 hours, etc.
[0065] Exemplarily, if the historical text image prompt word set does not include the historical text image prompt word, then the historical text image prompt word is added to the historical text image prompt word set. If the historical text image prompt word set includes the historical text image prompt word.
[0066] According to the above implementation, the historical text image prompt word set can be updated in a timely manner, so as to improve the recognition accuracy when recognizing the artificial intelligence generated image later.
[0067] Figure 4 It is a structural block diagram of an artificial intelligence generated image recognition device based on generative watermark according to an embodiment of the present invention.
[0068] like Figure 4 As shown, the artificial intelligence generated image recognition device based on generative watermark includes: A watermark extraction module 410 is used to extract a watermark from the first image to obtain a first watermark information vector and a first image latent vector; A first semantic analysis module 420, configured to perform semantic analysis on the first watermark information vector to obtain first image semantic information and first non-image semantic information; A second semantic parsing module 430, configured to perform semantic parsing on the first image latent vector to obtain second image semantic information; The image generation module 440 is used to encode the first image semantic information to obtain a prompt word vector when the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, and generate a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information, and perform noise reduction processing on each of the noise images based on the prompt word vector to obtain a plurality of second images; The first image recognition module 450 is used to determine that the first image is an artificial intelligence generated image when the image similarity between the first image and each of the second images meets a preset image similarity condition.
[0069] In one implementation, the image generation module 440 includes: A semantic classification unit, used for performing semantic classification on the first image semantic information to obtain a plurality of sub-semantic information; A text set determination unit, used for determining, based on each sub-semantic information, a synonymous text set corresponding to each sub-semantic information; A text combination unit, used to extract a text from each of the synonymous text sets, and combine the extracted texts based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information to obtain a text image prompt word; A first information determining unit, configured to determine a generation time and a generation platform of the first image based on the first non-image semantic information; A second information determining unit, configured to determine noise information based on the generation time and the generation platform; The noise image generating unit is used to generate a plurality of noise images based on each of the text image prompt words and the noise information.
[0070] In one implementation, the image generation module 340 includes: an initial image vector determining unit, configured to encode the noisy image using an image encoder to obtain an initial image latent vector; A target image vector determination unit, configured to input the initial image latent vector and the prompt word vector into a denoising network to obtain a target image latent vector output by the denoising network; a watermark vector determining unit, configured to determine watermark text information based on second non-image semantic information related to noise information in the noise image and the first image semantic information, and encode the watermark text information to obtain a second watermark information vector; A feature fusion unit is used to use a feature fusion network to perform feature fusion on the target image latent vector and the second watermark information vector to generate the second image.
[0071] In one embodiment, the denoising network includes N upsampling layers connected in sequence and N downsampling layers connected in sequence, and the N upsampling layers are connected to the N downsampling layers through an intermediate layer; When the upsampling layer samples the input image latent vector, based on the prompt word vector and the sampling time step embedding vector corresponding to the upsampling layer, the input image latent vector is upsampled through a cross attention mechanism to obtain the image latent vector output by the upsampling layer; When the downsampling layer samples the input image latent vector, the input image latent vector is downsampled through a cross-attention mechanism based on the prompt word vector and the sampling time step embedding vector corresponding to the upsampling layer to obtain the image latent vector output by the downsampling layer.
[0072] In one embodiment, the above device further comprises: a similarity calculation module, configured to calculate the similarity between the first image semantic information and the third image semantic information corresponding to each text image prompt word in a preset historical text image prompt word set, when the semantic similarity between the first image semantic information and the second image semantic information meets a first semantic similarity condition; The second image recognition module is used to determine that the first image is an artificial intelligence generated image when the similarity between the first image semantic information and the third image semantic information meets a second semantic similarity condition.
[0073] In one embodiment, the above device further comprises: A historical prompt word acquisition module is used to obtain the historical text image prompt words in the artificial intelligence generation platform from each artificial intelligence generation platform according to a preset frequency; The prompt word set generation module is used to update the historical text image prompt word set based on the historical text image prompt words.
[0074] For the description of specific functions and examples of each module and submodule of the system in the embodiment of the present invention, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0075] In the technical solution of the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0076] According to an embodiment of the present invention, the present invention also provides a system and a readable storage medium.
[0077] Figure 5 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0078] like Figure 5 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 to a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0079] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0080] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as an artificial intelligence generated image recognition method based on a generative watermark. For example, in some embodiments, an artificial intelligence generated image recognition method based on a generative watermark may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the artificial intelligence generated image recognition method based on the generative watermark described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute an artificial intelligence generated image recognition method based on a generative watermark in any other appropriate manner (eg, by means of firmware).
[0081] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0082] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0083] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0084] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0085] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0086] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0087] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and this document does not limit this.
[0088] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of the present invention should be included in the protection scope of the present invention.
Claims
1. An artificial intelligence generated image recognition method based on generative watermark, characterized in that: include: Extracting a watermark from the first image to obtain a first watermark information vector and a first image latent vector; Performing semantic analysis on the first watermark information vector to obtain first image semantic information and first non-image semantic information; Performing semantic analysis on the first image latent vector to obtain semantic information of the second image; When the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, the first image semantic information is encoded to obtain a prompt word vector, and a plurality of noise images are generated based on the first image semantic information and noise information related to the first non-image semantic information, and each of the noise images is subjected to denoising processing based on the prompt word vector to obtain a plurality of second images; When the image similarity between the first image and each of the second images meets a preset image similarity condition, it is determined that the first image is an artificial intelligence generated image.
2. The method according to claim 1, characterized in that The step of generating a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information comprises: Performing semantic classification on the first image semantic information to obtain a plurality of sub-semantic information; Based on each sub-semantic information, determining a synonymous text set corresponding to each sub-semantic information; Extracting a text from each of the synonymous text sets respectively, and combining the extracted texts based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information to obtain a text image prompt word; Determining, based on the first non-image semantic information, a generation time and a generation platform of the first image; Determining noise information based on the generation time and the generation platform; Based on the respective text image prompt words and the noise information, a plurality of noise images are generated.
3. The method according to claim 1, characterized in that The step of performing noise reduction processing on each of the noise images based on the prompt word vector to obtain a plurality of second images includes: Encoding the noisy image using an image encoder to obtain an initial image latent vector; Inputting the initial image latent vector and the prompt word vector into a denoising network to obtain a target image latent vector output by the denoising network; Determining watermark text information based on second non-image semantic information related to noise information in the noise image and the first image semantic information, and encoding the watermark text information to obtain a second watermark information vector; A feature fusion network is used to perform feature fusion on the target image latent vector and the second watermark information vector to generate the second image.
4. The method according to claim 3, characterized in that The denoising network includes N upsampling layers connected in sequence and N downsampling layers connected in sequence, wherein the N upsampling layers are connected to the N downsampling layers through an intermediate layer; When the upsampling layer samples the input image latent vector, based on the prompt word vector and the sampling time step embedding vector corresponding to the upsampling layer, the input image latent vector is upsampled through a cross attention mechanism to obtain the image latent vector output by the upsampling layer; When the downsampling layer samples the input image latent vector, the input image latent vector is downsampled through a cross-attention mechanism based on the prompt word vector and the sampling time step embedding vector corresponding to the upsampling layer to obtain the image latent vector output by the downsampling layer.
5. The method according to claim 1, characterized in that Also includes: When the semantic similarity between the first image semantic information and the second image semantic information meets the first semantic similarity condition, calculating the similarity between the first image semantic information and the third image semantic information corresponding to each text image prompt word in the preset historical text image prompt word set; When the similarity between the first image semantic information and the third image semantic information meets a second semantic similarity condition, it is determined that the first image is an artificial intelligence generated image.
6. The method according to claim 5, characterized in that Also includes: Obtaining historical cultural image prompt words in the artificial intelligence generation platform from each artificial intelligence generation platform according to a preset frequency; Based on the historical text image prompt words, the historical text image prompt word set is updated.
7. An artificial intelligence generated image recognition device based on generative watermark, characterized in that: include: A watermark extraction module, used to extract the watermark from the first image to obtain a first watermark information vector and a first image latent vector; A first semantic parsing module, configured to perform semantic parsing on the first watermark information vector to obtain first image semantic information and first non-image semantic information; A second semantic parsing module, configured to perform semantic parsing on the first image latent vector to obtain second image semantic information; an image generation module, configured to encode the first image semantic information to obtain a prompt word vector when the semantic similarity between the first image semantic information and the second image semantic information does not meet a preset first semantic similarity condition, and to generate a plurality of noise images based on the first image semantic information and noise information related to the first non-image semantic information, and to perform noise reduction processing on each of the noise images based on the prompt word vector to obtain a plurality of second images; The first image recognition module is used to determine that the first image is an artificial intelligence generated image when the image similarity between the first image and each of the second images meets a preset image similarity condition.
8. The device according to claim 7, characterized in that The image generation module comprises: A semantic classification unit, used for performing semantic classification on the first image semantic information to obtain a plurality of sub-semantic information; A text set determination unit, used for determining, based on each sub-semantic information, a synonymous text set corresponding to each sub-semantic information; A text combination unit, used to extract a text from each of the synonymous text sets, and combine the extracted texts based on the arrangement order of the sub-semantic information corresponding to each text in the first image semantic information to obtain a text image prompt word; A first information determining unit, configured to determine a generation time and a generation platform of the first image based on the first non-image semantic information; A second information determining unit, configured to determine noise information based on the generation time and the generation platform; The noise image generating unit is used to generate a plurality of noise images based on each of the text image prompt words and the noise information.
9. An artificial intelligence generated image recognition system based on generative watermark, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1-6.
Citation Information
Patent Citations
Image digital watermark embedding method, image digital watermark extracting method and image digital watermark embedding system
CN118632012A
Image category judgment method and device, equipment and storage medium
CN119540651A
Cross-modal image-watermark joint generation and detection device and method thereof
US12125119B1