Image processing method and device, computer readable medium and computer equipment

By calculating the differential image between the sample image and its modified version, and training the image recognition model based on the differential image, the problem of difficult to recognize the generated images in the prior art is solved, and higher recognition accuracy and generalization capabilities are achieved.

CN120047731APending Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510107357.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify machine-generated images, especially when faced with images generated by new generation techniques.

Method used

By obtaining a differential image between the sample image and its modified version, and generating input data of the image recognition model based on the differential image, the model is trained to identify whether the sample image is a machine-generated image.

Benefits of technology

It realizes effective recognition of machine-generated images, and improves the generalization ability and recognition accuracy of image recognition models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047731A_ABST
    Figure CN120047731A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method and device, a computer readable medium and computer equipment. The image processing method comprises the following steps: acquiring a sample image, and modifying the sample image to obtain a modified image; calculating a difference image between the sample image and the modified image according to the sample image and the modified image; input data of a to-be-trained image recognition model is generated based on the difference image, a prediction result for the input data is output through the image recognition model, and the prediction result is used for representing whether the sample image is a machine generated image or not; and according to loss data between the prediction result and the label of the sample image, adjusting model parameters of the image recognition model to obtain a trained image recognition model. According to the technical scheme provided by the embodiment of the invention, the machine generated image can be effectively recognized, and meanwhile, the image recognition model has relatively high generalization ability and recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer and communication technology, and in particular to an image processing method, apparatus, computer-readable medium and computer equipment. Background Art

[0002] With the rapid development of AI Generated Content (AIGC) technology, it is becoming increasingly important to identify whether an image is machine-generated. This technology is critical for protecting original works, screening high-quality materials, and collecting secondary creative materials. However, due to the rapid evolution of generation technology, the classification methods based on training of generation models proposed in related technologies are difficult to cope with images generated by new generation technologies. Therefore, how to effectively identify machine-generated images is a technical problem that needs to be solved urgently. Summary of the invention

[0003] The embodiments of the present application provide an image processing method, apparatus, computer-readable medium and computer equipment, which can effectively identify machine-generated images, and also enable the image recognition model to have strong generalization ability and recognition accuracy.

[0004] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0005] According to one aspect of an embodiment of the present application, there is provided an image processing method, comprising: acquiring a sample image, and a modified image obtained by modifying the sample image; calculating a differential image between the sample image and the modified image based on the sample image and the modified image; generating input data of an image recognition model to be trained based on the differential image, and outputting a prediction result for the input data through the image recognition model to be trained, wherein the prediction result is used to characterize whether the sample image is a machine-generated image; adjusting model parameters of the image recognition model to be trained based on loss data between the prediction result and a label of the sample image to obtain a trained image recognition model.

[0006] According to one aspect of an embodiment of the present application, there is provided an image processing method, comprising: acquiring an image to be identified; inputting the image to be identified into a trained image recognition model, and obtaining a recognition result for the image to be identified output by the trained image recognition model, wherein the recognition result is used to characterize whether the image to be identified is a machine-generated image; wherein the trained image recognition model is trained according to the image processing method described in the aforementioned embodiment.

[0007] According to one aspect of the embodiments of the present application, an image processing apparatus is provided, including: an acquisition unit configured to acquire a sample image and a modified image obtained by modifying the sample image; a calculation unit configured to calculate a difference image between the sample image and the modified image according to the sample image and the modified image; a processing unit configured to generate input data for an image recognition model to be trained based on the difference image, and output a prediction result for the input data through the image recognition model to be trained, where the prediction result is used to characterize whether the sample image is a machine-generated image; and a training unit configured to adjust model parameters of the image recognition model to be trained according to loss data between the prediction result and a label of the sample image to obtain a trained image recognition model.

[0008] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: extract signal-to-noise ratio data of the difference image to obtain a signal-to-noise ratio image corresponding to the difference image, and use the signal-to-noise ratio image as the input data for the image recognition model to be trained, so as to output a prediction result for the input data through the image recognition model to be trained.

[0009] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: extract a feature image corresponding to the difference image, where the feature image has the same size as the difference image; for each pixel point in the feature image, count pixel point features in a neighboring region of each pixel point to obtain a statistical feature corresponding to each pixel point; and generate a signal-to-noise ratio image corresponding to the difference image according to the statistical feature corresponding to each pixel point.

[0010] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: add a padding layer matching the size of the convolution kernel at the boundary of the difference image according to the size of the convolution kernel used to extract the feature image to obtain a processed difference image; and perform convolution processing on the difference image using the convolution kernel to obtain the feature image.

[0011] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: count a pixel mean and a standard deviation in a neighboring region of each pixel point; calculate a ratio between the pixel mean and the standard deviation, and use the ratio as the statistical feature corresponding to each pixel point.

[0012] In some embodiments of the present application, based on the foregoing solution, the image recognition model includes a feature extraction network and a classification network; the processing unit is configured to: use the signal-to-noise ratio image as the input data of the feature extraction network, and obtain the signal-to-noise ratio features output by the feature extraction network; generate the input data of the classification network according to the signal-to-noise ratio features, so as to output the prediction result through the classification network.

[0013] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: perform a fusion process on the signal-to-noise ratio features and the semantic features of the sample image to obtain fusion features, and use the fusion features as the input data of the classification network; or use the signal-to-noise ratio features as the input data of the classification network.

[0014] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: perform a normalization process on the semantic features to obtain processed semantic features; perform a first fusion process on the signal-to-noise ratio features and the processed semantic features to obtain the result of the first fusion process; perform a second fusion process on the result of the first fusion process and the processed semantic features to obtain the fusion features.

[0015] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: perform a regularization process with a norm of 1 on the semantic features to perform a normalization process on the semantic features to obtain the processed semantic features.

[0016] In some embodiments of the present application, based on the foregoing solution, the feature dimensions of the processed semantic features and the signal-to-noise ratio features both include: a width dimension, a height dimension, and a channel dimension; the processing unit is configured to: merge the width dimension feature and the height dimension feature in the signal-to-noise ratio features to convert the signal-to-noise ratio features into a two-dimensional signal-to-noise ratio feature matrix; merge the width dimension feature and the height dimension feature in the processed semantic features to convert the processed semantic features into a two-dimensional semantic feature matrix; perform a first multiplication on the two-dimensional signal-to-noise ratio feature matrix and the two-dimensional semantic feature matrix, and use the result of the first multiplication as the result of the first fusion process.

[0017] In some embodiments of the present application, based on the foregoing solution, the processing unit is configured to: perform a second multiplication on the result of the first fusion process and the two-dimensional semantic feature matrix to obtain the result of the second multiplication; convert the result of the second multiplication into a three-dimensional feature matrix to obtain the fusion features.

[0018] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: divide the training data into multiple batches, where each batch contains at least one sample image; calculate the loss data of each batch according to the loss data corresponding to each sample image, and adjust the model parameters of the image recognition model to be trained according to the loss data of each batch; if the model parameters of the image recognition model to be trained are adjusted according to the loss data of the multiple batches, then complete one round of training process of the image recognition model to be trained; if the training process of the set number of rounds is completed for the image recognition model to be trained, then determine that the trained image recognition model is obtained.

[0019] According to one aspect of the embodiments of the present application, there is provided an image processing apparatus, including: an acquisition unit configured to acquire an image to be recognized; a recognition unit configured to input the image to be recognized into the trained image recognition model to obtain a recognition result of the trained image recognition model for the image to be recognized, where the recognition result is used to characterize whether the image to be recognized is a machine-generated image; wherein, the trained image recognition model is trained according to the image processing method in the foregoing embodiments.

[0020] According to one aspect of the embodiments of the present application, there is provided a computer-readable medium having a computer program stored thereon, where the computer program, when executed by a processor, implements the image processing method as described in the foregoing embodiments.

[0021] According to one aspect of the embodiments of the present application, there is provided a computer device, including: one or more processors; a storage device for storing one or more computer programs, when the one or more computer programs are executed by the one or more processors, enabling the computer device to implement the image processing method as described in the foregoing embodiments.

[0022] According to one aspect of the embodiments of the present application, there is provided a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads and executes the computer program from the computer-readable storage medium, enabling the computer device to execute the image processing method provided in the above various alternative embodiments.

[0023] In the technical solutions provided by some embodiments of the present application, a difference image between a sample image and a modified image obtained by modifying the sample image can be calculated. Then, input data for a to-be-trained image recognition model is generated based on the difference image, and a prediction result for the input data is output through the image recognition model. Furthermore, based on the loss data between the prediction result and the label of the sample image, the model parameters of the to-be-trained image recognition model are adjusted to obtain a trained image recognition model. It can be seen that the technical solution of the embodiment of the present application enables capturing the subtle differences between the sample image and the modified image by calculating the difference image between them. Such subtle differences often contain the information introduced when modifying the image, which helps the image recognition model to more accurately identify machine-generated images. At the same time, in the embodiment of the present application, generating the input data for the to-be-trained image recognition model based on the difference image also enables the image recognition model to extract general feature representations based on the difference image, thereby enabling the image recognition model to have stronger generalization ability and still maintain a high recognition accuracy when facing newly emerging image generation technologies.

[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 The schematic diagram of the image processing process in the related art is shown;

[0026] Figure 2 The schematic diagram of an exemplary scenario to which the technical solution of the embodiment of the present application can be applied is shown;

[0027] Figure 3 The schematic diagram of an exemplary scenario to which the technical solution of the embodiment of the present application can be applied is shown;

[0028] Figure 4 The flowchart of the image processing method according to an embodiment of the present application is shown;

[0029] Figure 5 The flowchart of the image processing method according to an embodiment of the present application is shown;

[0030] Figure 6 The flowchart of the image processing method according to an embodiment of the present application is shown;

[0031] Figure 7 The schematic diagram of the principle according to an embodiment of the present application is shown;

[0032] Figure 8 The flowchart of the image processing method according to an embodiment of the present application is shown;

[0033] Figure 9 The flowchart of an image processing method according to an embodiment of the present application is shown;

[0034] Figure 10 The flowchart of a feature fusion process according to an embodiment of the present application is shown;

[0035] Figure 11 The schematic diagram of an exemplary scenario to which the technical solution of the embodiment of the present application can be applied is shown;

[0036] Figure 12 The block diagram of an image processing apparatus according to an embodiment of the present application is shown;

[0037] Figure 13 The block diagram of an image processing apparatus according to an embodiment of the present application is shown;

[0038] Figure 14 The schematic diagram of the structure of a computer system of a computer device suitable for implementing the embodiment of the present application is shown. Detailed implementation manners

[0039] Now, the exemplary embodiments will be described in a more comprehensive manner with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to these examples; on the contrary, these embodiments are provided so that the present application is more comprehensive and complete, and the concept of the exemplary embodiments is fully conveyed to those skilled in the art.

[0040] In addition, the features, structures, or characteristics described in the present application can be combined in any suitable manner in one or more embodiments. In the following description, there are many specific details so that the embodiments of the present application can be fully understood. However, those skilled in the art should be aware that when implementing the technical solution of the present application, not all the detailed features in the embodiments are required, one or more specific details can be omitted, or other methods, elements, devices, steps, etc. can be adopted.

[0041] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0042] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0043] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0044] It should be noted that: "a plurality of" mentioned in this article means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0045] It can be understood that before and during the collection of relevant data (such as image data, etc.) in this application, a prompt interface or pop-up window can be displayed, which is used to prompt the user that their relevant data is currently being collected, so that this application only starts to execute the relevant steps of obtaining relevant data after obtaining the confirmation operation issued by the user for this prompt interface or pop-up window. Otherwise (that is, when the confirmation operation issued by the user for this prompt interface or pop-up window is not obtained), the relevant steps of obtaining relevant data are ended, that is, relevant data is not obtained. In other words, all the data collected in this application is collected with the consent and authorization of the user, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.

[0046] Artificial intelligence content generation (AIGC) refers to various forms of content automatically generated through artificial intelligence technology, including but not limited to text, images, audio, and video, etc. With the development of deep learning and generative models (such as generative adversarial networks (GANs), variational autoencoders (VAEs), diffusion models, etc.), AIGC has become a very active and rapidly developing field. And with the rapid development of AIGC technology, it has become increasingly important to identify whether an image is machine-generated. This technology is of crucial significance for protecting original works, screening high-quality materials, and collecting secondary creation materials.

[0047] In the related art, generally, available natural images (such as captured images, images intercepted from videos, etc., which are non-machine-generated images) and machine-generated images are collected, and a conventional deep learning classification model (such as resnet50, swin transformer) is used to input the collected images into the deep learning classification model to train a classifier. During application, as Figure 1 shown, the original image to be recognized can be input into the trained deep learning classification model. The deep learning classification model includes a feature network for extracting image features and a classification network for outputting prediction results, and thus the trained deep learning classification model can be used to identify whether the original image is a machine-generated image.

[0048] Since the feature network in the related art generally uses a semantic recognition network, the extracted image features are semantic features, and the recognition of whether an image is a machine-generated image can be regarded as a recognition task related to distribution. This leads to a mismatch between the mentioned image features and the recognition task, and further results in poor recognition effects for cross-image generators. For example, the recognition effect for machine-generated images generated by SD1.4 is relatively good, but the recognition effect for machine-generated images generated by BigGan is relatively poor. At the same time, during the training process of the deep learning classification model, only the collected images are used as input, and the semantic information therein is too large, which masks the distribution information related to forged images, so that the model cannot effectively learn the information related to forgery through learning adjustment, and further leads to poor recognition effects of the model.

[0049] Based on this, the embodiments of the present application provide a new image processing method. The method can calculate a differential image between a sample image and a modified image obtained by modifying the sample image, and generate input data for an image recognition model to be trained based on the differential image, so that the model can capture the subtle differences between the sample image and the modified image, and such subtle differences often contain the information introduced when modifying the image, which helps the image recognition model to more accurately identify machine-generated images. At the same time, generating the input data for the image recognition model to be trained based on the differential image in the embodiments of the present application also enables the image recognition model to extract general feature representations based on the differential image, and further enables the image recognition model to have stronger generalization ability.

[0050] The technical solutions of the embodiments of the present application can be applied in various scenarios. For example, they can be applied in scenarios such as original work protection, medical image verification, detection of false news pictures, authenticity recognition of biometric images, etc.

[0051] As an example, in the scenario of authenticity recognition of biometric images, such as Figure 2As shown, the terminal device 210 can be a device for collecting palmprint images, and users can perform payment operations by swiping their palmprints. Specifically, when the palmprint image collection conditions are met, the terminal device 210 can collect the user's palmprint image; among them, the palmprint image collection conditions include but are not limited to: the user's order is confirmed, it is detected that the palmprint image entry is triggered, a registration operation of the palmprint image is detected, the camera captures a collectable palmprint image within the collection area, and so on.

[0052] The terminal device 210 can send the collected palmprint image to the server 220 so that the server 220 can verify the palmprint image sent by the terminal device 210 based on the registered palmprint information stored in the database; or the terminal device 210 can also verify the collected palmprint image based on its own database. If the server 220 determines that the palmprint image sent by the terminal device 210 has a high similarity with the registered palmprint information stored in the database (such as the similarity is greater than or equal to 95% etc.), then the account associated with the registered palmprint information can be determined accordingly, and then the payment can be deducted from this account to complete the operation of palmprint payment. After that, the server 220 can return the result information of palmprint payment to the terminal device 210.

[0053] It should be noted that when verifying the palmprint image, the technical solution of this embodiment of the present application can be used to detect whether the palmprint image is a forged image, that is, to detect whether the palmprint image sent by the terminal device 210 to the server 220 is a machine-generated image (such as a machine-generated forged image, a machine-modified palmprint image, etc.). Only when it is determined that the palmprint image sent by the terminal device 210 to the server 220 comes from a real human body (that is, a non-machine-generated image), can the similarity comparison be performed.

[0054] As another example, in the application scenario of original work protection, as Figure 3 shown, a video player program is running on the terminal device 310, and the terminal device 310 can pull a video stream from the server 320 for playback. In this process, the terminal device 310 or the server 320 can use the technical solution of this embodiment of the present application to detect whether the image frames in the pulled video content are machine-generated images. If it is detected that the image frames are machine-generated images, then a prompt message 330 can be displayed on the interface of the video player program to prompt the user that this content is suspected to be AI-generated.

[0055] It should be noted that in the embodiments of the present application, the terminal device (such as Figure 2 the terminal device 210 shown in Figure 3 andFigure 2 the server 220 shown in Figure 3 the server 320) shown in may be a server that provides various services, which may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. Among them, the terminal device and the server can communicate through a wired communication link or a wireless communication link.

[0056] The implementation details of the technical solution of the embodiments of the present application are elaborated in detail below:

[0057] Figure 4 shows a flowchart of an image processing method according to an embodiment of the present application. The image processing method can be executed by a computer device, which can be a terminal device, a server, or other devices. Refer to Figure 4 as shown, the image processing method at least includes S410 to S440, which are introduced in detail as follows:

[0058] In S410, a sample image and a modified image obtained by modifying the sample image are acquired.

[0059] In some optional embodiments, the sample image may be an image obtained for training an image recognition model. Since the image recognition model is mainly used to identify whether an image is a machine-generated image, the sample image may include natural images (such as images captured by a camera, images intercepted from a video, etc., non-machine-generated images) and machine-generated images. Optionally, the machine-generated image refers to an image generated by artificial intelligence technology, such as an image generated by a generative model such as Generative Adversarial Networks (GANs), Variational Autoencoders (VAEs), Diffusion Models, etc.

[0060] In some optional embodiments, the modified image is an image obtained by modifying the sample image. For example, the modified image can be obtained by modifying the sample image with an image modification tool. Optionally, the image modification tool can be, for example, the StableDiffusion 1.4 model, the StableDiffusion 1.5 model, the StableDiffusion XL model, etc.

[0061] In S420, according to the sample image and the modified image, a difference image between the sample image and the modified image is calculated.

[0062] In some alternative embodiments, calculating the difference image between the sample image and the modified image may be obtaining the difference image by subtracting the modified image from the sample image. For example, for each pixel position, the difference between the corresponding pixel values in the two images can be calculated, and then the difference image can be obtained. Optionally, if the sizes of the sample image and the modified image are inconsistent, the sizes of the sample image and the modified image can be adjusted to the same size first, and then the difference image is calculated; meanwhile, it is also necessary to ensure that the sample image and the modified image are in the same color space. If the sample image and the modified image are in different color spaces (such as the RGB color space, grayscale image, etc.), then the sample image and the modified image need to be unified into the same color space.

[0063] Continue to refer to Figure 4 As shown, in S430, input data for the image recognition model to be trained is generated based on the difference image, and a prediction result for the input data is output through the image recognition model to be trained, and this prediction result is used to characterize whether the sample image is a machine-generated image.

[0064] In some alternative embodiments, the process of generating input data for the image recognition model to be trained based on the difference image may be: extracting the signal-to-noise ratio data of the difference image to obtain the signal-to-noise ratio image corresponding to the difference image, and then using the signal-to-noise ratio image as the input data for the image recognition model to be trained, so as to output a prediction result for the input data through the image recognition model to be trained. Optionally, the difference image can also be directly used as the input data for the image recognition model to be trained.

[0065] Optionally, the signal-to-noise ratio data of the difference image mainly refers to a feature obtained by further processing the difference image, which is mainly used to highlight the regions with significant changes and suppress the minor changes caused by noise.

[0066] In some alternative embodiments, when extracting the signal-to-noise ratio data of the difference image, a feature image corresponding to the difference image can be extracted. The size of the feature image is the same as that of the difference image. Then, for each pixel point in the feature image, the pixel point features in the neighboring region of each pixel point are statistically analyzed to obtain the statistical feature corresponding to each pixel point. Furthermore, based on the statistical feature corresponding to each pixel point, the signal-to-noise ratio image corresponding to the difference image is generated. Optionally, the neighboring region of each pixel point can be a region with a set size centered on this pixel point, such as a 3×3 region or a 5×5 region centered on each pixel point.

[0067] Optionally, the pixel features within the neighboring region of each pixel can be statistically analyzed, which can be the pixel mean, variance, standard deviation, median, etc. of the pixels within the neighboring region of each pixel. As an example, the pixel mean and standard deviation within the neighboring region of each pixel can be statistically analyzed, and then the ratio between the pixel mean and the standard deviation can be calculated, and this ratio can be used as the statistical feature corresponding to each pixel. In other embodiments of the present application, the ratio can also be adjusted (such as adding a certain value, subtracting a certain value, multiplying by a certain ratio, etc.), and the adjusted value can be used as the statistical feature corresponding to each pixel.

[0068] In some optional embodiments, when extracting the feature image of the difference image, according to the size of the convolution kernel used for extracting the feature image, a padding layer matching the size of the convolution kernel can be added to the boundary of the difference image to obtain the processed difference image, and then the convolution kernel is used to perform convolution processing on the difference image to obtain the feature image. For example, if the size of the convolution kernel used is 3×3, then one row or one column of pixel points can be filled respectively at the boundary of the difference image (i.e., above, below, left, and right), and then convolution processing is performed; if the size of the convolution kernel used is 5×5, then 2 rows or 2 columns of pixel points can be filled respectively at the boundary of the difference image (i.e., above, below, left, and right), and then convolution processing is performed.

[0069] In some optional embodiments, the image recognition model can include a feature extraction network and a classification network, where the feature extraction network is used to extract image features; the classification network is mainly used to classify the image according to the input image features (such as signal-to-noise ratio features, etc.), and determine whether the image is a natural image (i.e., a non-machine-generated image) or a machine-generated image. Specifically, by learning the relationship between the image features and the categories, the classification network can effectively classify the image into different categories (such as natural images or machine-generated images). Based on the image recognition model in this embodiment, the signal-to-noise ratio image can be used as the input data of the feature extraction network, and the signal-to-noise ratio features output by the feature extraction network can be obtained, and then the input data of the classification network can be generated according to the signal-to-noise ratio features, so as to output the prediction result through the classification network.

[0070] Optionally, as Figure 5As shown, after calculating the difference image between the sample image and the modified image, the signal-to-noise ratio data of the difference image can be extracted to obtain the signal-to-noise ratio image corresponding to the difference image. Then, the signal-to-noise ratio image is used as the input data of the feature extraction network, and the signal-to-noise ratio feature output by the feature extraction network is used as the input data of the classification network, so as to output the prediction result through the classification network. The technical solution of this embodiment enables the classification network to learn more general rules by using the signal-to-noise ratio feature as the input data of the classification network, rather than relying on the features generated by a specific generator, which makes the model have stronger generalization ability and can thus identify images generated by different generators.

[0071] Optionally, as Figure 6 shown, after calculating the difference image between the sample image and the modified image, the signal-to-noise ratio data of the difference image can be extracted to obtain the signal-to-noise ratio image corresponding to the difference image. Then, the signal-to-noise ratio image is used as the input data of the feature extraction network, and at the same time, the semantic features of the sample image are extracted through the semantic extraction network. After that, the signal-to-noise ratio feature output by the feature extraction network and the semantic features of the sample image are fused to obtain the fused feature, and then the fused feature is used as the input data of the classification network, so as to output the prediction result through the classification network. In this embodiment, since the signal-to-noise ratio data of the difference image and the semantic features of the sample image capture the important information of the image from different perspectives, that is, the signal-to-noise ratio feature can highlight the subtle changes and noise patterns in the image, while the semantic feature provides a high-level content description. Therefore, by fusing the semantic features of the sample image with the signal-to-noise ratio image corresponding to the difference image, a more comprehensive and rich feature representation can be formed, and thus the performance and prediction accuracy of the classification network can be improved.

[0072] In some optional embodiments, when fusing the signal-to-noise ratio feature with the semantic features of the sample image, the semantic features can be normalized to obtain the processed semantic features. Then, the signal-to-noise ratio feature and the processed semantic features are fused for the first time to obtain the result of the first fusion process. After that, the result of the first fusion process and the processed semantic features are fused for the second time to obtain the fused feature. In this embodiment, fusing the signal-to-noise ratio feature with the processed semantic features for the first time enables the correlation between the signal-to-noise ratio feature and the processed semantic features to be captured. And by fusing the result of the first fusion process with the processed semantic features for the second time, the information of the signal-to-noise ratio feature and the semantic feature can be further integrated to form a more rich and comprehensive feature representation. In other words, the technical solution of the embodiment of the present application not only considers the correlation between features, but also combines the global information of the semantic features, making the fused feature more descriptive.

[0073] In some alternative embodiments, the process of normalizing the semantic features may be to perform regularization processing with a modulus length of 1 on the semantic features to normalize the semantic features and obtain the processed semantic features. In this embodiment, since different feature extraction methods may generate feature vectors with different scales, such as the signal-to-noise ratio feature and the semantic feature may have different numerical ranges, by normalizing the semantic features to unit length, it can be ensured that the signal-to-noise ratio feature and the processed semantic feature are on the same scale, avoiding a certain type of feature from dominating the final feature fusion result due to a large numerical range, and making the fused feature representation more balanced and effective.

[0074] In some alternative embodiments, the feature dimensions of the processed semantic features and the signal-to-noise ratio features can both be three-dimensional, that is, including: width dimension, height dimension, and channel dimension. Then, when performing the first fusion processing on the signal-to-noise ratio feature and the processed semantic feature, the width dimension feature and the height dimension feature in the signal-to-noise ratio feature can be combined to convert the signal-to-noise ratio feature into a two-dimensional signal-to-noise ratio feature matrix, and then the width dimension feature and the height dimension feature in the processed semantic feature can be combined to convert the processed semantic feature into a two-dimensional semantic feature matrix, and then the two-dimensional signal-to-noise ratio feature matrix and the two-dimensional semantic feature matrix are multiplied for the first time, and the result of the first multiplication is used as the result of the first fusion processing. The technical solution of this embodiment converts the three-dimensional signal-to-noise ratio feature into a two-dimensional signal-to-noise ratio feature matrix and converts the three-dimensional processed semantic feature into a two-dimensional semantic feature matrix, so that the computational complexity of the matrix multiplication operation can be reduced and the performance of the image recognition model can be improved.

[0075] In some alternative embodiments, the process of performing the second fusion processing on the result of the first fusion processing and the processed semantic feature may be to multiply the result of the first fusion processing and the two-dimensional semantic feature matrix for the second time to obtain the result of the second multiplication, and then convert the result of the second multiplication into a three-dimensional feature matrix to obtain the fused feature. The technical solution of this embodiment, since the result of the first fusion processing contains the signal-to-noise ratio feature and the preliminarily extracted semantic feature, and the second fusion processing further combines deeper semantic information, this multi-level feature fusion can provide a more rich and comprehensive feature representation, enhance the model's understanding of the image content, and thus improve the accuracy of the classification network.

[0076] Continue to refer to Figure 4 As shown, in S440, according to the loss data between the prediction result and the label of the sample image, the model parameters of the image recognition model to be trained are adjusted to obtain the trained image recognition model.

[0077] In some alternative embodiments, when training a to-be-trained image recognition model, the training data can be divided into multiple batches, with each batch containing at least one sample image. Then, according to the loss data corresponding to each sample image, the loss data for each batch is calculated, and based on the loss data for each batch, the model parameters of the to-be-trained image recognition model are adjusted. If the model parameters of the to-be-trained image recognition model are adjusted according to the loss data of multiple batches obtained by dividing the training data, it indicates that one round of training for the to-be-trained image recognition model is completed. If the to-be-trained image recognition model completes a set number of rounds of training, the trained image recognition model is determined. Optionally, the loss data between the prediction result and the label of the sample image can be cross-entropy loss data. When calculating the loss data for each batch, the loss data corresponding to the sample images included in a batch (i.e., the cross-entropy loss data for each sample image) can be averaged to obtain the loss data for each batch. Then, the stochastic gradient descent (SGD) method can be used to backpropagate the loss data for each batch into the to-be-trained image recognition model to obtain the gradients of the model parameters and update the model parameters.

[0078] In some alternative embodiments, after obtaining the trained image recognition model, the to-be-recognized image can be input into the trained image recognition model, and the recognition result for the to-be-recognized image output by the trained image recognition model is obtained. This recognition result is used to indicate whether the to-be-recognized image is a machine-generated image. Optionally, the image recognition model can provide a visualization interface to the user through a computer device. The user can submit the to-be-recognized image in this visualization interface. Then, the computer device can call the trained image recognition model to identify whether the to-be-recognized image is a machine-generated image and return the recognition result to the visualization interface provided to the user.

[0079] The following Figures 7 to 11 , elaborates on the implementation details of the technical solution of the embodiments of the present application in detail again.

[0080] The technical solution of the embodiments of the present application is mainly used to identify whether an image is a natural image (such as a non-machine-generated image obtained by shooting) or a machine-generated image (for ease of description, denoted as an AI image). Specifically, as Figure 7As shown, due to the difference in the information distribution between natural images and AI images, the differential images generated after the same forgery process for the two types of images also have this difference. That is to say, there is a certain distance between the forged images of natural images and the forged images of AI images, which is reflected in belonging to different clusters during clustering processing. Therefore, the information of the differential images can be learned through a network to identify whether the image is an AI image.

[0081] In the embodiment of the present application, it is necessary to perform training processing on the image recognition model. Specifically, as Figure 8 shown, first, a forged image is generated for the sample image (i.e., the original image shown in Figure 8 ) using a generator. The differential image is obtained by performing differential calculation on the original image and the forged image. Then, the signal-to-noise ratio data of the differential image is extracted, and this signal-to-noise ratio data is used as the information difference and input into the information network (Info Net, which is a feature extraction network). Furthermore, the signal-to-noise ratio feature output by the information network is fused with the semantic feature of the original image (the feature output from the content network Content Net), and the finally fused feature is input into the classifier for classification. Optionally, Figure 8 the information network and the classification network in

[0082] In a variant embodiment of the present application, as Figure 9 shown, the semantic feature of the original image can be removed, and directly the signal-to-noise ratio feature output by the information network is used as the input of the classification network. Optionally, Figure 9 the information network and the classification network in

[0083] In some alternative embodiments, during the data preparation stage of training the image recognition model, natural images and AI images can be collected as sample images, with labels 0 and 1 (the label 0 indicates a natural image, and the label 1 indicates an AI image; or the label 1 indicates a natural image, and the label 0 indicates an AI image). Referring to Figure 8 and Figure 9 shown, the sample images can be used as the original images to generate forged images. When generating the forged images, the stablediffusion1.4 model can be used to perform inpainting on the original images to generate forged images, or the stablediffusion1.5 or stablediffusion XL model can also be used to perform inpainting on the original images.

[0084] Continuing to refer to Figure 8 and Figure 9As shown, the input of the information network is the signal-to-noise ratio feature of the differential image (which can be represented by the signal-to-noise ratio image). When calculating the signal-to-noise ratio feature of the differential image, the following calculation process can be adopted: Subtract the forged image from the original image to obtain the differential image, and then perform a 3×3 convolution operation on the differential image. To ensure that the output feature image has the same size as the differential image, before performing the 3×3 convolution operation on the differential image, one row or one column of pixel points can be added to the edges of the differential image (i.e., the upper, lower, left, and right sides), that is, the padding of the convolution is 1, and then the convolution operation is performed. Optionally, other convolution kernel sizes can also be used to perform the convolution operation on the differential image. For example, if a 5×5 convolution is used, then before performing the 5×5 convolution operation on the differential image, two rows or two columns of pixel points can be added to the edges of the differential image (i.e., the upper, lower, left, and right sides), that is, the padding of the convolution is 2.

[0085] Optionally, the convolution kernel size of the information network can also be a learnable parameter to be updated during the training process of the model.

[0086] In some alternative embodiments, after obtaining the feature image of the differential image, the mean and variance can be calculated for each pixel point in the 3×3 region adjacent to it (i.e., the 3×3 region centered on this pixel point. In other embodiments of the present application, other sizes of adjacent regions can also be used), and then the mean mean and standard deviation std (the standard deviation is equal to the square root of the variance) at each pixel point position of the feature image are obtained. Furthermore, for each pixel point position in the feature image, the mean at this pixel point position is divided by the variance (mean / std) to obtain the signal-to-noise ratio image.

[0087] In some alternative embodiments, Figure 8 and Figure 9 the information network shown in is a feature extraction network, and its main purpose is to extract the features in the signal-to-noise ratio image that can be used to distinguish the forged image. Optionally, the information network can adopt a network structure in which multiple multi-layer perceptrons (MLPs) are connected in sequence, such as a network structure in which three non-linear fully connected layers are connected in sequence. Optionally, if the information network adopts a network structure in which three non-linear fully connected layers are connected in sequence, the output dimensions of these three non-linear fully connected layers can be set to 256, 512, and 1024 respectively. The last non-linear fully connected layer needs to be aligned with the semantic features because the signal-to-noise ratio features need to perform related operations with the semantic features, and the semantic features are 1024-dimensional (only for example. In other words, it is necessary to ensure that the output dimension of the last non-linear fully connected layer is consistent with the dimension of the semantic features).

[0088] It should be noted that the number of non-linear fully connected layers adopted by the information network can be set according to actual needs, or a network structure with multiple connected convolutional layers can be used to replace the fully connected layer. The embodiments of the present application do not limit this.

[0089] In some alternative embodiments, Figure 8 and Figure 9 the classification network shown in is mainly used to classify an image according to the input image features (such as signal-to-noise ratio features, etc.), and determine whether the image is a natural image (i.e., a non-machine-generated image) or a machine-generated image. Optionally, the classification network can adopt 1 layer of fully connected layer, with an input dimension of 1024 and an output dimension of 2, that is, the output is divided into 2 categories. An output of 0 indicates a natural image, and an output of 1 indicates an AI image; or an output of 1 indicates a natural image, and an output of 0 indicates an AI image. Optionally, a multi-layer perceptron (the structure of each layer of perceptron is: linear layer + non-linear activation unit) can also be inserted before the classification network for non-linear transformation processing of features. For example, the following cross-layer perceptron structure can be adopted: FC1+relu1+FC2+relu2+FC. Among them, relu represents a non-linear activation unit (i.e., Rectified Linear Unit).

[0090] In some alternative embodiments, considering that different image contents have an impact on the signal-to-noise ratio (as shown in the figure, when an image captures an apple and an indoor scene in the distance, since the apple bulges towards the camera while the distance bulges away from the camera, the relevant information captured by the camera for the two is very different), therefore, after extracting the signal-to-noise ratio features, as Figure 8 shown, the signal-to-noise ratio features and the semantic features of the original image can be fused, and then the fused features obtained by the fusion processing are input into the classification network.

[0091] Optionally, as Figure 10As shown, the fusion process of the signal-to-noise ratio feature and the semantic feature may include: First, perform regularization (norm) processing with a norm of 1 on the semantic feature to obtain f2. Then, perform the following correlation operations on the signal-to-noise ratio feature f1(w1, h1, 1024) and the f2(w2, h2, 1024) after regularization processing: First, merge the first two dimensions (i.e., the width dimension w and the height dimension h) of the signal-to-noise ratio feature f1 and f2 into one dimension, that is, they become (w1×h1, 1024) and (w2×h2, 1024) vectors respectively; then perform matrix multiplication operations, that is, it is equivalent to multiplying 1024 (w1×h1, 1) matrices by 1024 (1, w2×h2) matrices one by one to obtain the correlation map m1 of the signal-to-noise ratio feature and the semantic feature, with the image dimension: (w1×h1, w2×h2); then perform a second matrix calculation on the correlation map m1 and f2 (converted to two dimensions as (w2×h2, 1024)) to obtain a feature of (w1×h1, 1024) dimension, and expand and restore it to (w1, h1, 1024) dimension for output.

[0092] It should be noted that the semantic feature can be extracted by using the image branch of the Contrastive Language–Image Pretraining (CLIP) model (i.e., Figure 8 the content network shown in), and output 1024-dimensional features. Optionally, other models can also be used to extract semantic features, such as using the Bootstrapping Language-Image Pretraining (BLIP) model to extract semantic features. Among them, CLIP is mainly pre-trained on a large-scale dataset through a contrastive loss function, so that the model can map similar images and texts to close positions in the embedding space; BLIP mainly uses self-supervised learning methods to learn effective feature representations from a large number of image-text pairs without a large amount of manually labeled data.

[0093] In some alternative embodiments, Figure 8 and Figure 9The information network and classification network shown are the modules to be trained in this application, while the content network Content net does not need to update its parameters. Optionally, a deep learning network training method can be used for training. Before training, forged images are generated for N original images (i.e., sample images). During training, M rounds (e.g., 10) of iterations are performed on the full set of image data (i.e., N images). During each round of iteration, the full set of image data is divided into batches of bs samples each, and the parameters of a new model (i.e., the information network and classification network) are updated for each batch. When all batches (N / bs) have been trained once, that is, when all sample images have been trained once in the model, it is called one round of iteration. Optionally, the training process for each batch is as follows:

[0094] Parameter initialization before the first batch of the first round: For all parameters to be learned, they are initialized using a normal distribution, with a learning rate of 0.0005. After every 5 rounds of learning, the learning rate becomes 0.1 times the original. A total of 10 rounds of training are performed (it should be noted that the values in this example are only for illustration and can be adjusted according to actual needs in practical applications); bs sample images are extracted and input into the model, and each sample image and its forged image are input and the output is obtained in the Figure 8 or Figure 9 shown manner; then the classification Cross entropy loss for all samples in this batch is calculated, and the mean value is calculated to obtain the batch loss; furthermore, the method of stochastic gradient descent is used to backpropagate the loss into the model to obtain the gradients of the model parameters and update the model parameters.

[0095] In some alternative embodiments, after the model training is completed, the trained model can be deployed on a cloud server to provide services. In this case, as Figure 11 shown, the cloud server can provide an input interface. Users can input an image of a certain business or any network image through this input interface, and then the cloud server can call the trained model to perform inference and generate a prediction result, outputting whether it is an AI image.

[0096] In the technical solution of the above embodiments of the present application, since the information of the differential image is used as the input data of the model, it is possible to better obtain the tiny differences in the image and avoid the semantic information masking problem caused by directly inputting the target image. At the same time, since the signal-to-noise ratio feature is used as the input of the classification network, and the signal-to-noise ratio feature often indicates the distribution of the data, the distribution information of the differential image can be characterized, so that more effective information can be obtained for AIGC recognition. In addition, since the technical solution of the embodiments of the present application can perform AI image detection across generators, the technical solution of the embodiments of the present application can be used for general generated image recognition (a model trained based on SD1.4, used to recognize the generation effect of samples of any other generator, and the average accuracy rate is increased by 18%), and it can also have a high recognition effect.

[0097] The following introduces the device embodiments of the present application, which can be used to execute the image processing method in the above embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the embodiments of the above image processing method of the present application.

[0098] Figure 12 The block diagram of an image processing device according to an embodiment of the present application is shown. The image processing device can be applied to a computer device, which can be a terminal device, a server, or other devices.

[0099] Referring to Figure 12 As shown, an image processing device 1200 according to an embodiment of the present application includes: an acquisition unit 1202, a calculation unit 1204, a processing unit 1206, and a training unit 1208.

[0100] Among them, the acquisition unit 1202 is configured to acquire a sample image and a modified image obtained by modifying the sample image; the calculation unit 1204 is configured to calculate a differential image between the sample image and the modified image according to the sample image and the modified image; the processing unit 1206 is configured to generate input data of a to-be-trained image recognition model based on the differential image, and output a prediction result for the input data through the to-be-trained image recognition model, where the prediction result is used to characterize whether the sample image is a machine-generated image; the training unit 1208 is configured to adjust the model parameters of the to-be-trained image recognition model according to the loss data between the prediction result and the label of the sample image to obtain a trained image recognition model.

[0101] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: extract the signal-to-noise ratio data of the differential image to obtain a signal-to-noise ratio image corresponding to the differential image, and use the signal-to-noise ratio image as input data of the image recognition model to be trained, so as to output a prediction result for the input data through the image recognition model to be trained.

[0102] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: extract a feature image corresponding to the differential image, where the size of the feature image is the same as that of the differential image; for each pixel point in the feature image, count the pixel point features in the neighboring region of each pixel point to obtain a statistical feature corresponding to each pixel point; and generate a signal-to-noise ratio image corresponding to the differential image according to the statistical feature corresponding to each pixel point.

[0103] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: add a padding layer matching the size of the convolution kernel to the boundary of the differential image according to the size of the convolution kernel used to extract the feature image, to obtain a processed differential image; and perform convolution processing on the differential image using the convolution kernel to obtain the feature image.

[0104] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: count the pixel mean and standard deviation in the neighboring region of each pixel point; calculate the ratio between the pixel mean and the standard deviation, and use the ratio as the statistical feature corresponding to each pixel point.

[0105] In some embodiments of the present application, based on the foregoing solution, the image recognition model includes a feature extraction network and a classification network; the processing unit 1206 is configured to: use the signal-to-noise ratio image as input data of the feature extraction network, and obtain the signal-to-noise ratio features output by the feature extraction network; generate input data of the classification network according to the signal-to-noise ratio features, so as to output the prediction result through the classification network.

[0106] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: perform fusion processing on the signal-to-noise ratio features and the semantic features of the sample image to obtain fusion features, and use the fusion features as input data of the classification network; or use the signal-to-noise ratio features as input data of the classification network.

[0107] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: normalize the semantic features to obtain processed semantic features; perform a first fusion process on the signal-to-noise ratio features and the processed semantic features to obtain a result of the first fusion process; and perform a second fusion process on the result of the first fusion process and the processed semantic features to obtain the fusion features.

[0108] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: perform regularization processing with a norm of 1 on the semantic features to normalize the semantic features and obtain the processed semantic features.

[0109] In some embodiments of the present application, based on the foregoing solution, the feature dimensions of the processed semantic features and the signal-to-noise ratio features both include: a width dimension, a height dimension, and a channel dimension; the processing unit 1206 is configured to: merge the width dimension feature and the height dimension feature in the signal-to-noise ratio features to convert the signal-to-noise ratio features into a two-dimensional signal-to-noise ratio feature matrix; merge the width dimension feature and the height dimension feature in the processed semantic features to convert the processed semantic features into a two-dimensional semantic feature matrix; and perform a first multiplication on the two-dimensional signal-to-noise ratio feature matrix and the two-dimensional semantic feature matrix, and use the result of the first multiplication as the result of the first fusion process.

[0110] In some embodiments of the present application, based on the foregoing solution, the processing unit 1206 is configured to: perform a second multiplication on the result of the first fusion process and the two-dimensional semantic feature matrix to obtain a result of the second multiplication; and convert the result of the second multiplication into a three-dimensional feature matrix to obtain the fusion features.

[0111] In some embodiments of the present application, based on the foregoing solution, the training unit 1208 is configured to: divide the training data into multiple batches, with each batch containing at least one sample image; calculate the loss data for each batch according to the loss data corresponding to each sample image, and adjust the model parameters of the image recognition model to be trained according to the loss data of each batch; if the model parameters of the image recognition model to be trained are adjusted according to the loss data of the multiple batches, then complete one round of training process for the image recognition model to be trained; and if the training process for the image recognition model to be trained is completed for a set number of rounds, then determine that the trained image recognition model is obtained.

[0112] Figure 13The block diagram of an image processing apparatus according to an embodiment of the present application is shown. The image processing apparatus can be applied to a computer device, which can be a terminal device, a server, or other devices.

[0113] Referring to Figure 13 As shown, an image processing apparatus 1300 according to an embodiment of the present application includes: an acquisition unit 1302 and an identification unit 1304.

[0114] Among them, the acquisition unit 1302 is configured to acquire an image to be identified; the identification unit 1304 is configured to input the image to be identified into a trained image recognition model to obtain an identification result of the image to be identified output by the trained image recognition model, and the identification result is used to characterize whether the image to be identified is a machine-generated image; wherein, the trained image recognition model is trained according to the image processing method described in the foregoing embodiment.

[0115] Figure 14 The structural schematic diagram of a computer system of a computer device suitable for implementing the embodiments of the present application is shown. The computer device can be the computer device used to execute the image processing method in the foregoing embodiment.

[0116] It should be noted that Figure 14 The computer system 1400 of the computer device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present application.

[0117] As Figure 14 shown, the computer system 1400 may include a central processing unit (CPU) 1401, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1402 or the program loaded from the storage section 1408 into the random access memory (RAM) 1403, such as executing the method described in the foregoing embodiment. In the RAM 1403, various programs and data required for system operation are also stored. The CPU 1401, ROM 1402, and RAM 1403 are connected to each other through a bus 1404. The input / output (I / O) interface 1405 is also connected to the bus 1404.

[0118] The following components can be connected to the I / O interface 1405: an input part 1406 including a keyboard, a mouse, etc.; an output part 1407 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage part 1408 including a hard disk, etc.; and a communication part 1409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication part 1409 performs communication processing via a network such as the Internet. The drive 1410 is also connected to the I / O interface 1405 as needed. A removable medium 1411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1410 as needed so that a computer program read from it can be installed into the storage part 1408 as needed.

[0119] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program is used to execute the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1409, and / or installed from the removable medium 1411. When the computer program is executed by the central processing unit (CPU) 1401, various functions defined in the system of the present application are executed.

[0120] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a computer program, and this computer program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable computer program is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and a computer program.

[0122] The units involved in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, constitute a limitation on the units themselves.

[0123] As another aspect, the present application also provides a computer-readable medium, which may be included in the computer device described in the above embodiments; or may exist separately without being assembled into the computer device. The above computer-readable medium carries one or more computer programs, and when the above one or more computer programs are executed by a computer device, the computer device implements the methods described in the above embodiments.

[0124] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-described modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0125] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented in software or in a manner of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computer device to execute the methods according to the embodiments of the present application. For example, the computer device can execute Figure 4 the image processing method shown.

[0126] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0127] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. An image processing method, characterized in that: include: Acquire a sample image and modify the sample image to obtain a modified image; Calculating a difference image between the sample image and the modified image according to the sample image and the modified image; generating input data of an image recognition model to be trained based on the differential image, and outputting a prediction result for the input data through the image recognition model to be trained, wherein the prediction result is used to characterize whether the sample image is a machine-generated image; According to the loss data between the prediction result and the label of the sample image, the model parameters of the image recognition model to be trained are adjusted to obtain the trained image recognition model.

2. The image processing method according to claim 1, characterized in that: Generating input data of an image recognition model to be trained based on the differential image, and outputting a prediction result for the input data through the image recognition model to be trained, including: The signal-to-noise ratio data of the differential image is extracted to obtain a signal-to-noise ratio image corresponding to the differential image, and the signal-to-noise ratio image is used as input data of the image recognition model to be trained, so as to output a prediction result for the input data through the image recognition model to be trained.

3. The image processing method according to claim 2, characterized in that: Extracting signal-to-noise ratio data of the differential image to obtain a signal-to-noise ratio image corresponding to the differential image includes: Extracting a feature image corresponding to the differential image, wherein the feature image has the same size as the differential image; For each pixel in the feature image, counting pixel features in a neighborhood of each pixel to obtain statistical features corresponding to each pixel; A signal-to-noise ratio image corresponding to the differential image is generated according to the statistical features corresponding to each pixel point.

4. The image processing method according to claim 3, characterized in that: Extracting a characteristic image of the differential image includes: According to the size of the convolution kernel used to extract the feature image, a padding layer matching the size of the convolution kernel is added to the boundary of the difference image to obtain a processed difference image; The difference image is convolved using the convolution kernel to obtain the feature image.

5. The image processing method according to claim 3, characterized in that: Counting pixel features in a neighborhood of each pixel to obtain statistical features corresponding to each pixel, including: Counting the pixel mean and standard deviation in the neighborhood of each pixel; The ratio between the pixel mean and the standard deviation is calculated, and the ratio is used as the statistical feature corresponding to each pixel point.

6. The image processing method according to claim 2, characterized in that: The image recognition model includes a feature extraction network and a classification network; Using the signal-to-noise ratio image as input data of the image recognition model to be trained, so as to output a prediction result for the input data through the image recognition model to be trained, comprising: Using the signal-to-noise ratio image as input data of the feature extraction network, and obtaining the signal-to-noise ratio feature output by the feature extraction network; The input data of the classification network is generated according to the signal-to-noise ratio feature, so as to output the prediction result through the classification network.

7. The image processing method according to claim 6, characterized in that: Generating input data of the classification network according to the signal-to-noise ratio feature includes: fusing the signal-to-noise ratio feature with the semantic feature of the sample image to obtain a fused feature, and using the fused feature as input data of the classification network; or The signal-to-noise ratio feature is used as input data of the classification network.

8. The image processing method according to claim 7, characterized in that: The signal-to-noise ratio feature is fused with the semantic feature of the sample image to obtain a fused feature, including: Normalizing the semantic features to obtain processed semantic features; Performing a first fusion process on the signal-to-noise ratio feature and the processed semantic feature to obtain a result of the first fusion process; The result of the first fusion processing is subjected to a second fusion processing with the processed semantic features to obtain the fusion features.

9. The image processing method according to claim 8, characterized in that: The semantic features are normalized to obtain processed semantic features, including: Regularization processing with a modulus length of 1 is performed on the semantic feature to normalize the semantic feature to obtain the processed semantic feature.

10. The image processing method according to claim 8, characterized in that: The feature dimensions of the processed semantic features and the feature dimensions of the signal-to-noise ratio features both include: a width dimension, a height dimension, and a channel dimension; The signal-to-noise ratio feature and the processed semantic feature are subjected to a first fusion process to obtain a result of the first fusion process, including: Merging the width dimension feature and the height dimension feature in the signal-to-noise ratio feature to convert the signal-to-noise ratio feature into a two-dimensional signal-to-noise ratio feature matrix; Merging the width dimension feature and the height dimension feature in the processed semantic feature to convert the processed semantic feature into a two-dimensional semantic feature matrix; The two-dimensional signal-to-noise ratio feature matrix and the two-dimensional semantic feature matrix are multiplied for the first time, and a result of the first multiplication is used as a result of the first fusion processing.

11. The image processing method according to claim 8, characterized in that: The result of the first fusion process is subjected to a second fusion process with the processed semantic feature to obtain the fusion feature, including: Multiplying the result of the first fusion processing by the two-dimensional semantic feature matrix for a second time to obtain a result of the second multiplication; The result of the second multiplication is converted into a three-dimensional feature matrix to obtain the fused feature.

12. The image processing method according to any one of claims 1 to 11, characterized in that: According to the loss data between the prediction result and the label of the sample image, the model parameters of the image recognition model to be trained are adjusted to obtain the trained image recognition model, including: Divide the training data into multiple batches, each batch contains at least one sample image; Calculating the loss data of each batch according to the loss data corresponding to each sample image, and adjusting the model parameters of the image recognition model to be trained according to the loss data of each batch; If the model parameters of the image recognition model to be trained are adjusted according to the loss data of the multiple batches, then a round of training process of the image recognition model to be trained is completed; If the set rounds of training process are completed for the image recognition model to be trained, it is determined that the trained image recognition model is obtained.

13. An image processing method, characterized in that: include: Obtain an image to be recognized; Inputting the image to be identified into a trained image recognition model to obtain a recognition result output by the trained image recognition model for the image to be identified, wherein the recognition result is used to indicate whether the image to be identified is a machine-generated image; Wherein, the trained image recognition model is obtained according to the image processing method described in any one of claims 1 to 12.

14. An image processing device, characterized in that: include: an acquisition unit, configured to acquire a sample image and a modified image obtained by modifying the sample image; a calculation unit configured to calculate a difference image between the sample image and the modified image based on the sample image and the modified image; a processing unit configured to generate input data of an image recognition model to be trained based on the differential image, and output a prediction result for the input data through the image recognition model to be trained, wherein the prediction result is used to characterize whether the sample image is a machine-generated image; The training unit is configured to adjust the model parameters of the image recognition model to be trained according to the loss data between the prediction result and the label of the sample image, so as to obtain the trained image recognition model.

15. An image processing device, characterized in that: include: An acquisition unit, configured to acquire an image to be recognized; a recognition unit configured to input the image to be recognized into a trained image recognition model, and obtain a recognition result for the image to be recognized output by the trained image recognition model, wherein the recognition result is used to indicate whether the image to be recognized is a machine-generated image; Wherein, the trained image recognition model is obtained according to the image processing method described in any one of claims 1 to 12.

16. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 13 is implemented.

17. A computer device, characterized in that: include: one or more processors; A memory for storing one or more computer programs, which, when executed by the one or more processors, enables the computer device to implement the image processing method according to any one of claims 1 to 13.

18. A computer program product, characterized in that The computer program product comprises a computer program, which is stored in a computer-readable storage medium. A processor of a computer device reads and executes the computer program from the computer-readable storage medium, so that the computer device executes the image processing method according to any one of claims 1 to 13.