Image authenticity detection model training method and device, equipment and storage medium

By fine-tuning the low-rank decomposition matrix and image processing of the pre-trained model, and building an association between the low-rank adaptation model and the classification head, the problems of image authenticity detection model training complexity and accuracy are solved, and efficient and accurate image authenticity detection is achieved.

CN120599401APending Publication Date: 2025-09-05BEIJING REALAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510673265.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing image authenticity detection models are complex to train and lack detection accuracy, especially when dealing with complex and varied forgery methods. In addition, multi-model training requires a lot of time and resources.

Method used

By obtaining a pre-trained visual feature extraction model, using a low-rank decomposition matrix to fine-tune the model, building a low-rank adaptation model and associating it with a classification head, and combining it with a training image set for training, a true and false image detection model with multiple low-rank adaptation models and classification heads is formed, and decision fusion is used to improve detection accuracy.

Benefits of technology

It reduces the complexity and resource consumption of model training, improves the efficiency and flexibility of model adjustment, and enhances the accuracy of image true and false detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599401A_ABST
    Figure CN120599401A_ABST
Patent Text Reader

Abstract

The invention discloses an image authenticity detection model training method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a pre-trained visual feature extraction model, and carrying out the model fine tuning of the visual feature extraction model based on a low-rank decomposition matrix, and obtaining a plurality of low-rank adaptation models; constructing a classification head for performing dichotomy processing on the authenticity of the image, and determining an association relationship between the low-rank adaptation model and the classification head; acquiring an image set for training, and performing image processing on each image in the image set to obtain an image group corresponding to each image in the image set; and training the low-rank adaptation model and the classification head according to the image group, and obtaining an image authenticity detection model formed based on the trained low-rank adaptation model and the trained classification head after the training is completed. Through loading and fine tuning of the pre-training model, the complexity of model training is reduced, and through feature extraction based on multiple different dimensions, the accuracy of image detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image detection technology, and in particular to a training method, device, electronic device and storage medium for an image authenticity detection model. Background Art

[0002] With the rapid advancement of generative AI technology, editing images and multimedia content has ushered in an era of unprecedented convenience and efficiency. Advanced models such as Midjourney, SDXL, and DALL E 2 have made it easy for even non-professional users to create high-quality image content. However, this double-edged sword of technology has also brought new challenges regarding the authenticity and misleading nature of information. Fake images created using AI technology are becoming increasingly difficult to detect with the naked eye, and this has begun to cause public concern.

[0003] Currently, mainstream fake image detection methods rely primarily on deep neural networks to extract visual features, such as convolutional neural networks (CNNs) and vision transformers (ViTs). These methods exploit forged features in images to distinguish true from false. Examples include building fake image detectors based on a single deep neural network model and combining different models (such as CNNs and ViTs) to form a more powerful ensemble detector. However, both approaches present certain challenges. The former is limited in handling complex and varied forgeries, resulting in lower detection accuracy. The latter, however, requires training and maintaining multiple models, which consumes significant time and resources. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a training method, device, electronic device and storage medium for an image authenticity detection model to solve the technical problems in the related art that the training of the image authenticity detection model is highly complex and the image detection is not accurate enough.

[0005] In a first aspect, an embodiment of the present application provides a method for training an image authenticity detection model, comprising:

[0006] Obtaining a pre-trained visual feature extraction model, and fine-tuning the visual feature extraction model based on a low-rank decomposition matrix to obtain several low-rank adaptation models;

[0007] Constructing a classification head for performing binary classification processing on the true and false images, and determining the correlation relationship between the low-rank adaptation model and the classification head;

[0008] Acquire a training image set, and perform image processing on each image in the image set to obtain an image group corresponding to each image in the image set;

[0009] The low-rank adaptive model and the classification head are trained according to the image group, and when the training is completed, an image true and false detection model based on the trained low-rank adaptive model and the trained classification head is obtained.

[0010] In a second aspect, an embodiment of the present application provides a training device for an image authenticity detection model, comprising:

[0011] A fine-tuning processing module is used to obtain a pre-trained visual feature extraction model and fine-tune the visual feature extraction model based on a low-rank decomposition matrix to obtain a plurality of low-rank adaptation models;

[0012] An association processing module, configured to construct a classification head for performing binary classification processing on the authenticity of an image, and determine an association relationship between the low-rank adaptation model and the classification head;

[0013] An image acquisition module is used to acquire an image set for training and perform image processing on each image in the image set to obtain an image group corresponding to each image in the image set;

[0014] A training processing module is used to train the low-rank adaptive model and the classification head according to the image group, and obtain an image true and false detection model based on the trained low-rank adaptive model and the trained classification head when the training is completed.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the steps in the training method of the image authenticity detection model described in any one of the above items.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in the training method of the image authenticity detection model described in any one of the above items.

[0017] The embodiment of the present application provides a training method, device, electronic device and storage medium for an image authenticity detection model. During training, a pre-trained visual feature extraction model is obtained, and the visual feature extraction model is fine-tuned based on a low-rank decomposition matrix to obtain several low-rank adaptive models. At the same time, each low-rank adaptive model corresponds to a classification head for binary classification processing. Then, a set of images to be trained is obtained, and image processing is performed on each image in the image set to obtain an image group corresponding to each image in the image set. The low-rank adaptive model and the classification head are trained according to the image group, and when the training is completed, an image authenticity detection model composed of the trained low-rank adaptive model and the trained classification head is obtained. By adjusting the pre-trained model, the complexity and resource consumption of the model training are reduced. At the same time, fine-tuning the model based on the low-rank decomposition matrix improves the efficiency of the model adjustment and the flexibility of subsequent deployment. In addition, when making predictions, the output results of each classification head are fused for decision-making, so as to detect the authenticity of the image based on multiple different dimensions, thereby improving the accuracy of subsequent authenticity detection of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a flowchart of the steps of the training method of the image authenticity detection model provided in the embodiment of the present application;

[0019] Figure 2 This is a flow chart of the steps for performing external parameter calibration provided in an embodiment of the present application;

[0020] Figure 3 This is a flow chart of the steps of filtering the second point cloud data provided in an embodiment of the present application;

[0021] Figure 4 This is another flowchart of the steps for training a point cloud image authenticity detection model provided in an embodiment of the present application;

[0022] Figure 5 This is a structural diagram of a training device for an image authenticity detection model provided in an embodiment of the present application;

[0023] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0024] Figure 7 This is another structural diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] It should be understood that the various steps described in the method embodiments disclosed herein may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0027] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.

[0028] In related technologies, mainstream fake image detection methods rely primarily on deep neural networks to extract visual features, such as convolutional neural networks (CNNs) and vision transformers (ViTs), to mine forged features in images to distinguish true from false. Examples include building fake image detectors based on a single deep neural network model and combining different models (such as CNNs and ViTs) to form a more powerful integrated detector. However, both approaches present certain challenges. The former is limited in handling complex and diverse forgeries, while the latter, due to the need to train and maintain multiple models, requires significant time and resources.

[0029] In order to solve the technical problems existing in the related art, the embodiment of the present application provides a training method for an image true and false detection model, see Figure 1 , Figure 1 This is a flow chart of the steps of the training method of the image authenticity detection model provided in an embodiment of the present application, and the method includes steps 101 to 104.

[0030] Step 101: obtain a pre-trained visual feature extraction model, and fine-tune the visual feature extraction model based on a low-rank decomposition matrix to obtain several low-rank adaptation models.

[0031] Among them, when the model is trained based on the method of the present application to obtain an image authenticity detection model, LoRA (Low-Rank Adaptation) is used to fine-tune the extremely small-scale low-rank matrix parameters in the pre-trained model, thereby achieving effective fine-tuning of the model, retaining the rich knowledge representation of the pre-trained model, and reducing computational complexity and resource consumption. Then, by designing diverse training strategies and configurations, multiple LoRA model variants with unique learning characteristics are trained to capture the forgery features of different dimensions in the image data, so that the trained image authenticity detection model can analyze and process the authenticity of the image based on different dimensions.

[0032] In one embodiment, when constructing an image authenticity detection model for detecting the authenticity of an image, a pre-trained visual feature extraction model is first obtained, and the visual feature extraction model is fine-tuned, and then the fine-tuned visual feature extraction model is used as a partial model for extracting features from the image in the image authenticity detection model.

[0033] Specifically, a pre-trained visual feature extraction model is obtained, and the visual feature extraction model is fine-tuned based on the LoRA method. A low-rank decomposition matrix is ​​constructed to add the low-rank decomposition matrix to the loaded visual feature extraction model to achieve model fine-tuning of the visual feature extraction model. At the same time, multiple different low-rank decomposition matrices are used to fine-tune the parameters of the pre-trained visual feature extraction model to obtain several low-rank adaptation models.

[0034] For example, a multi-LoRA model set is created on the pre-trained visual feature extraction model E For each network layer of Model E, LoRA i Multiple parameter layers will be randomly selected to add low-rank decomposition matrices. The rank of the matrix is ​​randomly selected as an integer r in the interval [8,16]. After randomly adding the low-rank decomposition matrix, a low-rank adaptation model will be obtained, and the data required for the low-rank adaptation model is based on the construction of the image true and false detection model.

[0035] Step 102: construct a classification head for performing binary classification processing on the true or false images, and determine the correlation relationship between the low-rank adaptation model and the classification head.

[0036] In one embodiment, after completing the fine-tuning of the pre-trained visual feature extraction model based on the set low-rank decomposition matrix, a classification head for binary classification is also constructed. In fact, the low-rank adaptation model obtained by fine-tuning can extract features from images in different dimensions, and when performing true or false detection of images, the extracted features are processed to determine the true or false detection result corresponding to the image. Therefore, in order to realize feature-based true or false detection of images, a classification head for binary classification is constructed to perform binary classification based on the obtained image features to distinguish between real images and false images.

[0037] For example, after fine-tuning the pre-trained visual feature extraction model, a low-rank adaptive model for feature extraction in different dimensions can be obtained, so that when detecting the authenticity of an image, analysis and judgment can be performed based on features of multiple different dimensions, which can effectively improve the accuracy of image authenticity detection.

[0038] To perform analysis based on different dimensions, a corresponding classification head must be built for each low-rank adaptive model, and the corresponding relationship between each low-rank adaptive model and the classification head is unique. After the classification head is built and associated with the low-rank adaptive model, the image features in different dimensions can be analyzed and processed, improving the accuracy of true and false image detection.

[0039] Further, refer to Figure 2 , Figure 2 This is a schematic diagram of a model structure of an image true / false detection model provided in an embodiment of the present application. The image true / false detection model is composed of several low-rank adaptive models and classification heads. In the image true / false detection model, each low-rank adaptive model corresponds to one classification head, and the image true / false detection model is obtained by combining multiple low-rank adaptive models and classification heads. Figure 2 When the image authenticity detection model shown performs authenticity detection and analysis on an image, the image to be detected is input into the trained low-rank adaptive model to extract features of the image to be detected in different dimensions, and each feature extracted by the low-rank adaptive model is input into the corresponding trained classification head for binary classification processing, that is, the binary classification result corresponding to each classification head will be obtained, that is, the detection result, and then the decision integration module is performed to obtain the authenticity detection of the image to be detected, wherein the decision integration module fuses the output results of each classification head to obtain the final prediction result.

[0040] Step 103: Acquire a training image set, and perform image processing on each image in the image set to obtain an image group corresponding to each image in the image set.

[0041] In one embodiment, before model training, an image set containing both real and fake images is obtained. The image set used for model training is then labeled as real or fake, resulting in the image set used for training. After obtaining the image set, each image in the image set is processed accordingly to obtain an image group corresponding to each image in the image set.

[0042] Among them, the image group contains multiple images and includes an image in the image set, that is, the obtained image group contains an image in the image set and several images after the image is processed, and then the low-rank adaptation model and classification head obtained in step 101 and step 102 are trained and processed based on the obtained several image groups.

[0043] Further, refer to Figure 3 , Figure 3 This is a flow chart of the steps of obtaining an image group provided in an embodiment of the present application, wherein the steps include steps 301 to 304.

[0044] Step 301: obtaining pre-labeled real images and fake images to form an image set for training;

[0045] Step 302, determining the number of models of the low-rank adaptive model;

[0046] Step 303: perform image augmentation processing on each image in the image set according to the number of models to obtain an augmented image corresponding to each image;

[0047] Step 304 : Combine the augmented image with the corresponding original image to obtain an image group for each image, wherein the number of images in the image group is the same as the number of models, and the original image is the image in the image set that has been augmented.

[0048] Specifically, when constructing an image group, image enhancement and other processing are performed based on an image in the training image to obtain a certain number of enhanced images, and then the image that has been enhanced is combined with the enhanced image obtained after the processing to obtain an image group corresponding to each image in the training image.

[0049] Exemplarily, when each image in the training image is processed to obtain an image group, the number of images contained in the image group is related to the model structure of the image authenticity detection model, specifically, the number of low-rank adaptive models and classification heads contained in the model. Therefore, when the training image is obtained and processed to obtain the image group, the number of models of the low-rank adaptive model is determined, and then an image is selected in the training image as the original image, and the original image is enhanced multiple times based on the obtained number of models to obtain a certain number of enhanced images. When performing image enhancement, the processing means include rotation, cropping, etc., which are not specifically limited. For example, when the number of low-rank adaptive models is 4, the number of enhanced images after enhancement processing is 3, and then they are combined with the original image to obtain an image group containing 4 images.

[0050] In practical applications, after training the low-rank adaptation model and the classification head based on the training images, it is necessary to extract and analyze the features of the image in multiple different dimensions. Therefore, the images input into each low-rank adaptation model need to be the same or similar images. Therefore, in addition to obtaining the enhanced image corresponding to the original image through the enhancement method described above, the image can also be copied to obtain an image group.

[0051] Step 104 , training the low-rank adaptive model and the classification head according to the image group, and obtaining an image true / false detection model based on the trained low-rank adaptive model and the trained classification head when the training is completed.

[0052] In one embodiment, after obtaining the corresponding image group by processing the training images, the image group low-rank adaptation model and classification head will be used for training, and then when the training is completed, the image authenticity detection model used for subsequent image authenticity detection is constructed based on the trained low-rank adaptation model and the trained classification head.

[0053] Exemplarily, when training the low-rank adaptation model and the classification head, there is an association between the low-rank adaptation model and the classification head, wherein the low-rank adaptation model is used to extract features of the image, and the classification head is used to perform binary classification based on the extracted features. Finally, the results obtained by each classification head are fused to obtain the true and false detection results of the original image, that is, the prediction results, and then the true and false detection results obtained by the model prediction processing are compared with the actual labels (true and false labels) of the original image. The corresponding method is used to determine whether the model training is completed, such as by using the set loss function to determine, where the loss function can be as follows:

[0054]

[0055] Among them, y i is the sample xi The true label, p i is the true and false detection result after fusion (i.e., predicted probability).

[0056] When the value of the loss function meets the set convergence condition, the training is determined to be completed. Otherwise, the training continues until the loss function value meets the convergence condition.

[0057] After the training is completed, when constructing an image true or false detection model for detecting true or false images, the structure of the constructed model can be as follows: Figure 2 As shown, it specifically includes several trained low-rank adaptation models and several classification heads, as well as a decision fusion module, which outputs the true or false prediction results of the image, such as prediction probability.

[0058] Furthermore, when training the low-rank adaptive model and the classification head based on the image group, the low-rank adaptive model and the classification head need to be initialized first, specifically including: initializing the low-rank adaptive model and the classification head to obtain the initialized low-rank adaptive model and the initialized classification head; constructing the initialized low-rank adaptive model and the initialized classification head based on the association relationship to obtain several training models, wherein each training model contains a low-rank adaptive model and a classification head; training the training model according to the image group, and obtaining an image true and false detection model obtained by model construction based on the trained training model when the training is completed.

[0059] Specifically, when performing initialization processing, the initialization method can be selected from zero initialization, normal distribution initialization, uniform distribution initialization, Kaiming initialization and Xavier initialization, and then the parameters in the low-rank adaptation model and the classification head are initialized based on the selected method.

[0060] After completing the parameter initialization processing, since the low-rank adaptation model and the classification head are one-to-one corresponding, the true and false detection of the image can be completed based on the features of different dimensions. Therefore, after completing the initialization processing, based on the determined association relationship between the low-rank adaptation model and the classification head, several training models are constructed, and then the training models are trained using the combined image group. Among them, when training the training model, all training models composed of low-rank adaptation models and classification heads are trained at the same time to determine whether the training is completed based on the result output of the decision integration module and the convergence of the model.

[0061] When the training is completed, each low-rank adaptation model in the image true or false detection model and each parameter in each classification head are optimized. In addition, when performing decision integration processing, the output results of each classification head are fused in a weighted manner. At this time, the weights of the classification heads can be made equal to the default weight values, and then the prediction or detection results corresponding to the image are obtained through weighted calculation.

[0062] Therefore, when training the training model according to the image group, it includes: inputting each image in the image group into the training model respectively, and obtaining the prediction probability of each model in the training model for the image group; obtaining the true or false prediction results of the image group according to the prediction probability; judging the training progress of the training model according to the true or false prediction results and the true or false results of the image group, and when it is determined that the training progress is completed, obtaining the image true or false detection model constructed based on the trained training model, wherein the training progress includes training completed and training incomplete.

[0063] That is, during training, for the images contained in the image group, an output result (the output result of the classification head) will be obtained after being processed by any training model. At the same time, for each image group, the labels of the original images contained therein are known. During the training process, the loss function value is calculated based on the predicted labels and the known labels to determine whether the training is completed, and the training progress is whether the training is completed. When the training progress is completed, that is, the value of the loss function recorded above meets the convergence condition, it will be determined that the training is completed, that is, the adjustment of the parameters of the low-rank adaptation model and the classification head in the training model is completed, otherwise the training and parameter adjustment will continue.

[0064] In addition, during the training process of the training model, when obtaining the true or false prediction results of the image group based on the prediction probability, it includes: obtaining the default weight of the classification head in the training model, and performing weighted calculation based on the default weight and the prediction probability to obtain the true or false prediction probability of the image group.

[0065] Among them, each training model will output a result, and then the output results of each training model are weighted by weighted calculation to obtain the prediction results of each image in the training image, specifically a prediction value. In practical applications, the absolute calculation result corresponding to the real image can be set to 0, and the absolute calculation result of the false image can be set to 1. That is, when making a prediction, when the obtained value is close to 0, it means that the image is a real image. Conversely, when the obtained value is close to 1, it means that the image is a false image.

[0066] The result obtained after weighted fusion is a specific number between 0 and 1, and 0.5 is usually used as the dividing line between true and false. When the result is less than 0.5, it is determined to be a real image, otherwise it is determined to be a false image.

[0067] Through the above-mentioned training method, the pre-trained visual feature extraction model is first loaded, and the model is fine-tuned based on the low-rank decomposition matrix to obtain multiple low-rank adaptation models for extracting different dimensions, and a corresponding classification head is constructed for each low-rank adaptation model. Then, the low-rank adaptation model and the classification head of the training image are used for training. By adjusting the pre-trained model, the complexity and resource consumption of the model training are reduced. At the same time, the model is fine-tuned based on the low-rank decomposition matrix, which improves the efficiency of model adjustment and the flexibility of subsequent deployment. After completing the training of the low-rank adaptation model and the classification head, the model is constructed based on the trained multiple low-rank adaptation models and classification heads, and decision fusion is performed based on the output results of multiple classification heads to realize the detection of the authenticity of the image based on multiple different dimensions, thereby improving the accuracy of the subsequent detection of the authenticity of the image.

[0068] Furthermore, after completing the training of the low-rank adaptation model and the classification head to obtain the image true and false detection model, the weighted weights of each classification head can be adjusted during decision fusion before use to meet the intuitiveness and accuracy of subsequent detection. Figure 4 , Figure 4 This is another flow chart of the steps of the training method of the image authenticity detection model provided in an embodiment of the present application, wherein the steps include steps 401 to 404.

[0069] Step 401: construct a weight combination of classification heads in the image authenticity detection model, wherein each weight in the weight combination corresponds to a classification head;

[0070] Step 402: Adjust the image authenticity detection model based on the weight combination to obtain an adjusted image authenticity detection model.

[0071] Step 403: Acquire a test image set, and perform a model test on the adjusted image authenticity detection model based on the test image set to determine the prediction fit of the weight combination;

[0072] In step 404, a weight combination for adjusting the classification head in the training model is obtained according to the predicted fit, and the image authenticity detection model is fine-tuned based on the weight combination to obtain an adjusted image authenticity detection model.

[0073] For example, after obtaining the image authenticity detection model based on the process of steps 101 to 104, the weight values ​​for weighted calculation can be adjusted before use to enable more intuitive and accurate judgment. For example, when detecting a real image based on the image authenticity detection model, and taking 0.5 as the dividing line between authenticity and falsehood, if the detection result for the real image is 0.4 at this time, it can be determined that the detection result is correct. However, in the actual processing process, abnormal situations are inevitable. For example, the detection result of a real image is 0.49. Since it is close to the edge of the dividing line, the recognition of its authenticity is low at this time. Since the weighted weights when weighting the output results of each classification head are default values, they can be adjusted at this time so that the output results of the real image processed based on the image authenticity detection model are closer to 0, and the output results of the false image are closer to 1.

[0074] Furthermore, the image true and false detection model obtained by training can be tested before use. Specifically, several groups of weight combinations for the classification head can be set, and then the image true and false detection model can be fine-tuned based on the set weight combinations. At this time, the specific adjustment objects can be Figure 2 The decision integration module shown in the figure adjusts the weight value of each classification head in the decision integration module, and then uses the set test image set to test the adjusted image true and false detection model.

[0075] For each weight combination, a prediction fit degree will be obtained during the test, which is specifically the distance between 0 and 1 when the prediction is correct. For example, when the prediction is accurate for a real image, if the output result of the decision integration module is closer to 0, the prediction fit degree is high, otherwise it means the prediction fit degree is low. Similarly, when the prediction is accurate for a fake image, if the output result of the decision integration module is closer to 1, the prediction fit degree is high, otherwise it means the prediction fit degree is low.

[0076] When determining the weight combination based on the predicted fit, the set of weight combinations with the highest predicted fit is used as the weight combination for fine-tuning, that is, the decision integration module in the image authenticity detection model is fine-tuned based on the weight combination, and then after fine-tuning, an image authenticity detection model that can be used for subsequent use is obtained.

[0077] It should be noted that the weight combination can be set manually or continuously adjusted based on the testing process, with no specific restrictions. When continuously adjusted based on the testing process, the appropriate adjustment can be determined based on the distance between the predicted value and the absolute value (0 and 1). For example, when the distance to the mean is less than 0.2, the adjustment is terminated, i.e., the model is fine-tuned based on the currently determined weight combination.

[0078] Furthermore, after completing the training and fine-tuning of the image authenticity detection model, it can be used in subsequent image authenticity detection. When performing authenticity detection, it includes: obtaining the image to be detected, and inputting the image to be detected into the image authenticity detection model to obtain a predicted probability value of the image to be detected; according to the predicted probability value, obtaining the authenticity detection result of the image to be detected; if the predicted probability value is in the true interval, then determining that the authenticity detection result of the image to be detected is a true image; if the predicted probability value is in the false interval, then determining that the authenticity detection result of the image to be detected is a false image; if the predicted probability value is outside the true interval and the false interval, then determining that the authenticity detection result of the image to be detected is a blurred image.

[0079] For example, when performing authenticity detection on an image, the image to be tested is input into the image authenticity detection model to obtain a predicted probability value corresponding to the image to be tested, specifically a value between 0 and 1. The predicted probability value is then compared with different numerical intervals set for image authenticity determination, such as the true interval and the false interval. When the obtained predicted probability value is within the true interval, the image is determined to be true, and when the predicted probability value is within the false interval, the image is determined to be false. The specific setting values ​​for the true interval and the false interval are set based on actual conditions.

[0080] In addition, in the actual true-false detection process, there may be cases where the detection is relatively vague. For example, for a false image that is extremely similar to a real image, the output result of the model may be relatively centered during the detection, such as 0.495. Since 0.5 is the dividing line between true and false, images with detection results within this range can be defined as blurred images to prompt users or operators to pay attention to them.

[0081] In summary, the above embodiment provides a training method for an image authenticity detection model. During training, a pre-trained visual feature extraction model is obtained, and the visual feature extraction model is fine-tuned based on a low-rank decomposition matrix to obtain several low-rank adaptation models. At the same time, each low-rank adaptation model corresponds to a classification head for binary classification processing. Then, a set of images to be trained is obtained, and image processing is performed on each image in the image set to obtain an image group corresponding to each image in the image set. The low-rank adaptation model and the classification head are trained according to the image group, and when the training is completed, an image authenticity detection model composed of the trained low-rank adaptation model and the trained classification head is obtained. By adjusting the pre-trained model, the complexity and resource consumption of the model training are reduced. At the same time, fine-tuning the model based on the low-rank decomposition matrix improves the efficiency of the model adjustment and the flexibility of subsequent deployment. In addition, when making predictions, the output results of each classification head are fused for decision-making, so as to detect the authenticity of the image based on multiple different dimensions, thereby improving the accuracy of subsequent authenticity detection of the image.

[0082] According to the method described in the above embodiment, this embodiment will be further described from the perspective of a training device for an image authenticity detection model. The training device for the image authenticity detection model can be implemented as an independent entity or integrated into an electronic device, such as a terminal. The terminal may include a mobile phone, a tablet computer, etc.

[0083] See Figure 5 , Figure 5 This is a structural diagram of a training device for an image authenticity detection model provided in an embodiment of the present application. Figure 5 As shown, the training device 500 of the image authenticity detection model provided in the embodiment of the present application includes:

[0084] A fine-tuning processing module 501 is used to obtain a pre-trained visual feature extraction model and fine-tune the visual feature extraction model based on a low-rank decomposition matrix to obtain several low-rank adaptation models;

[0085] An association processing module 502 is used to construct a classification head for performing binary classification processing on the true or false images, and to determine the association relationship between the low-rank adaptive model and the classification head;

[0086] An image acquisition module 503 is used to acquire an image set for training and perform image processing on each image in the image set to obtain an image group corresponding to each image in the image set;

[0087] The training processing module 504 is used to train the low-rank adaptive model and the classification head according to the image group, and obtain an image true and false detection model based on the trained low-rank adaptive model and the trained classification head when the training is completed.

[0088] In one embodiment, the image acquisition module 503 is further configured to:

[0089] Obtain pre-labeled real images and fake images to form an image set for training;

[0090] Determine the number of models for low-rank adaptation models;

[0091] Perform image augmentation processing on each image in the image set according to the number of models to obtain an augmented image corresponding to each image;

[0092] The augmented image is combined with the corresponding original image to obtain an image group for each image, wherein the number of images in the image group is the same as the number of models, and the original image is the image in the image set that has been augmented.

[0093] In one embodiment, the training processing module 504 is further configured to:

[0094] Initializing the low-rank adaptation model and the classification head to obtain an initialized low-rank adaptation model and an initialized classification head;

[0095] The initialized low-rank adaptive model and the initialized classification head are constructed based on the association relationship to obtain several training models, where each training model contains a low-rank adaptive model and a classification head;

[0096] The training model is trained according to the image group, and when the training is completed, an image true and false detection model is obtained by constructing a model based on the trained training model.

[0097] In one embodiment, the training processing module 504 is further configured to:

[0098] Input each image in the image group into the training model to obtain the predicted probability of each model in the training model for the image group;

[0099] Obtain true or false prediction results of the image group based on the prediction probability;

[0100] The training progress of the training model is judged according to the true and false prediction results and the true and false actual results of the image group, and when the training progress is determined to be completed, an image true and false detection model constructed based on the trained training model is obtained, wherein the training progress includes training completed and training incomplete.

[0101] In one embodiment, the training processing module 504 is further configured to:

[0102] Get the default weights of the classification head in the training model, and perform weighted calculation based on the default weights and prediction probabilities to obtain the true or false prediction probabilities of the image group.

[0103] In one embodiment, the training apparatus 500 for the image authenticity detection model further includes a weight fine-tuning module for:

[0104] Construct a weighted combination of classification heads in the image authenticity detection model, where each weight in the weighted combination corresponds to a classification head;

[0105] Adjust the image true / false detection model based on the weight combination to obtain an adjusted image true / false detection model;

[0106] Obtain a test image set and perform a model test on the adjusted image authenticity detection model based on the test image set to determine the degree of prediction fit of the weight combination;

[0107] According to the degree of prediction fit, a weight combination for adjusting the classification head in the training model is obtained, and based on the weight combination, the image true and false detection model is fine-tuned to obtain the adjusted image true and false detection model.

[0108] In one embodiment, the image authenticity detection model training device 500 further includes a detection processing module for:

[0109] Obtain the image to be detected, and input the image to be detected into the image true or false detection model to obtain the predicted probability value of the image to be detected;

[0110] According to the predicted probability value, the authenticity detection result of the image to be detected is obtained;

[0111] If the predicted probability value is in the true interval, the authenticity detection result of the image to be detected is determined to be a true image;

[0112] If the predicted probability value is in the false interval, the authenticity detection result of the image to be detected is determined to be a false image;

[0113] If the predicted probability value is outside the true interval and the false interval, it is determined that the authenticity detection result of the image to be detected is a blurred image.

[0114] Also, see Figure 6 , Figure 6 This is a structural diagram of an electronic device provided in an embodiment of the present application. The electronic device may be a mobile terminal such as a smart phone, a tablet computer, or the like. Figure 6 As shown, the electronic device 600 includes a processor 601 and a memory 602. The processor 601 is electrically connected to the memory 602.

[0115] The processor 601 is the control center of the electronic device 600. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or loading applications stored in the memory 602 and calling data stored in the memory 602, it executes various functions of the electronic device 600 and processes data, thereby monitoring the electronic device 600 as a whole.

[0116] In this embodiment, the processor 601 in the electronic device 600 will load the instructions corresponding to the processes of one or more applications into the memory 602 in accordance with the following steps, and the processor 601 will run the application stored in the memory 602, thereby implementing any step in the training method of the image authenticity detection model provided in the above embodiment.

[0117] The electronic device 600 can implement the steps in any embodiment of the training method of the image authenticity detection model provided in the embodiments of the present application. Therefore, it can achieve the beneficial effects that can be achieved by the training method of any image authenticity detection model provided in the embodiments of the present application. Please refer to the previous embodiments for details and will not be repeated here.

[0118] See Figure 7 , Figure 7 is another structural diagram of the electronic device provided in the embodiment of the present application, such as Figure 7 As shown, Figure 7 The electronic device 700 can be a mobile terminal such as a smartphone or a laptop computer.

[0119] RF circuit 710 is used to receive and transmit electromagnetic waves, converting electromagnetic waves into electrical signals, thereby communicating with a communication network or other devices. RF circuit 710 may include various existing circuit components for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, memory, and the like. RF circuit 710 can communicate with various networks such as the Internet, an intranet, or a wireless network, or communicate with other devices via a wireless network. Such wireless networks may include cellular telephone networks, wireless local area networks, or metropolitan area networks. The wireless networks may utilize various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE802.11g, and / or IEEE802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messaging, and any other suitable communication protocols, including those currently undeveloped.

[0120] The memory 720 can be used to store software programs and modules, such as the program instructions / modules corresponding to the training method of the image authenticity detection model in the above-mentioned embodiment. The processor 780 executes various functional applications and the training method of the image authenticity detection model by running the software programs and modules stored in the memory 720.

[0121] The memory 720 may include a high-speed random access memory (RAM) and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 720 may further include a memory remotely located relative to the processor 780, and such remote memory may be connected to the electronic device 700 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0122] The input unit 730 can be used to receive uploaded digital or character information and generate keyboard, mouse, joystick, optical, or trackball signal input related to user settings and function control. Specifically, the input unit 730 may include a touch-sensitive surface 731 and other input devices 732. The touch-sensitive surface 731, also known as a touch display or touchpad, can detect user touch operations on or near it (for example, operations performed by a user using a finger, stylus, or any other suitable object or accessory on or near the touch-sensitive surface 731) and drive corresponding connected devices according to a pre-set program. Optionally, the touch-sensitive surface 731 may include a touch detection device and a touch controller. The touch detection device detects the user's touch position and detects signals generated by the touch operation, transmitting the signals to the touch controller. The touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 780. It can also receive and execute commands from the processor 780. In addition, the touch-sensitive surface 731 can be implemented using various types, including resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 731, the input unit 730 may further include other input devices 732. Specifically, the other input devices 732 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, and a joystick.

[0123] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device 700. These graphical user interfaces can be composed of graphics, text, icons, videos, or any combination thereof. The display unit 740 may include a display panel 741. Optionally, the display panel 741 can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), or the like. Furthermore, the touch-sensitive surface 731 can cover the display panel 741. When the touch-sensitive surface 731 detects a touch operation on or near it, it transmits the information to the processor 780 to determine the type of touch event. The processor 780 then provides a corresponding visual output on the display panel 741 based on the type of touch event. Although the touch-sensitive surface 731 and the display panel 741 are shown in the figure as two independent components to implement input and output functions, in some embodiments, the touch-sensitive surface 731 and the display panel 741 can be integrated to implement input and output functions.

[0124] The electronic device 700 may also include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor may generate an interrupt when the flip cover is closed or closed. As a type of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the electronic device 700 may also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be described here.

[0125] Audio circuit 760, speaker 761, and microphone 762 provide an audio interface between the user and electronic device 700. Audio circuit 760 can convert received audio data into electrical signals and transmit them to speaker 761, which then converts them into sound signals for output. Microphone 762, on the other hand, converts collected sound signals into electrical signals, which are then received by audio circuit 760 and converted into audio data. The audio data is then processed by processor 780 and transmitted via RF circuit 710 to, for example, another terminal. Alternatively, the audio data can be output to memory 720 for further processing. Audio circuit 760 may also include an earphone jack to allow communication between external headphones and electronic device 700.

[0126] Electronic device 700 can help users receive requests, send information, etc. through a transmission module 770 (e.g., a Wi-Fi module), providing users with wireless broadband Internet access. Although transmission module 770 is shown in the figure, it is understandable that it is not a required component of electronic device 700 and can be omitted as needed without changing the essence of the invention.

[0127] Processor 780 is the control center of electronic device 700. It connects all components of the phone using various interfaces and circuits. By running or executing software programs and / or modules stored in memory 720 and accessing data stored in memory 720, it executes various functions of electronic device 700 and processes data, thereby providing overall monitoring of the electronic device. Optionally, processor 780 may include one or more processing cores. In some embodiments, processor 780 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 780.

[0128] The electronic device 700 also includes a power supply 790 (e.g., a battery) for supplying power to various components. In some embodiments, the power supply can be logically connected to the processor 780 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 790 can also include any components such as one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0129] Although not shown, the electronic device 700 further includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal further includes a memory and one or more programs, wherein the one or more programs are stored in the memory and are configured so that one or more processors execute the one or more programs to implement any step of the training method of the image authenticity detection model provided in the above embodiment.

[0130] During specific implementation, the above modules can be implemented as independent entities, or can be arbitrarily combined and implemented as the same or several entities. The specific implementation of the above modules can be found in the previous method embodiments and will not be repeated here.

[0131] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be accomplished by instructions, or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present application provides a storage medium storing a plurality of instructions that, when executed by a processor, can implement any step in the training method for the image authenticity detection model provided in the above embodiments.

[0132] The storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0133] Since the instructions stored in the storage medium can execute the steps in any embodiment of the training method of the image authenticity detection model provided in the embodiments of the present application, the beneficial effects that can be achieved by the training method of any image authenticity detection model provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0134] The above is a detailed introduction to the training method, device, electronic device and storage medium of an image authenticity detection model provided by the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of ​​the present application. At the same time, for those skilled in the art, according to the ideas of the present application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present application. Moreover, for those of ordinary skill in the art, without departing from the principles of the present application, several improvements and modifications can be made, and these improvements and modifications are also considered to be within the scope of protection of the present application.

Claims

1. A training method for an image authenticity detection model, characterized in that: include: Obtaining a pre-trained visual feature extraction model, and fine-tuning the visual feature extraction model based on a low-rank decomposition matrix to obtain several low-rank adaptation models; Constructing a classification head for performing binary classification processing on the true and false images, and determining the correlation relationship between the low-rank adaptation model and the classification head; Acquire a training image set, and perform image processing on each image in the image set to obtain an image group corresponding to each image in the image set; The low-rank adaptive model and the classification head are trained according to the image group, and when the training is completed, an image true and false detection model based on the trained low-rank adaptive model and the trained classification head is obtained.

2. The method according to claim 1, wherein The acquiring of the image set for training and performing image processing on each image in the image set to obtain an image group corresponding to each image in the image set includes: Obtain pre-labeled real images and fake images to form an image set for training; determining the number of models of the low-rank adaptive model; Performing image augmentation processing on each image in the image set according to the number of models to obtain an augmented image corresponding to each image; The augmented image is combined with the corresponding original image to obtain an image group for each image, wherein the number of images in the image group is the same as the number of the models, and the original image is the image in the image set that has been augmented.

3. The method according to claim 1, wherein The low-rank adaptive model and the classification head set are trained according to the image group, and an image true / false detection model based on the trained low-rank adaptive model and the trained classification head is obtained upon completion of the training, including: Initializing the low-rank adaptive model and the classification head to obtain an initialized low-rank adaptive model and an initialized classification head; Constructing the initialized low-rank adaptive model and the initialized classification head based on the association relationship to obtain a plurality of training models, wherein each training model includes a low-rank adaptive model and a classification head; The training model is trained according to the image group, and when the training is completed, an image true and false detection model is obtained by constructing a model based on the trained training model.

4. The method according to claim 3, wherein The training model is trained according to the image group, and when the training is completed, a true or false image detection model is obtained by constructing a model based on the trained training model, including: Inputting each image in the image group into the training model respectively to obtain the prediction probability of each model in the training model for the image group; Obtaining true or false prediction results of the image group according to the prediction probability; The training progress of the training model is judged according to the true and false prediction results and the true and false results of the image group, and when it is determined that the training progress is completed, an image true and false detection model constructed based on the trained training model is obtained, wherein the training progress includes training completed and training incomplete.

5. The method according to claim 4, wherein Obtaining true or false prediction results of the image group according to the prediction probability includes: Obtaining the default weights of the classification heads in the training model, and performing weighted calculation based on the default weights and the prediction probabilities to obtain true or false prediction probabilities for the image group.

6. The method according to claim 1, wherein The method further comprises: Constructing a weighted combination of classification heads in the image authenticity detection model, wherein each weight in the weighted combination corresponds to a classification head; Performing model adjustment on the image authenticity detection model based on the weight combination to obtain an adjusted image authenticity detection model; Acquire a test image set, and perform a model test on the adjusted image authenticity detection model based on the test image set to determine the prediction fit of the weight combination; A weight combination for adjusting the classification head in the training model is obtained according to the predicted fit degree, and the image true and false detection model is fine-tuned based on the weight combination to obtain an adjusted image true and false detection model.

7. The method according to claim 1, wherein The method further comprises: Acquire an image to be detected, and input the image to be detected into the image true or false detection model to obtain a predicted probability value of the image to be detected; Obtaining a authenticity detection result of the image to be detected according to the predicted probability value; If the predicted probability value is in the true interval, determining that the authenticity detection result of the image to be detected is a true image; If the predicted probability value is in the false interval, determining that the authenticity detection result of the image to be detected is a false image; If the predicted probability value is outside the true interval and the false interval, it is determined that the authenticity detection result of the image to be detected is a blurred image.

8. A training device for an image authenticity detection model, characterized in that: include: A fine-tuning processing module is used to obtain a pre-trained visual feature extraction model and fine-tune the visual feature extraction model based on a low-rank decomposition matrix to obtain a plurality of low-rank adaptation models; An association processing module, configured to construct a classification head for performing binary classification processing on the authenticity of an image, and determine an association relationship between the low-rank adaptation model and the classification head; An image acquisition module is used to acquire an image set for training and perform image processing on each image in the image set to obtain an image group corresponding to each image in the image set; A training processing module is used to train the low-rank adaptive model and the classification head according to the image group, and obtain an image true and false detection model based on the trained low-rank adaptive model and the trained classification head when the training is completed.

9. An electronic device, characterized in that: The electronic device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 7 are implemented.