An Artificial Intelligence Fake Content Detection Method Based on the Joint Decision-making of Multiple Expert Models
By building a joint decision-making framework for multi-expert models, combining multi-expert collaboration mechanism and robustness optimization, the problems of poor robustness and low fault tolerance in the existing technology are solved, and efficient detection and comprehensive judgment of diversified umbrella types are achieved.
Patent Information
- Application Number
- CN202510310700.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing forged content detection technology is poor in the face of real scenarios, and the fault tolerance of methods based on single-model discrimination is low, making it difficult to adapt to diverse forgery types and common post-processing operations.
Design an artificial intelligence forged content detection method for joint decision-making of multi-expert models, build a joint decision-making framework for general and specialized expert models, combine multi-expert collaboration mechanisms and robust optimization to improve the robustness and adaptability of the detection system.
Through joint decision-making of multiple expert models, the detection system's adaptability and robustness to diversified forgery types is improved, and it can efficiently detect multi-type forgery content in complex scenarios, providing system-level solutions.
Smart Images

Figure CN119832552B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically refers to an artificial intelligence forged content detection method for joint decision-making of multiple expert models. Background Art
[0002] With the rapid development of generative artificial intelligence technology, various types of deepfake content have spread widely in social networks. Although these technologies have positive applications in creative design and life entertainment, the forged content also brings ethical, privacy, and security risks to society. Malicious content such as false news and sensitive topics has greatly affected social stability and national security. Therefore, it is imperative to invent detection algorithms for forged content. In a real environment, the forgery methods are complex and diverse, making it impossible for a specialized single detection model to effectively identify all forged content. Moreover, the forged content is often spread multiple times through social networks and often undergoes various post-processing operations, such as compression, cropping, scaling, etc., further increasing the difficulty of detecting forged content.
[0003] Generative artificial intelligence content (AIGC) is an important branch in the field of artificial intelligence, which generates multi-modal content such as images, texts, audios, and videos through deep learning technology. The AIGC technology has not only greatly promoted the development of fields such as creative design and film and television production but also shown broad prospects in application scenarios such as medical imaging and virtual reality. In recent years, with the innovation of technologies such as generative adversarial networks (GANs) and diffusion models, the AIGC technology has developed rapidly, and the quality, fidelity, and diversity of the generated content have been continuously improved.
[0004] As a representative forgery generation method, the generative adversarial network consists of a generator and a discriminator. Through the adversarial training of the two, the generator is continuously optimized to generate realistic content, while the discriminator improves its detection ability by distinguishing between the generated content and the real content. This adversarial mechanism has brought the content generated by GAN to a new height in terms of quality and authenticity. In recent years, the StyleGAN series of models has become one of the representative works of GANs. It introduces technologies such as hierarchical feature control and style encoding, enabling the generated images to not only have higher resolutions but also achieve fine-grained content editing, such as changing the expression, age, or background style of a face. In addition, the BigGAN model has significantly improved the diversity and clarity of the generated content through the optimization of the model architecture and larger-scale training. The innovation of these methods has greatly expanded the application scenarios of GANs, such as virtual human production, anime image generation, image restoration, and super-resolution reconstruction. However, this improvement in generation ability has also facilitated the production of deepfake content, making it more difficult to identify forged images and videos.
[0005] Diffusion models are a type of generation method based on step-by-step denoising. Through a gradual restoration process from the noise distribution to the target data distribution, they can generate realistic images. The principle is to gradually add Gaussian noise to the input data to generate noisy data, and then remove the noise through the inverse process to finally restore the target data. In recent years, diffusion models have become an emerging main force in AIGC technology due to their stable training process and excellent generation quality. Methods such as DALL-E and Stable Diffusion are typical representatives of diffusion models. They combine the Transformer architecture and use large-scale multimodal data for training, and can generate target images according to text descriptions. Such diffusion models have surpassed traditional GAN models in the ability to generate complex textures and fine details, and are widely used in advertising design, game scene creation, etc. However, after the content generated by diffusion models undergoes compression, cropping, or other post-processing, its forgery features become more concealed, posing higher requirements for forgery detection.
[0006] Existing methods for detecting generative models can be divided into two categories: methods based on handcrafted features and methods based on deep learning.
[0007] The handcrafted feature methods analyze the forgery traces or specific patterns in AIGC content, extract features with forgery detection capabilities, and use traditional machine learning models for classification. For example, by studying the forgery traces in the frequency domain of images, inconsistent lighting and geometric relationships, or detecting statistical anomalies brought by specific generative models, such as color distribution, edge smoothness, pixel consistency, etc. The main feature of this type of method is that it relies on the knowledge of domain experts to design feature engineering, and has good interpretability. Its advantages include intuitive and clear feature design, good performance on small-scale datasets, and low computational complexity. However, the handcrafted feature methods cannot adapt to new forgery models, and when the content undergoes post-processing such as compression, cropping, and scaling, the detection effect will decrease significantly. In addition, with the improvement of the generation quality of AIGC, the forgery features tend to become more concealed, and it is difficult to meet the high-precision detection requirements only relying on handcrafted features.
[0008] Neural network-based methods automatically extract the latent features of AIGC content through deep learning models and train a forgery detection network with classification capabilities. For example, based on convolutional neural networks (CNNs) to detect artifacts in forged images, and time series models (such as RNNs) to analyze the temporal characteristics of forged videos. Compared with manual feature methods, the advantage of neural network methods lies in their powerful feature extraction ability and adaptability to multi-modal complex data. Deep features in the neural network can be automatically extracted only through training, without relying on complex manual feature design. Moreover, deep models can learn subtle patterns in forged content that are difficult for humans to detect and have better detection ability for high-quality AIGC content. In addition, neural network methods can use large-scale datasets for end-to-end training to optimize performance. However, its disadvantages include high demand for computing resources, long training and inference times, and weak interpretability of the model. More importantly, when the forged content undergoes operations such as compression and cropping, the detection performance of neural network methods will also decline. At the same time, the generalization ability of deep models to unseen forgery types or generation models may be limited.
[0009] According to the above description, the existing technical solutions mainly have the following defects:
[0010] The current forgery detection technology has poor robustness in the face of real-world scenarios. Since forged content is usually spread through social media, and common social media usually performs post-processing operations such as compressing multimedia content, and is often subject to operations such as zooming and cropping by users during the dissemination process. This series of post-processing operations reduces the robustness of the model and results in poor detection performance.
[0011] Existing methods usually rely on single-model discrimination, resulting in a low error tolerance rate. Due to the diverse forms of forged content, a single model usually has difficulty covering the entire feature space, making the model vulnerable to adversarial samples. And since forged content spreads rapidly and causes great harm, the omission of the model often brings huge losses.
[0012] Current detection methods usually only target specific scenarios and lack consideration of system-level design. Due to the rapid development and innovation of forgery techniques, a single detection method is difficult to cover all application scenarios. Moreover, the detection methods do not consider system interface design and scalability design, making it difficult for new detection methods to be integrated into existing systems, resulting in difficulties in iterating the detection system and being unfavorable for the long-term maintenance of the system.
[0013] Therefore, the present invention designs an artificial intelligence forged content detection system with joint decision-making of multiple expert models, constructs a joint decision-making framework for expert models from a system-level perspective, gives full play to the advantages of different models in identifying forged content, and improves the adaptability of the detection system to diverse forgery types. At the same time, for the three common post-processing operations of cropping, compression, and scaling, their features are integrated into the expert models to improve the robustness and accuracy of the detection system when facing post-processed content. This innovative method helps to achieve more robust forgery content identification in complex and changing real-world scenarios and also provides a system-level solution for the task of forged content detection. Summary of the Invention
[0014] In view of the deficiencies of the prior art, the present invention proposes an artificial intelligence forged content detection method with joint decision-making of multiple expert models, constructs a joint decision-making framework for expert models from a system-level perspective, gives full play to the advantages of different models in identifying forged content, and improves the adaptability of the detection system to diverse forgery types. At the same time, for the three common post-processing operations of cropping, compression, and scaling, their features are integrated into the expert models to improve the robustness and accuracy of the detection system when facing post-processed content. By combining the multi-expert cooperation mechanism and robustness optimization, the present invention aims to establish a forged content detection system with high scalability and reliability, which can achieve efficient detection and comprehensive judgment of multiple types of forged content in complex application scenarios.
[0015] To solve the above technical problems, the technical solution of the present invention is as follows:
[0016] An artificial intelligence forged content detection method with joint decision-making of multiple expert models, comprising the following steps:
[0017] Step 1, construct and train a general expert model group, the general expert model group includes several general expert models, the general expert models use ConvNeXt as the backbone network, and the cross-entropy loss function is used as a constraint during the training of the general expert models;
[0018] Step 2, construct and train a specialized expert model group, the specialized expert model group includes several specialized expert models, the specialized expert models use ConvNeXt as the backbone network, the training data set for training the specialized expert models contains real samples and forged samples, the forged samples are marked with corresponding forgery type labels, and during the training process, a corresponding specialized expert model is trained for each forgery type label, and the cross-entropy loss function is used as a constraint;
[0019] Step 3: Obtain the multimedia content to be detected. First, it will be detected by the specialized expert model group. For the multimedia content to be detected, it will be input into each type of specialized expert model. When a certain type of specialized expert model detects that it exceeds the forgery threshold, it will directly jump to the confidence calculation link; if the content to be detected passes the detection of all specialized expert models, it will enter the general expert model detection process;
[0020] Step 4: After passing the detection of the specialized expert model, enter the general expert model group for detection. Input the content to be detected into the general expert model group. If the forgery threshold is reached, it will be marked and then enter the confidence calculation link;
[0021] Step 5: When the content to be detected goes through the detection processes of the first two types of expert model groups, perform confidence calculation and finally feedback to the user.
[0022] Preferably, a robust expert model is introduced into the general expert model group and the specialized expert model corresponding to each type of forgery type.
[0023] Preferably, the robust expert model uses ConvNeXt as the backbone network and uses the cross-entropy loss function as a constraint during training.
[0024] Preferably, the general expert model, the specialized expert model, and the robust expert model all adopt any one of the three variant models of ConvNeXt-Small, ConvNeXt-Base, and ConvNeXt-XL.
[0025] Preferably, in step 1, the training method of the general expert model is: obtain the training data set , and preprocess the training data set through horizontal flipping, vertical flipping, and rotation. During training, given the samples in the training set , 0 represents a real image, 1 represents a forged image, is the prediction probability of the model, and the loss function can be expressed by the following formula:
[0026] ;
[0027] Among them, represents the number of samples. By minimizing this loss function, when a real sample is input, the value tends to 0, and when a forged sample is input, the value approaches 1, that is, to distinguish real content and forged content.
[0028] Preferably, for the robust expert model introduced in the general expert model group, the training data set is obtained through data augmentation to get the training data set , and the training data set Train the robust expert model introduced in the general expert model group, and its training method is the same as that of the general expert model.
[0029] Preferably, in step 2, obtain the training data set , and preprocess the training data set by horizontal flipping, vertical flipping, and rotation. The training data set includes real samples and forged samples. The forged type label corresponding to the forged sample is denoted as . Divide all the forged sample data sets into several subsets according to the forged label , that is . During the training process, train a corresponding specialized expert model for each forged type .
[0030] Preferably, for the robust expert model introduced by the specialized expert model corresponding to each forged type, obtain the training data set through data augmentation to obtain the training data set . Apply the training data set to train the robust expert model introduced by the specialized expert model corresponding to each forged type, and its training method is the same as that of the general expert model.
[0031] Preferably, the real samples in the training data set are shared with the training data set .
[0032] Preferably, in step 5, when the input content is determined to be forged content by the specialized expert model or the robust expert model corresponding to the forged type, the confidence adopts the average confidence of this group of specialized expert models.
[0033] Preferably, the calculation method of the confidence is: Denote that this group of specialized expert models contains sub-models . Take the average of the confidence of the models in the group, and the calculation method is as follows:
[0034] ;
[0035] where is the input sample, the output confidence of the th expert model in the group, is the total number of the specialized expert models and the robust expert models in this group.
[0036] Preferably, when the input content is not judged as forged content by the specialized expert model or its corresponding robust expert model, the output confidence level is obtained by taking the average of the output confidence levels of the general expert model group and its corresponding robust expert model. The specific formula is as follows:
[0037] ;
[0038] where is the input sample, is the output confidence level of the general expert model, is the total number of the general expert model group and its corresponding robust expert model.
[0039] The present invention designs a multi-expert model training strategy for forged content detection, improves the limitations of a single model in detecting forged content, gives full play to the advantages of different models in identifying forged content, and enhances the adaptability of the detection system to diverse forgery types;
[0040] designs a forged content discrimination mechanism based on the collaboration of multi-expert models, constructs a joint decision-making framework for expert models from a system-level perspective, and gives full play to the advantages of specialized expert models and general expert models in detection, so as to improve the overall detection speed and performance of the system.
[0041] Integrates three post-processing operations of scaling, cropping, and compression into the expert model, making the system more compatible with real scenarios and improving the robustness of the system.
[0042] (1) Adopt multi-expert model collaboration to detect forged content, improve the limitations of a single model in detecting forged content, and enhance the generalization of the system.
[0043] (2) Add specialized models for compression, cropping, and scaling operations to cope with common user behaviors in real scenarios and improve the robustness of the system.
[0044] (3) The invention is not only designed for a single scenario, but from a system-level perspective. The unique organizational structure of the system allows it to dynamically add and change models during operation, enabling the system to face emerging new forgery methods in the future. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0046] Figure 1Schematic flow diagram of the general expert model in the embodiments of the present invention;
[0047] Figure 2 Schematic flow diagram of the specialized patent model in the embodiments of the present invention;
[0048] Figure 3 Schematic flow diagram of the joint decision-making of multiple expert models in the embodiments of the present invention. Detailed implementation manners
[0049] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0050] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.
[0051] In the description of the present invention, it should be noted that, unless otherwise clearly defined and limited, the terms "mounted", "connected", "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection, an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations. Embodiment
[0052] This embodiment provides an artificial intelligence forged content detection method for joint decision-making of multiple expert models, which consists of two parts: training of multiple expert models and joint decision-making of multiple expert models.
[0053] (1) Training of multiple expert models: The training of multiple expert models aims to formulate a suitable training method, divide different expert models, and prepare suitable data for each model to be trained, so as to achieve the generalization detection and specialization detection of forged content, and improve the reliability and robustness of the system.
[0054] (1.1) General expert model training: This type of expert model aims at detecting forged content in general scenarios and relies on large-scale and diverse general datasets for training. The model extracts texture features from the data and analyzes the differences between forged content and real content, so as to accurately distinguish whether the multimedia content is forged. As Figure 1 shown, the specific training steps are as follows:
[0055] (1.1.1) First, perform preprocessing operations on the training data , including horizontal flipping, vertical flipping, rotation and other operations. These operations generate diverse data samples, provide richer learning materials for subsequent model training, and improve the robustness and adaptability of the model.
[0056] (1.1.2) Then, send the data into the model for training. Since the goal at this stage is to train several general expert models to identify various types of forged content, and forged content usually differs from real content in texture information, ConvNeXt is used as the backbone network. The ConvNeXt network is based on convolutional operations and can well capture local texture details in images, identifying the subtle differences between forged content and real content. At the same time, there are various variants of ConvNeXt, which are suitable for different computing resources and task requirements. In the present invention, three variants, namely ConvNeXt-Small, ConvNeXt-Base, and ConvNeXt-XL, are adopted to provide corresponding solutions for applications of different scales.
[0057] For the problem of forged content detection, this task can be regarded as a binary classification problem: real content and forged content. Therefore, the cross-entropy loss function is adopted as the constraint during training. Given , 0 represents a real image, 1 represents a forged image, is the predicted probability of the model, and the loss function can be expressed by the following formula:
[0058]
[0059] where represents the number of samples. By minimizing this loss function, when a real sample is input, the value tends to 0, and when a forged sample is input, the value approaches 1, that is, to distinguish real content and forged content.
[0060] It is understandable that after being optimized by the loss function, the neural network can distinguish samples of different categories. Minimizing the loss function is the way of optimization. The loss is used to evaluate the degree of error of the training samples, that is, the training quality of the neural network. The greater the loss, the more training samples are misjudged by the neural network, that is, the training is not good. Minimizing the loss function means that the training samples have the fewest errors, so the better the proof effect, which represents that the neural network is better trained.
[0061] (1.2) Specialized expert model training: This type of expert model is mainly for the detection of specific forgery types, providing reliable judgments in specific scenarios, so that the system can maintain reliable judgments when facing various known forgery types, such as Figure 2 As shown, the specific steps are as follows:
[0062] (1.2.1) First, perform preprocessing operations on the training data including horizontal flipping, vertical flipping, rotation and other operations. These operations generate diverse data samples, providing richer learning materials for subsequent model training and improving the robustness and adaptability of the model.
[0063] (1.2.2) Then, send the data into the model for training. Since the forged content usually has differences in texture information from the real content, ConvNeXt is used as the backbone network. And consistent with what is mentioned in (1.1.2), three variants of ConvNeXt-Small, ConvNeXt-Base, and ConvNeXt-XL are adopted to provide corresponding solutions for different scale applications.
[0064] Since the goal at this stage is to train several specialized expert models to identify specified types of forged content, specific data needs to be prepared for each expert model. The specific division rules are as follows: For real samples, all expert models share this part of the data so that they can widely learn the general features of real samples; for forged samples, the forged type label corresponding to the forged sample is denoted as , and all forged sample data sets are divided into several subsets according to the forged label , that is , and during the training process, a corresponding specialized expert model is trained for each forged type , and the corresponding data sets used include and . The loss function used in training is the cross-entropy loss function, and the formula is as follows:
[0065]
[0066] Where Indicates the number of samples. Given , 0 represents the real image, 1 represents the forged image, is the prediction probability of the model. By minimizing this loss function, the model can effectively distinguish between real and forged content.
[0067] (2) Joint decision-making of multi-expert models: The joint detection of multi-expert models aims to jointly detect the input image by a certain mechanism, combine each expert model, and give the final detection confidence through a certain strategy. The specific process is as Figure 2 shown:
[0068] (2.1) Detection by specialized expert models: The multimedia content to be detected will first be detected by specialized expert models. Such models have extremely high detection accuracy for known forgery types. The specific process is as follows: For the multimedia content to be detected, input it into each type of specialized expert model . When the detection of a certain type of specialized expert model exceeds the forgery threshold, directly jump to the confidence calculation link. If the content to be detected passes the detection of all specialized expert models, it enters the general expert model detection process.
[0069] It should be noted that the forgery threshold of the specialized expert model is the probability value output by the specialized expert model, that is, the confidence level. In this embodiment, the forgery threshold of the specialized expert model is set to 0.5.
[0070] (2.2) Detection by general expert models: After passing the specialized expert detection model, enter the general expert model detection. Such expert models are designed for general scenario detection. Compared with the specialized expert models, although the detection ability of known forgery types is weakened, they can analyze general forgery traces and better detect general forged content. Input the content to be detected into the general expert model. If the forgery threshold is reached, mark it and enter the confidence calculation link.
[0071] It should be noted that the forgery threshold of the general expert model is the probability value output by the general expert model, that is, the confidence level. In this embodiment, the forgery threshold of the general expert model is set to 0.5.
[0072] (2.3) Confidence calculation: When the content to be detected passes the detection processes of the first two types of expert models, the confidence level is obtained through a certain mechanism and finally fed back to the user. The specific confidence calculation method is as follows:
[0073] When the input content is determined to be forged content by the specialized expert model, the confidence level adopts the average value of the confidence levels of this group of expert models. Specifically, denote this group of expert models as including sub-models . Take the average of the model confidence levels in the group. The formula is as follows:
[0074]
[0075] Where is the input sample, is the output confidence level of the th expert model in the group, and is the total number of specialized expert models in this group.
[0076] When the input content is not judged as forged content by the specialized expert model, the output confidence level is obtained by taking the average of the output confidence levels of the general expert models. The specific formula is as follows:
[0077]
[0078] Where is the input sample, is the output confidence level of the general expert model, and is the total number of general expert models in this group. Embodiment
[0079] The difference between this embodiment and Embodiment 1 is that robust expert models are introduced into the general expert model group and the specialized expert models corresponding to each type of forgery.
[0080] (1.3) Robust expert model training: This type of model mainly targets forged content that has undergone post-processing operations such as cropping, compression, and scaling. It is mainly used to enhance the application effect of the first two types of expert models in complex scenarios, making the system more adaptable to real scenarios such as social platforms. This type of expert model is usually integrated into the system as a variant of the first two types of expert models. Therefore, during training, for the robust expert models introduced into the general expert model group, the training dataset is obtained through data augmentation to get the training dataset , and the training dataset is used to train the robust expert models introduced into the general expert model group.
[0081] For the robust expert models introduced into the specialized expert models corresponding to each forgery type, the training dataset is obtained through data augmentation to get the training dataset , and the training dataset is used to train the robust expert models introduced into the specialized expert models corresponding to each forgery type.
[0082] The specific training steps are as follows:
[0083] (1.3.1)First, prepare the training data and perform preprocessing operations on the training data, including horizontal flipping, vertical flipping, rotation, etc. This is consistent with (1.1.1) and (1.2.1).
[0084] (1.3.2)Perform specified enhancement operations on the data. In real scenarios, compression, cropping, and scaling are the most common post - processing operations and have a great impact on the detection performance. Therefore, during the training of this type of expert model, corresponding post - processing operations are applied to the data and then trained to enable the model to learn the forgery traces after these post - processing operations, enhancing the robustness of the model in real application scenarios.
[0085] (1.3.3)Feed the enhanced data into the model for training. The training model uses ConvNeXt as the backbone network. And as mentioned in (1.1.2), three variants, ConvNeXt - Small, ConvNeXt - Base, and ConvNeXt - XL, are adopted to provide corresponding solutions for applications of different scales.
[0086] For the joint decision - making stage of the multi - expert model, the multimedia content to be detected will first be detected by the specialized expert model group. For the multimedia content to be detected, it will be input into each type of specialized expert model and its corresponding robust expert model. When a certain type of specialized expert model or its corresponding robust expert model detects that it exceeds the forgery threshold, it will directly jump to the confidence calculation link; if the content to be detected passes the detection of all the specialized expert model groups, it will enter the general expert model group detection process;
[0087] When the input content is determined to be forged content by the specialized expert model or the robust expert model corresponding to the forgery type, the confidence adopts the average value of the confidence of this group of specialized expert models.
[0088] The calculation method of the confidence is as follows: Denote this group of specialized expert models containing sub - models , take the average value of the confidence of the models in the group, and the calculation method is as follows:
[0089] ;
[0090] where is the input sample, the output confidence of the th expert model in the group, is the total number of this group of specialized expert models and robust expert models.
[0091] After being detected by the specialized expert model, it enters the detection of the general expert model group. The content to be detected is input into the general expert model group. If the general expert model or its corresponding robust expert model in the general expert model group reaches the forgery threshold, it is marked and then enters the confidence calculation link.
[0092] When the input content is not judged as forged content by the specialized expert model or its corresponding robust expert model, the confidence level is output. It is obtained by taking the average value of the output confidence levels of the general expert model group and its corresponding robust expert model. The specific formula is as follows:
[0093] ;
[0094] Among them is the input sample, is the output confidence level of the general expert model, is the total number of the general expert model group and its corresponding robust expert model in this group.
[0095] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations can be made to these embodiments including components, and they still fall within the protection scope of the present invention.
Claims
1. An artificial intelligence forged content detection method based on joint decision-making of multiple expert models, characterized in that: The steps include: Step 1: construct and train a general expert model group, wherein the general expert model group includes several general expert models, wherein the general expert model uses ConvNeXt as a backbone network, and a cross entropy loss function is used as a constraint during the training of the general expert model; Step 2, construct and train a specialized expert model group, wherein the specialized expert model group includes several specialized expert models, wherein the specialized expert model uses ConvNeXt as a backbone network, and the training data set used to train the specialized expert model includes real samples and forged samples, wherein the forged samples are marked with corresponding forged type labels, and during the training process, a corresponding specialized expert model is trained for each forged type label, and a cross entropy loss function is used as a constraint; Step 3: Obtain the multimedia content to be detected. First, it will be tested by the specialized expert model group. For the multimedia content to be detected, it will be input into each type of specialized expert model. When the detection of a certain type of specialized expert model exceeds the forgery threshold, it will directly jump to the confidence calculation link; if the content to be detected passes the detection of all specialized expert models, it will enter the general expert model detection process; Step 4: After passing the specialized expert model detection, enter the general expert model group detection, input the content to be detected into the general expert model group, and if it reaches the forgery threshold, it will be marked and enter the confidence calculation stage; Step 5: After the content to be detected has passed the detection process of the first two types of expert model groups, the confidence calculation is performed and finally fed back to the user.
2. According to the method of claim 1, the method is characterized by: The general expert model group and the specialized expert model corresponding to each forgery type are respectively introduced into the robust expert model.
3. According to claim 2, the artificial intelligence forged content detection method with joint decision-making of multiple expert models is characterized in that: The robust expert model uses ConvNeXt as the backbone network and uses the cross entropy loss function as a constraint during training.
4. According to claim 3, the artificial intelligence forged content detection method with joint decision-making of multiple expert models is characterized in that: The general expert model, specialized expert model and robust expert model all adopt any one of the three variant models of ConvNeXt-Small, ConvNeXt-Base and ConvNeXt-XL.
5. According to claim 4, the artificial intelligence forged content detection method with joint decision-making of multiple expert models is characterized in that: In step 1, the training method of the general expert model is: obtain a training data set A, and pre-process the training data set A by horizontal flipping, vertical flipping and rotation. During training, given a sample y in the training set i ∈{0,1}, 0 represents a real image, 1 represents a fake image, is the predicted probability of the model, and the loss function can be expressed by the following formula: Among them, N represents the number of samples. By minimizing the loss function, when real samples are input, the value tends to 0, and when forged samples are input, the value tends to 1, that is, distinguishing real content from forged content.
6. The artificial intelligence forged content detection method of multiple expert models joint decision-making according to claim 5 is characterized in that: For the robust expert model introduced in the general expert model group, the training data set A is obtained by data enhancement to obtain the training data set A′, and the training data set A′ is used to train the robust expert model introduced in the general expert model group. The training method is the same as the training method of the general expert model.
7. The artificial intelligence forged content detection method of multiple expert models joint decision-making according to claim 6 is characterized in that: In step 2, a training data set B is obtained and preprocessed by horizontal flipping, vertical flipping and rotation. The training data set B includes real samples and forged samples. The forged type label corresponding to the forged sample is denoted as L i , all the forged sample data sets D fake Press Fake Label L i Divide into several subsets, namely D fake ={D1,D2,…,D n }, during the training process, a corresponding specialized expert model M is trained for each forgery type i i .
8. The artificial intelligence forged content detection method of multiple expert models joint decision-making according to claim 7 is characterized in that: For the robust expert model introduced by the specialized expert model corresponding to each forgery type, the training data set B is obtained by data augmentation to obtain the training data set B′, and the training data set B′ is used to train the robust expert model introduced by the specialized expert model corresponding to each forgery type. The training method is the same as the training method of the general expert model.
9. The artificial intelligence forged content detection method of multiple expert models joint decision-making according to claim 7 is characterized in that: The real samples in the training data set B are shared with the training data set A.
10. The artificial intelligence forged content detection method of multiple expert models joint decision-making according to claim 8 is characterized in that: In step 5, when the input content is determined to be forged content by the specialized expert model corresponding to the forgery type or its corresponding robust expert model, the confidence C1 adopts the average confidence value of the group of specialized expert models.
11. The artificial intelligence forged content detection method of multiple expert models joint decision-making according to claim 10 is characterized in that: The confidence C1 is calculated as follows: i Contains sub-models {M i1 ,M i2 ,…,M in }, take the average model confidence in the group, and the calculation method is as follows: Where x is the input sample, M ij (x) is the output confidence of the jth expert model in the group, and n is the total number of specialized expert models and robust expert models in the group.
12. The artificial intelligence forged content detection method of multiple expert models joint decision-making according to claim 11 is characterized in that: When the input content is not judged as forged content by the specialized expert model or its corresponding robust expert model, the output confidence C2 is obtained by averaging the output confidences of the general expert model group and its corresponding robust expert model. The specific formula is as follows: Where x is the input sample, M k (x) is the output confidence of the general expert model, and m is the total number of general expert models and robust expert models in this group.
Citation Information
Patent Citations
Image detection model training method and device, equipment and storage medium
CN116958637A
Image authenticity detection model training method and image authenticity detection method
CN119274024A