Model training method, multimedia data classification method and related device
By using the feature difference training method of teacher network and student network in the field of AI, the problem of low accuracy of AI-generated image recognition in the prior art is solved, and more accurate recognition of multimedia data is achieved.
Patent Information
- Application Number
- CN202311577190.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the classification model used to identify images generated by AI-generated images has a low accuracy in recognition of images generated by unknown generators, making it difficult to achieve accurate recognition.
By obtaining the characteristics of multimedia data and inputting it into the teacher network and student network, the student network is trained based on the characteristic difference values output by the teacher network and student network, so that it can narrow or increase the characteristic difference when processing non-AI generation and AI generation of multimedia data, thereby improving the recognition accuracy.
The recognition accuracy of AI-generated multimedia data is improved, so that the model can more effectively deal with various types of AI-generated multimedia data.
Smart Images

Figure CN120031102A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Artificial Intelligence (AI) technology, and in particular to a model training method, a multimedia data classification method and related devices. Background Art
[0002] With the development of AI technology, more and more AI models are playing an important role in people's lives. For example, image classification models can help people classify photos in smartphones; voice recognition models can provide people with functions such as human-computer dialogue or voice remote control. However, although AI models have brought a lot of convenience to human life, they also have the potential to have a negative impact on society.
[0003] Specifically, current AI generation models such as image generation models and text generation models can generate realistic images or texts based on human instructions, so that the generated images or texts are used to create malicious fake news.
[0004] In order to reduce the negative impact of AI-generated images on society, classification models are used in related technologies to identify AI-generated images. The classification model is trained by using real images and AI-generated images to form a training set. However, the classification model in the related technology has a low accuracy rate in classifying images generated by unknown generators (i.e., AI-generated images that do not appear in the training set), which makes it difficult to accurately identify AI-generated images in some scenarios. Summary of the invention
[0005] The present application provides a model training method that can improve the recognition accuracy of the trained model for AI-generated multimedia data.
[0006] The first aspect of the present application provides a model training method for training a model for identifying multimedia data types in the field of AI. The model training method comprises: first, acquiring multimedia data and extracting features of the multimedia data, where the multimedia data is training data in a training set. Extracting features of multimedia data may refer to performing feature extraction on multimedia data through a pre-trained feature extraction network, or may refer to converting multimedia data into a feature matrix that can be used as a neural network input.
[0007] Then, the features of the multimedia data are input into the teacher network and the student network respectively, and the first feature output by the teacher network and the second feature output by the student network are obtained. Wherein, both the teacher network and the student network can be neural networks that can perform feature processing on the features of multimedia data, such as deep neural networks, attention networks (such as Transformer networks) or convolutional neural networks. In addition, the teacher network is used to guide the training of the student network. During the training process of the student network, the parameters of the teacher network are fixed.
[0008] Secondly, based on the difference between the first feature output by the teacher network and the second feature output by the student network, the student network is trained, wherein the difference is used to obtain the classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI.
[0009] The training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
[0010] That is to say, when the input of the teacher network and the student network is the features of multimedia data generated by non-AI, the student network needs to reduce the difference between the features output by itself and the features output by the teacher network during training. When the input of the teacher network and the student network is the features of multimedia data generated by AI, the student network needs to increase the difference between the features output by itself and the features output by the teacher network during training.
[0011] In this solution, the extracted multimedia data is input into the teacher network and the student network respectively, and the student network is trained based on the difference between the features output by the teacher network and the student network, so that the student network can output features similar to the output features of the teacher network when processing non-AI generated multimedia data, and output features that are significantly different from the output features of the teacher network when processing AI generated multimedia data. That is, the student network and the teacher network jointly serve as detectors of multimedia data, and the output of the detector is ultimately used to determine the classification result of the multimedia data. The difference between the output of the student network and the output of the teacher network is the output of the detector. During the training process of the detector, the training focus is targeted on expanding the output difference of the detector when processing real multimedia data and AI generated multimedia data, so that the detector finally trained can effectively deal with various types of AI generated multimedia data and improve the recognition accuracy of AI generated multimedia data.
[0012] In one possible implementation, the classification result obtained based on the difference value is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
[0013] That is, the subsequent classification of multimedia data is based on the difference between the features output by the teacher network and the features output by the student network, and different difference values may result in different classification results. In addition, the greater the difference between the features output by the teacher network and the features output by the student network, the greater the probability that the multimedia data is generated by AI; the smaller the difference between the features output by the teacher network and the features output by the student network, the smaller the probability that the multimedia data is generated by AI.
[0014] In this scheme, since the expression range of non-AI generated multimedia data is usually limited (for example, non-AI generated images are usually in line with physical common sense), and the expression range of AI generated multimedia data is relatively infinite (for example, AI can generate all kinds of fantastic images), by setting the probability that the multimedia data is generated by AI to have a positive correlation with the output difference value between the student network and the teacher network, it can be ensured that the student network can learn to output features similar to the output features of the teacher network when facing various non-AI generated multimedia data, ensuring that the multimedia data can be classified based on the output difference between the student network and the teacher network in the future, thereby ensuring the feasibility of the scheme.
[0015] In a possible implementation, there may be multiple different ways to extract features of multimedia data.
[0016] In the case where the multimedia data is not generated by AI, the multimedia data is input into the feature extraction network to obtain the features of the multimedia data output by the feature extraction network. That is, the multimedia data generated by non-AI is input into the feature extraction network to extract the features, and then input into the teacher network and the student network, which is equivalent to the input of the teacher network and the student network are both the output of the feature extraction network.
[0017] Alternatively, when the multimedia data is generated by AI, the multimedia data is input into the feature extraction network, and the features output by the feature extraction network are input into the feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network. That is, the multimedia data generated by AI is input into the feature extraction network to extract features, then input into the feature enhancement network to perform feature enhancement, and then input into the teacher network and the student network, which is equivalent to the input of the teacher network and the student network being the output of the feature enhancement network.
[0018] In this scheme, for AI-generated multimedia data, by setting a feature enhancement network after the feature extraction network, feature enhancement can be performed on the features of the AI-generated multimedia data, which is equivalent to performing AI generation processing again on the basis of the original features, thereby obtaining features of a wider range of AI-generated multimedia data, ensuring that the student network can recognize as many features of AI-generated multimedia data as possible during the training stage, thereby improving the generalization of the student network.
[0019] In a possible implementation, during the training of the student network, the student network and the feature enhancement network are trained alternately. Moreover, the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network. For example, after a round of iterative training is performed on the student network based on part or all of the training data, another round of iterative training is performed on the feature enhancement network based on part or all of the training data; after a round of iterative training is performed on the feature enhancement network, the training of the student network is continued, and so on, until multiple rounds of training of the student network and the feature enhancement network are completed.
[0020] That is to say, by training the feature enhancement network, the feature enhancement network can generate various features based on the features of the original AI-generated multimedia data, and when these features are used as inputs to the student network and the teacher network, it is difficult for the student network to output features that are sufficiently different from the output features of the teacher network, thereby helping the subsequent training of the student network to output features that are significantly different from the output features of the teacher network when facing various types of AI-generated multimedia data, ensuring that the student network can still have high performance when facing multimedia data generated by various unknown generators in the inference stage.
[0021] In one possible implementation, the teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.
[0022] That is to say, in the training phase of the teacher network, the features output by the teacher network will be further used to perform classification (for example, classification is performed on the features output by the teacher network based on the classifier) and obtain the corresponding classification results. In this way, when training the teacher network, by setting the training goal of the training phase to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results, the teacher network can output features with large differences for non-AI generated multimedia data and AI generated multimedia data, ensuring that there is sufficient difference between the features of the non-AI generated multimedia data output by the teacher network and the features of the AI generated multimedia data, thereby improving the accuracy of the student network obtained by subsequent training.
[0023] In one possible implementation, after the student network is trained, the features of the multimedia data can be input into the trained student network to obtain a third feature; then, the difference value between the first feature and the third feature is input into the classification network to obtain a classification result output by the classification network, and the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; finally, based on the classification result output by the classification network and the actual classification result corresponding to the multimedia data, the classification network is trained, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the actual classification result.
[0024] That is, the input of the classification network is the difference between the features output by the teacher network and the features output by the trained student network, and the output of the classification network is the classification result of the multimedia data, that is, the classification network is used to determine the classification result based on the difference value. After the teacher network and the student network are trained, the classification network can be trained based on the trained teacher network and the student network, so that the classification network can output accurate classification results based on the difference value.
[0025] In a possible implementation manner, the multimedia data is any one of the following data: image, video, text or voice.
[0026] The second aspect of the present application provides a method for classifying multimedia data, including: acquiring multimedia data and extracting features of the multimedia data; inputting the features of the multimedia data into a teacher network and a student network respectively to obtain a first feature output by the teacher network and a second feature output by the student network; determining a difference value between the first feature and the second feature, and inputting the difference value into a classification network to obtain a classification result output by the classification network, wherein the classification result is used to indicate whether the multimedia data is generated by AI.
[0027] In this scheme, by inputting the extracted multimedia data into the teacher network and the student network respectively, and performing classification of the multimedia data based on the difference between the features output by the teacher network and the student network, the classification network no longer performs classification directly based on the features of the multimedia data, but instead performs classification based on feature difference values that are easier to perform classification, thereby improving the recognition accuracy of AI-generated multimedia data.
[0028] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
[0029] In one possible implementation, the teacher network is a pre-trained neural network, the student network is trained under the guidance of the teacher network, and the training goal of the student network in the training phase is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
[0030] In one possible implementation, the training goal of the teacher network during the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.
[0031] In a possible implementation manner, the multimedia data is any one of the following data: image, video, text or voice.
[0032] The third aspect of the present application provides a model training device, including: an extraction module, used to obtain multimedia data and extract features of the multimedia data, the multimedia data being training data in a training set; a processing module, used to input the features of the multimedia data into a teacher network and a student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; a training module, used to train the student network based on a difference value between the first feature and the second feature; wherein the difference value is used to obtain a classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI, and the training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
[0033] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
[0034] In one possible implementation, the extraction module is specifically used to: when the multimedia data is not generated by AI, input the multimedia data into a feature extraction network to obtain features of the multimedia data output by the feature extraction network; or, when the multimedia data is generated by AI, input the multimedia data into a feature extraction network, and input the features output by the feature extraction network into a feature enhancement network to obtain features of the multimedia data output by the feature enhancement network.
[0035] In a possible implementation, the training module is also used to alternately train the student network and the feature enhancement network during the training process of the student network. The training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.
[0036] In one possible implementation, the teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.
[0037] In a possible implementation, the training module is also used to: after the student network is trained, input the features of the multimedia data into the trained student network to obtain a third feature; input the difference value between the first feature and the third feature into the classification network to obtain a classification result output by the classification network, and the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; based on the classification result output by the classification network and the actual classification result corresponding to the multimedia data, train the classification network, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the actual classification result.
[0038] In a possible implementation manner, the multimedia data is any one of the following data: image, video, text or voice.
[0039] In a fourth aspect, the present application provides a multimedia data classification device, including: an acquisition module, used to acquire multimedia data and extract features of the multimedia data; a processing module, used to input the features of the multimedia data into a teacher network and a student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; the processing module is also used to determine the difference between the first feature and the second feature, and input the difference into a classification network to obtain a classification result output by the classification network, and the classification result is used to indicate whether the multimedia data is generated by AI.
[0040] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
[0041] In one possible implementation, the teacher network is a pre-trained neural network, the student network is trained under the guidance of the teacher network, and the training goal of the student network in the training phase is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
[0042] In one possible implementation, the training goal of the teacher network during the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.
[0043] In a possible implementation manner, the multimedia data is any one of the following data: image, video, text or voice.
[0044] In a fifth aspect, the present application provides a model training device, which may include a processor, the processor and a memory are coupled, the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method of the first aspect or any implementation of the first aspect is implemented. For the steps in each possible implementation of the first aspect executed by the processor, please refer to the first aspect for details, which will not be repeated here.
[0045] In a sixth aspect, the present application provides a multimedia data classification device, which may include a processor, the processor and a memory are coupled, the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method of the second aspect or any implementation of the second aspect is implemented. For the steps in each possible implementation of the second aspect executed by the processor, the details can be referred to the second aspect, and will not be repeated here.
[0046] The seventh aspect of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes a method implemented in any one of the first or second aspects.
[0047] An eighth aspect of the present application provides a circuit system, the circuit system includes a processing circuit, and the processing circuit is configured to execute a method implemented in any one of the first or second aspects above.
[0048] The ninth aspect of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute a method implemented in any of the first or second aspects.
[0049] The tenth aspect of the present application provides a chip system, which includes a processor for supporting an electronic device to implement the functions involved in any implementation of the first aspect or the second aspect, for example, processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the electronic device. The chip system can be composed of a chip, or it can include a chip and other discrete devices.
[0050] The beneficial effects of the second to tenth aspects mentioned above can be referred to the introduction of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A schematic diagram of a system architecture 100 provided in an embodiment of the present application;
[0052] Figure 2 A flowchart of a model training method provided in an embodiment of the present application;
[0053] Figure 3 A training diagram of a student network provided in an embodiment of the present application;
[0054] Figure 4 A schematic diagram of determining a training target of a student network provided in an embodiment of the present application;
[0055] Figure 5 A schematic diagram of a student network training when the input is different types of data provided in an embodiment of the present application;
[0056] Figure 6 A schematic diagram of a training process of a student network provided in an embodiment of the present application;
[0057] Figure 7 A schematic diagram of a classification network training provided in an embodiment of the present application;
[0058] Figure 8 A training diagram of a teacher network provided in an embodiment of the present application;
[0059] Fig. 9 A flowchart of a multimedia data classification method provided in an embodiment of the present application;
[0060] Fig.10 A schematic diagram of performing classification on multimedia data in a reasoning process provided in an embodiment of the present application;
[0061] Fig.11 A schematic diagram of the structure of a model training device provided in an embodiment of the present application;
[0062] Fig.12 A schematic diagram of the structure of a multimedia data classification device provided in an embodiment of the present application;
[0063] Fig.13 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0064] Fig.14 A schematic diagram of the structure of a chip provided in an embodiment of the present application;
[0065] Fig.15 A schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION
[0066] In order to make the purpose, technical solutions and advantages of the present application clearer, the embodiments of the present application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of the present application, rather than all embodiments. It is known to those of ordinary skill in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0067] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchanged where appropriate, so that the embodiments can be implemented in a sequence other than that illustrated or described in the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are clearly listed, but may include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. The naming or numbering of the steps that appear in the present application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The process steps that have been named or numbered can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. There may be other division methods when it is implemented in actual applications. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. In addition, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed in multiple circuit units, and some or all of the units may be selected according to actual needs to achieve the purpose of the present application.
[0068] To facilitate understanding, some technical terms involved in the embodiments of the present application are first introduced below.
[0069] (1) Teacher Network
[0070] In this embodiment, the teacher network is a pre-trained neural network used to guide the training process of the student network, thereby realizing the training of the student network.
[0071] (2) Student Network
[0072] In this embodiment, the student network is also a neural network, and the training process of the student network needs to be completed under the guidance of the teacher network, that is, the training of the student network depends on the teacher network.
[0073] (3) Neural Network
[0074] A neural network may be composed of neural units, and a neural unit may refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input, and the output of the operation unit may be:
[0075]
[0076] Where s=1, 2, ...n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolution layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the characteristics of the local receptive field. The local receptive field can be an area composed of several neural units.
[0077] (4) Deep Neural Network (DNN)
[0078] Deep neural network, also known as multi-layer neural network, can be understood as a neural network with many hidden layers. There is no special metric for "many" here. From the position of different layers of DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since DNN has many layers, the coefficient W and the offset vector The number of these parameters is also very large. The definition of these parameters in DNN is as follows: Taking the coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscripts correspond to the output third layer index 2 and the input second layer index 4. In summary, the coefficients from the kth neuron in the L-1th layer to the jth neuron in the Lth layer are defined as It should be noted that the input layer does not have a W parameter. In a deep neural network, more hidden layers allow the network to better describe complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrix of all layers of the trained deep neural network (the weight matrix formed by many layers of vector W).
[0079] (5) Convolutional Neural Network (CNN)
[0080] A convolutional neural network is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor consisting of a convolutional layer and a subsampling layer. The feature extractor can be regarded as a filter, and the convolution process can be regarded as using a trainable filter to convolve with an input feature map. A convolutional layer refers to a neural unit layer in a convolutional neural network that performs convolution processing on the input signal. In a convolutional layer of a convolutional neural network, a neural unit can only be connected to some neural units in adjacent layers. A convolutional layer usually contains several feature planes, and each feature plane can be composed of some neural units arranged in a rectangular shape. The neural units in the same feature plane share weights, and the shared weights here are the convolution kernels.
[0081] The convolution kernel can be initialized in the form of a matrix of random size, and the convolution kernel can obtain reasonable weights through learning during the training process of the convolutional neural network. In addition, the direct benefit of shared weights is to reduce the connections between the layers of the convolutional neural network, while reducing the risk of overfitting.
[0082] (6) Recurrent Neural Network (RNN)
[0083] A recurrent neural network is a type of recursive neural network that takes sequence data as input, performs recursion in the direction of sequence evolution, and all nodes (recurrent units) are connected in a chain.
[0084] Recurrent neural networks have memory, parameter sharing, and Turing completeness, so they have certain advantages in learning the nonlinear features of sequences. Recurrent neural networks are used in natural language processing (NLP), such as speech recognition, language modeling, machine translation, and other fields, and are also used for various time series forecasting.
[0085] (7) Attention Network
[0086] Attention network is a network model that uses attention mechanism to improve the model training speed. At present, the typical attention network includes Transformer network. The model using attention mechanism can assign different weights to each part of the input sequence, so as to extract more important feature information from the input sequence, so that the model can finally obtain more accurate output.
[0087] (8) Loss function
[0088] In the process of training a neural network, because we hope that the output of the neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the target value we really want, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the neural network). For example, if the predicted value of the network is high, adjust the weight vector to make it predict a lower value, and continue to adjust until the neural network can predict the target value we really want or a value very close to the target value we really want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function, which are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the neural network becomes a process of minimizing this loss as much as possible.
[0089] (9) Back propagation algorithm
[0090] The neural network can use the error back propagation (BP) algorithm to correct the size of the parameters in the initial prediction model during the training process, so that the error loss of the prediction model becomes smaller and smaller. Specifically, the forward transmission of the input signal to the output will generate error loss, and the error loss information is back-propagated to update the parameters in the initial prediction model, so that the error loss converges. The back propagation algorithm is a back propagation movement dominated by error loss, aiming to obtain the optimal parameters of the prediction model, such as the weight matrix.
[0091] Specifically, during the model training process, the back propagation algorithm is usually used to calculate the gradient of each node in the model, so as to adjust the weight parameters of the node based on the gradient of each node, thereby reducing the loss function value of the model as much as possible. Among them, the gradient represents the rate of change of a function at a certain point. In addition, the gradient of each node in the model can be determined by taking partial derivatives.
[0092] (10) Gradient descent
[0093] Gradient descent is a first-order optimization algorithm that is often used in machine learning to recursively approximate the minimum deviation prediction model. To use gradient descent to find the local minimum of a function, it is necessary to iteratively search for points at a specified step distance in the opposite direction of the gradient (or approximate gradient) corresponding to the current point on the function. Gradient descent is one of the most commonly used methods for solving the prediction model parameters of machine learning algorithms, that is, unconstrained optimization problems.
[0094] Specifically, when solving the minimum value of the loss function, we can use the gradient descent method to iterate step by step to obtain the minimized loss function and prediction model parameter values. Conversely, if we need to solve the maximum value of the loss function, we need to use the gradient ascent method to iterate.
[0095] The applicant has found that the current method for identifying AI-generated images is usually based on a simple cross-entropy loss function to train a neural network, thereby obtaining an image classification model that can perform binary classification. However, for this conventional image classification model, if an image generated by a certain AI generator has appeared in the training set, the image classification model can obtain relatively accurate classification results when identifying other images generated by the AI generator; if an image generated by a certain AI generator has never appeared in the training set, the image classification model will find it difficult to obtain accurate classification results when identifying images generated by the AI generator. That is, the current image classification model is difficult to obtain good classification performance when processing images generated by unknown AI generators.
[0096] Based on this, the embodiment of the present application provides a model training method, by inputting the extracted multimedia data into the teacher network and the student network respectively, and training the student network based on the difference between the features output by the teacher network and the student network, so that the student network can output features similar to the output features of the teacher network when processing non-AI generated multimedia data, and output features that are significantly different from the output features of the teacher network when processing AI generated multimedia data. That is, the student network and the teacher network jointly serve as detectors of multimedia data, and the difference between the output of the student network and the teacher network is the output of the detector. In the training process of the detector, the training focus is targeted on expanding the output difference of the detector when processing real multimedia data and AI generated multimedia data, so that the detector finally trained can effectively deal with various types of AI generated multimedia data and improve the recognition accuracy of AI generated multimedia data.
[0097] The model training method provided in the embodiment of the present application can be applied to train models for classifying various multimedia data, such as images, videos, texts, or voices, so as to classify multimedia data as AI-generated or non-AI-generated. The multimedia data classification method provided in the embodiment of the present application can be applied to classify various multimedia data.
[0098] See also Figure 1 , Figure 1 A schematic diagram of a system architecture 100 provided in an embodiment of the present application. Figure 1 As shown, in the system architecture 100, the execution device 110 can be implemented by one or more servers. Optionally, the execution device 110 cooperates with other computing devices, such as data storage, routers, load balancers and other devices; the execution device 110 can be arranged on a physical site, or distributed on multiple physical sites. The execution device 110 can use the data in the data storage system 120, or call the program code in the data storage system 120 to implement the model training method and / or multimedia data classification method provided in the embodiment of the present application.
[0099] Users can operate their respective user devices (such as local device 101 and local device 102) to interact with execution device 110. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a laptop computer, and a smart car.
[0100] The local device of each user can interact with the execution device 110 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.
[0101] In one implementation, the execution device 110 is used to implement the model training method and the multimedia data classification method provided in the embodiment of the present application, and obtain a model for implementing multimedia data classification. In addition, when the local device 101 and the local device 102 need to classify multimedia data, the execution device 110 classifies the multimedia data provided by the user based on the trained model, and then returns the corresponding classification results to the local device 101 and the local device 102.
[0102] In another implementation, the execution device 110 is used to implement the model training method provided in the embodiment of the present application, and send the obtained model for implementing multimedia data classification to the local device 101 and the local device 102. In this way, when the local device 101 and the local device 102 need to classify multimedia data, the local device 101 and the local device 102 can classify the multimedia data provided by the user based on the trained model, and then obtain the corresponding classification result.
[0103] In another implementation, one or more aspects of the execution device 110 can be implemented by each local device. For example, the local device 101 can provide local data or feedback calculation results to the execution device 110, or execute the model training method and multimedia data classification method provided in the embodiments of the present application.
[0104] It should be noted that all functions of the execution device 110 may also be implemented by the local device. For example, the local device 101 implements the functions of the execution device 110 and provides services to its own user, or provides services to the user of the local device 102.
[0105] In general, the model training method and / or multimedia data classification method provided in the embodiments of the present application can be applied to electronic devices, such as the above-mentioned execution device 110, local device 101 or local device 102.
[0106] See also Figure 2 , Figure 2 A flow chart of a model training method provided in an embodiment of the present application. Figure 2 As shown, the model training method provided in the embodiment of the present application includes the following steps 201-203.
[0107] Step 201 , acquiring multimedia data and extracting features of the multimedia data, where the multimedia data is training data in a training set.
[0108] In this embodiment, the multimedia data is, for example, image, video, text, or voice data, and this embodiment does not specifically limit this.
[0109] Among them, extracting features of multimedia data can refer to performing feature extraction on multimedia data through a pre-trained feature extraction network, or it can refer to converting multimedia data into a feature matrix that can be used as a neural network input. This embodiment does not make specific limitations on this.
[0110] Step 202, input the features of the multimedia data into the teacher network and the student network respectively, and obtain the first feature output by the teacher network and the second feature output by the student network.
[0111] After obtaining the features of the multimedia data, feature processing is performed on the features of the multimedia data through the teacher network and the student network, respectively, so as to obtain the first feature and the second feature. Wherein, the teacher network and the student network can both be neural networks that can perform feature processing on the features of the multimedia data, such as deep neural networks, attention networks (such as Transformer networks) or convolutional neural networks, and this embodiment does not limit the structure of the teacher network and the student network. In addition, the structure of the teacher network and the structure of the student network can be the same or different.
[0112] Optionally, the teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results. That is to say, in the training phase of the teacher network, the features output by the teacher network will be further used to perform classification (for example, classification is performed on the features output by the teacher network based on the classifier) and obtain corresponding classification results. In this way, when training the teacher network, by setting the training goal of the training phase to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results, the teacher network can output features with large differences for non-AI generated multimedia data and AI generated multimedia data, ensuring that there is sufficient difference between the features of the non-AI generated multimedia data output by the teacher network and the features of the AI generated multimedia data, thereby improving the accuracy of the student network obtained by subsequent training.
[0113] Step 203: training the student network based on the difference between the first feature and the second feature.
[0114] In this embodiment, after obtaining the first feature output by the teacher network and the second feature output by the student network, the difference value between the first feature and the second feature is first obtained. Moreover, the difference value between the first feature and the second feature can be used to obtain the classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI. Specifically, the classification result can be used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value. That is, the subsequent classification of the multimedia data is performed based on the difference value between the features output by the teacher network and the features output by the student network, and different difference values may obtain different classification results. Moreover, the greater the difference value between the features output by the teacher network and the features output by the student network, the greater the probability that the multimedia data is generated by AI; the smaller the difference value between the features output by the teacher network and the features output by the student network, the smaller the probability that the multimedia data is generated by AI.
[0115] Since the difference between the features output by the teacher network and the features output by the student network is used to obtain the classification result, during the training phase of the student network, for different types of input data, it is necessary to train the features that are different between the features output by the student network and the features output by the teacher network. Specifically, the training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
[0116] That is to say, when the input of the teacher network and the student network is the features of multimedia data generated by non-AI, the student network needs to reduce the difference between the features output by itself and the features output by the teacher network during training. That is, when targeting multimedia data generated by non-AI, ensure that the features output by the student network are as close as possible to the features output by the teacher network. When the input of the teacher network and the student network is the features of multimedia data generated by AI, the student network needs to increase the difference between the features output by itself and the features output by the teacher network during training. That is, when targeting multimedia data generated by AI, ensure that the features output by the student network are as different as possible from the features output by the teacher network.
[0117] For example, see Figure 3 , Figure 3 A training diagram of a student network provided in an embodiment of the present application. Figure 3As shown, after the multimedia data is input into the feature extraction network, the features of the multimedia data output by the feature extraction network are obtained. Then, the features of the multimedia data are respectively input into the teacher network and the student network to obtain the first feature output by the teacher network and the second feature output by the student network. Secondly, the difference value is calculated for the first feature and the second feature to obtain the difference value between the first feature and the second feature, and the difference value is used to construct a loss function to train the student network.
[0118] See also Figure 4 , Figure 4 A schematic diagram of determining a training target for a student network provided in an embodiment of the present application. Figure 4 As shown, after acquiring the multimedia data, it can be determined whether the multimedia data is generated by AI. In the case where the multimedia data is generated by AI, the training goal of the student network is determined to reduce the difference between the features output by the student network and the features output by the teacher network; in the case where the multimedia data is not generated by AI, the training goal of the student network is determined to increase the difference between the features output by the student network and the features output by the teacher network.
[0119] In general, by training the student network in the above-mentioned training method, we can focus on widening the gap between the features output by the student network and the teacher network when facing real multimedia data and AI-generated multimedia data, so that the student network and the teacher network can output features as close as possible when facing real multimedia data, and can output features as different as possible when facing AI-generated multimedia data, so that the difference values used to obtain the classification results can have a sufficiently large gap when corresponding to real multimedia data and AI-generated multimedia data, and ensure that the subsequent student network also outputs features different from the teacher network as much as possible when facing AI-generated multimedia data generated by an unknown generator, thereby improving the recognition accuracy of AI-generated multimedia data.
[0120] Optionally, in order to enable the student network to recognize the characteristics of more types of AI-generated data during the training phase, different training processes are designed for non-AI-generated multimedia data and AI-generated multimedia data, respectively.
[0121] Exemplarily, in the above step 201, the features of the multimedia data are extracted, specifically including: when the multimedia data is generated by AI, the multimedia data is input into the feature extraction network to obtain the features of the multimedia data output by the feature extraction network. That is, the multimedia data that is not generated by AI is input into the feature extraction network to extract the features, and then input into the teacher network and the student network, which is equivalent to the input of the teacher network and the student network are both the output of the feature extraction network.
[0122] Alternatively, when the multimedia data is not generated by AI, the multimedia data is input into the feature extraction network, and the features output by the feature extraction network are input into the feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network. That is, the multimedia data generated by AI is input into the feature extraction network to extract features, then input into the feature enhancement network to perform feature enhancement, and then input into the teacher network and the student network, which is equivalent to the input of the teacher network and the student network being the output of the feature enhancement network.
[0123] For example, see Figure 5 , Figure 5 A schematic diagram of a student network training when the input is different types of data is provided in an embodiment of the present application. Figure 5 As shown in (a), when the input data is multimedia data generated by non-AI, the non-AI generated multimedia data is first input into the feature extraction network. After the features of the multimedia data output by the feature extraction network are obtained, the features of the multimedia data are input into the teacher network and the student network respectively, and the difference value between the features output by the teacher network and the student network is obtained to perform training of the student network based on the difference value.
[0124] like Figure 5 As shown in (b), when the input data is multimedia data generated by AI, the multimedia data generated by AI is first input into the feature extraction network. After obtaining the features output by the feature extraction network, the features output by the feature extraction network are input into the feature enhancement network. The feature enhancement network performs feature enhancement on the input features. Finally, the features of the multimedia data output by the feature enhancement network are respectively input into the teacher network and the student network, and the difference value between the features output by the teacher network and the student network is obtained, so as to perform training of the student network based on the difference value.
[0125] In this scheme, for AI-generated multimedia data, by setting a feature enhancement network after the feature extraction network, feature enhancement can be performed on the features of the AI-generated multimedia data, which is equivalent to performing AI generation processing again on the basis of the original features, thereby obtaining features of a wider range of AI-generated multimedia data, ensuring that the student network can recognize as many features of AI-generated multimedia data as possible during the training stage, thereby improving the generalization of the student network.
[0126] Optionally, during the training process of the student network, the student network and the feature enhancement network may be trained alternately. For example, after one round of iterative training of the student network is completed based on part or all of the training data, another round of iterative training of the feature enhancement network is performed based on part or all of the training data; after one round of iterative training of the feature enhancement network is completed, the training of the student network is continued, and so on, until multiple rounds of training of the student network and the feature enhancement network are completed.
[0127] Moreover, during the training process of the feature enhancement network, the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.
[0128] That is to say, by training the feature enhancement network, the feature enhancement network can generate various features based on the features of the original AI-generated multimedia data, and when these features are used as inputs to the student network and the teacher network, it is difficult for the student network to output features that are sufficiently different from the output features of the teacher network, thereby helping the subsequent training of the student network to output features that are significantly different from the output features of the teacher network when facing various types of AI-generated multimedia data, ensuring that the student network can still have high performance when facing multimedia data generated by various unknown generators in the inference stage.
[0129] For example, see Figure 6 , Figure 6 A schematic diagram of a student network training process provided in an embodiment of the present application. Figure 6 As shown in Figure 1, the training process of the student network actually includes three alternating stages. The first stage is to train the student network based on non-AI generated multimedia data; the second stage is to train the student network based on AI generated multimedia data; and the third stage is to train the feature enhancement network based on AI generated multimedia data.
[0130] In the first stage, the teacher network and the student network are connected after the feature extraction network. After the non-AI generated multimedia data is input into the feature extraction network, the features of the multimedia data output by the feature extraction network are input into the teacher network and the student network respectively, and the difference between the features output by the teacher network and the features output by the student network is calculated. Finally, a loss function is constructed based on the obtained difference value and the student network is trained based on the loss function. In addition, the training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network. In the first stage, the student network can be trained based on some or all of the non-AI generated multimedia data in the training set.
[0131] Exemplarily, in the first stage, the loss function for training the student network can be shown as the following formula.
[0132]
[0133] in, represents the loss function used to train the student network when the input is non-AI generated multimedia data; Represents the features of the teacher network output; represents the characteristics of the student network output; N represents the batch data size; b represents the batch of training data during training; N t (x r ) represents the teacher network N t Output characteristics; N s (x r ) represents the student network N s Output features; x r Represents non-AI generated multimedia data.
[0134] In the second stage, the feature enhancement network is connected after the feature extraction network, and the teacher network and the student network are connected after the feature enhancement network. After the multimedia data generated by AI is input into the feature extraction network, the features output by the feature extraction network will continue to be input into the feature enhancement network, and the features of the multimedia data output by the feature enhancement network will be input into the teacher network and the student network respectively. Then, the difference value is calculated for the features output by the teacher network and the features output by the student network, and the student network is trained based on the obtained difference value. Moreover, the training goal of the student network is to increase the difference value between the features output by the student network and the features output by the teacher network. Among them, in the second stage, the student network can be trained based on part or all of the AI-generated multimedia data in the training set.
[0135] Exemplarily, in the second stage, the loss function for training the student network can be shown as the following formula.
[0136]
[0137] in, represents the loss function used to train the student network when the input is AI-generated multimedia data; Represents the features of the teacher network output; Represents the characteristics of the student network output; N t (G(x f )) represents the teacher network N t Output characteristics; N s (G(x f )) represents the student network N s Output features; x f Represents multimedia data generated by AI; [x, 0] + represents max(x, 0); M is a hyperparameter, which can be set to 1 based on experience. In simple terms, the training process of the student network is to minimize the loss function The value of increases the difference between the features output by the student network and the features output by the teacher network.
[0138] In the third stage, the connection relationship and input data of the neural network are the same as those in the second stage. The difference is that after obtaining the difference between the output features of the teacher network and the output features of the student network, the feature enhancement network is trained based on the difference. In addition, the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network. In the second stage, the student network can be trained based on some or all of the AI-generated multimedia data in the training set.
[0139] Exemplarily, in the third stage, the loss function for training the feature enhancement network can be shown as the following formula.
[0140]
[0141] in, represents the loss function used to train the feature enhancement network when the input is non-AI generated multimedia data; Represents the features of the teacher network output; Represents the characteristics of the student network output; N t (G(x f )) represents the teacher network N t Output characteristics; N s (G(x f )) represents the student network N s Output features; x f represents multimedia data generated by AI; N represents the batch data size; b represents the batch of training data during training.
[0142] In the actual training process, the first stage of training may be performed first, then the second stage of training, and then the third stage of training. After the third stage of training is completed, the first stage, the second stage, and the third stage are repeated until each stage has executed a preset number of rounds, or the output of the student network has reached a preset condition.
[0143] It should be noted that the above description is about executing the first stage, the second stage and the third stage in sequence. In fact, the three stages can also be executed in sequence in other orders, for example, the second stage is executed first, then the first stage is executed, and then the third stage is executed. This embodiment does not specifically limit the execution order between the three stages.
[0144] Optionally, since the difference value between the output features of the student network and the output features of the teacher network is used to obtain the classification result, after completing the training of the student network, the classification network with the input as the difference value can be further trained to ensure that the classification network can obtain accurate classification results based on the input difference value.
[0145] Exemplarily, after the student network is trained, the features of the multimedia data used to train the student network can be input into the trained student network to obtain the third feature. Then, the difference between the first feature and the third feature is input into the classification network to obtain the classification result output by the classification network, and the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI. Finally, based on the classification result output by the classification network and the actual classification result corresponding to the multimedia data, the classification network is trained, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the actual classification result.
[0146] That is, the input of the classification network is the difference between the features output by the teacher network and the features output by the trained student network, and the output of the classification network is the classification result of the multimedia data, that is, the classification network is used to determine the classification result based on the difference value. After the teacher network and the student network are trained, the classification network can be trained based on the trained teacher network and the student network, so that the classification network can output accurate classification results based on the difference value.
[0147] Among them, the classification network can be a neural network that can perform binary classification, such as a deep neural network, an attention network (such as a Transformer network) or a convolutional neural network. This embodiment does not limit the structure of the classification network.
[0148] For example, see Figure 7 , Figure 7 A training diagram of a classification network provided in an embodiment of the present application. Figure 7 As shown, after the multimedia data is input into the feature extraction network, the features of the multimedia data output by the feature extraction network can be respectively input into the teacher network and the trained student network to obtain the first feature and the third feature respectively. Then, by obtaining the difference between the first feature and the third feature and inputting the difference into the classification network, the classification result output by the classification network is obtained. Finally, the loss function is constructed based on the classification result output by the classification network and the actual classification result of the multimedia data to realize the training of the classification network.
[0149] The above introduces the training process of the student network and the classification network. The following will introduce the training process of the teacher network. Among them, the training process of the teacher network is before training the student network and the classification network.
[0150] For example, see Figure 8 , Figure 8 A training diagram of a teacher network provided in an embodiment of the present application. Figure 8As shown, the teacher network is connected between the feature extraction network and the classifier. The feature extraction network is used to perform feature extraction on multimedia data and output the extracted features; then, the teacher network processes the features output by the feature extraction network and outputs the processed features. The classifier is a binary classification network, which is used to process the features output by the teacher network and output classification results. Based on the classification results output by the classifier and the actual classification results corresponding to the multimedia data, a loss function (such as a cross entropy loss function) can be constructed to train the teacher network so that the classifier outputs a classification result that is as close to the actual classification result as possible.
[0151] The above describes the model training method provided by the embodiment of the present application, and the model training method based on the above embodiment can realize the training of the teacher network, the student network and the classification network. The following will introduce the classification method of multimedia data provided by the embodiment of the present application.
[0152] For example, see Fig. 9 and Fig.10 , Fig. 9 A flowchart of a multimedia data classification method provided in an embodiment of the present application; Fig.10 A schematic diagram of performing classification on multimedia data in a reasoning process provided in an embodiment of the present application. Fig. 9 As shown, the multimedia data classification method includes the following steps 901-903.
[0153] Step 901: Acquire multimedia data and extract features of the multimedia data.
[0154] In this embodiment, the acquired multimedia data is multimedia data that needs to be identified as AI-generated during the inference process, such as images, videos, texts, or voice data. Among them, extracting features of multimedia data may refer to performing feature extraction on multimedia data through a pre-trained feature extraction network, or may refer to converting multimedia data into a feature matrix that can be used as a neural network input, and this embodiment does not specifically limit this.
[0155] Step 902, input the features of the multimedia data into the teacher network and the student network respectively, and obtain the first feature output by the teacher network and the second feature output by the student network.
[0156] In this embodiment, step 902 is similar to the above-mentioned step 202. Please refer to the above-mentioned step 202 for details, which will not be repeated here. In addition, the teacher network and the student network in this embodiment are both trained neural networks, for example, the teacher network and the student network trained based on the model training method introduced in the above embodiment.
[0157] Step 903, determine the difference value between the first feature and the second feature, and input the difference value into the classification network to obtain a classification result output by the classification network, where the classification result is used to indicate whether the multimedia data is generated by AI.
[0158] In other words, the features output by the teacher network and the features output by the student network will first calculate a difference value, and then the difference value will be processed based on the classification network to obtain the classification result of the multimedia data.
[0159] In this scheme, by inputting the extracted multimedia data into the teacher network and the student network respectively, and performing classification of the multimedia data based on the difference between the features output by the teacher network and the student network, the classification network no longer performs classification directly based on the features of the multimedia data, but instead performs classification based on feature difference values that are easier to perform classification, thereby improving the recognition accuracy of AI-generated multimedia data.
[0160] Optionally, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
[0161] Optionally, the teacher network is a pre-trained neural network, the student network is trained under the guidance of the teacher network, and the training goal of the student network in the training phase is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
[0162] Optionally, a training goal of the teacher network in the training phase is to reduce the difference between a classification result obtained based on features output by the teacher network and a true classification result.
[0163] The method provided in the embodiment of the present application is described in detail above. Next, the device provided in the embodiment of the present application for executing the above method will be introduced.
[0164] See also Fig.11 , Fig.11 This is a schematic diagram of the structure of a model training device provided in an embodiment of the present application. Fig.11As shown, the model training device provided by the embodiment of the present application includes: an extraction module 1101, used to obtain multimedia data and extract features of the multimedia data, where the multimedia data is training data in a training set; a processing module 1102, used to input the features of the multimedia data into a teacher network and a student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; a training module 1103, used to train the student network based on a difference value between the first feature and the second feature; wherein the difference value is used to obtain a classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI, and the training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
[0165] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
[0166] In one possible implementation, the extraction module 1101 is specifically used to: when the multimedia data is not generated by AI, input the multimedia data into a feature extraction network to obtain features of the multimedia data output by the feature extraction network; or, when the multimedia data is generated by AI, input the multimedia data into a feature extraction network, and input the features output by the feature extraction network into a feature enhancement network to obtain features of the multimedia data output by the feature enhancement network.
[0167] In a possible implementation, the training module 1103 is also used to alternately train the student network and the feature enhancement network during the training process of the student network, and the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.
[0168] In one possible implementation, the teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.
[0169] In one possible implementation, the training module 1103 is also used to: after the student network is trained, input the features of the multimedia data into the trained student network to obtain a third feature; input the difference value between the first feature and the third feature into the classification network to obtain a classification result output by the classification network, and the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; based on the classification result output by the classification network and the actual classification result corresponding to the multimedia data, train the classification network, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the actual classification result.
[0170] In a possible implementation manner, the multimedia data is any one of the following data: image, video, text or voice.
[0171] See also Fig.12 , Fig.12 The structure diagram of a multimedia data classification device provided in an embodiment of the present application is shown in FIG. Fig.12 As shown, the multimedia data classification device provided in the embodiment of the present application includes: an acquisition module 1201, used to acquire multimedia data and extract features of the multimedia data; a processing module 1202, used to input the features of the multimedia data into the teacher network and the student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; the processing module 1202 is also used to determine the difference value between the first feature and the second feature, and input the difference value into the classification network to obtain a classification result output by the classification network, and the classification result is used to indicate whether the multimedia data is generated by AI.
[0172] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
[0173] In one possible implementation, the teacher network is a pre-trained neural network, the student network is trained under the guidance of the teacher network, and the training goal of the student network in the training phase is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
[0174] In one possible implementation, the training goal of the teacher network during the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.
[0175] In a possible implementation manner, the multimedia data is any one of the following data: image, video, text or voice.
[0176] See also Fig.13 , Fig.13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1300 can be specifically a mobile phone, a tablet, a laptop, a smart wearable device, a server, etc., which is not limited here. Specifically, the electronic device 1300 includes: a receiving module 1301, a sending module 1302, a processor 1303 and a memory 1304 (wherein the number of processors 1303 in the electronic device 1300 can be one or more, Fig.13 In the example of FIG. 1301 , a processor 1303 is used, where the processor 1303 may include an application processor 13031 and a communication processor 13032. In some embodiments of the present application, the receiving module 1301, the sending module 1302, the processor 1303 and the memory 1304 may be connected via a bus or other means.
[0177] The memory 1304 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1303. A portion of the memory 1304 may also include a non-volatile random access memory (NVRAM). The memory 1304 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0178] The processor 1303 controls the operation of the electronic device. In a specific application, the various components of the electronic device are coupled together through a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are referred to as bus systems in the figure.
[0179] The method disclosed in the above embodiment of the present application can be applied to the processor 1303, or implemented by the processor 1303. The processor 1303 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 1303 or an instruction in the form of software. The above processor 1303 can be a general-purpose processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0180] The processor 1303 can implement or execute the methods, steps and logic diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as a hardware decoding processor for execution, or a combination of hardware and software modules in the decoding processor for execution. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1304, and the processor 1303 reads the information in the memory 1304 and completes the steps of the above method in combination with its hardware.
[0181] The receiving module 1301 can be used to receive input digital or character information, and generate signal input related to the relevant settings and function control of the electronic device. The sending module 1302 can be used to output digital or character information through the first interface; the sending module 1302 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the sending module 1302 can also include a display device such as a display screen.
[0182] The electronic device provided in the embodiment of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit, so that the chip in the electronic device executes the classification method of multimedia data described in the above embodiment, or so that the chip in the training device executes the model training method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device end, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0183] For details, please refer to Fig.14 , Fig.14 A schematic diagram of the structure of a chip provided in an embodiment of the present application, the chip can be expressed as a neural network processor NPU 1400, NPU 1400 is mounted on the host CPU (Host CPU) as a coprocessor, and the host CPU assigns tasks. The core part of the NPU is the operation circuit 1403, which is controlled by the controller 1404 to extract matrix data from the memory and perform multiplication operations.
[0184] In some implementations, the operation circuit 1403 includes multiple processing units (Process Engine, PE) inside. In some implementations, the operation circuit 1403 is a two-dimensional systolic array. The operation circuit 1403 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the operation circuit 1403 is a general-purpose matrix processor.
[0185] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The operation circuit takes the corresponding data of matrix B from the weight memory 1402 and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory 1401 and performs matrix operation with matrix B, and the partial result or final result of the matrix is stored in the accumulator 1408.
[0186] The unified memory 1406 is used to store input data and output data. The weight data is directly transferred to the weight memory 1402 through the direct memory access controller (DMAC) 1405. The input data is also transferred to the unified memory 1406 through the DMAC.
[0187] BIU stands for Bus Interface Unit, i.e., bus interface unit 1410 , which is used for interaction between AXI bus, DMAC and instruction fetch buffer (IFB) 1409 .
[0188] The bus interface unit 1410 (BIU) is used for the instruction fetch memory 1409 to obtain instructions from the external memory, and is also used for the storage unit access controller 1405 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0189] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1406 or to transfer weight data to the weight memory 1402 or to transfer input data to the input memory 1401.
[0190] The vector calculation unit 1407 includes multiple operation processing units, and further processes the output of the operation circuit 1403 when necessary, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, upsampling of feature planes, etc.
[0191] In some implementations, the vector calculation unit 1407 can store the processed output vector to the unified memory 1406. For example, the vector calculation unit 1407 can apply a linear function; or a nonlinear function to the output of the operation circuit 1403, such as linear interpolation of the feature plane extracted by the convolution layer, and then, for example, a vector of accumulated values to generate an activation value. In some implementations, the vector calculation unit 1407 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1403, for example, for use in a subsequent layer in a neural network.
[0192] An instruction fetch buffer 1409 connected to the controller 1404 is used to store instructions used by the controller 1404;
[0193] Unified memory 1406, input memory 1401, weight memory 1402 and instruction fetch memory 1409 are all on-chip memories. External memories are private to the NPU hardware architecture.
[0194] The processor mentioned in any of the above places may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.
[0195] See also Fig.15 , Fig.15 This is a schematic diagram of a computer-readable storage medium provided in an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the above Figure 2 or Fig. 9 The disclosed methods may be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of manufacture.
[0196] Fig.15 Schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.
[0197] In one embodiment, computer readable storage medium 1500 is provided using signal bearing medium 1501. Signal bearing medium 1501 may include one or more program instructions 1502, which when executed by one or more processors may provide the above-mentioned Figure 2 or Fig. 9 Describes the functionality or part of the functionality.
[0198] In some examples, the signal bearing medium 1501 may include a computer readable medium 1503 such as, but not limited to, a hard drive, a compact disk (CD), a digital video disk (DVD), a digital tape, a memory, a ROM or RAM, and the like.
[0199] In some embodiments, the signal bearing medium 1501 may include a computer recordable medium 1504, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, etc. In some embodiments, the signal bearing medium 1501 may include a communication medium 1505, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, etc.). Thus, for example, the signal bearing medium 1501 may be communicated by a wireless form of the communication medium 1505 (e.g., a wireless communication medium that complies with the IEEE 802.X standard or other transmission protocol).
[0200] The one or more program instructions 1502 may be, for example, computer executable instructions or logic implementation instructions. In some examples, the computing device of the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1502 communicated to the computing device via one or more of the computer readable medium 1503, the computer recordable medium 1504, and / or the communication medium 1505.
[0201] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0202] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the method of each embodiment of the present application.
[0203] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0204] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on the computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device or data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that contains one or more available media integration. Available media can be magnetic media, (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state hard disk (SSD)), etc.
Claims
1. A model training method, It is characterized in that include: Acquire multimedia data and extract features of the multimedia data, wherein the multimedia data is training data in a training set; Inputting the features of the multimedia data into the teacher network and the student network respectively, obtaining a first feature output by the teacher network and a second feature output by the student network; Training the student network based on the difference between the first feature and the second feature; Among them, the difference value is used to obtain the classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI. The training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
2. The method according to claim 1, It is characterized in that The classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
3. The method according to claim 1 or 2, It is characterized in that The extracting and obtaining the features of the multimedia data includes: In the case where the multimedia data is not generated by AI, inputting the multimedia data into a feature extraction network to obtain features of the multimedia data output by the feature extraction network; Alternatively, in the case where the multimedia data is generated by AI, the multimedia data is input into the feature extraction network, and the features output by the feature extraction network are input into a feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network.
4. The method according to claim 3, It is characterized in that The method further comprises: During the training process of the student network, the student network and the feature enhancement network are trained alternately, and the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.
5. The method according to any one of claims 1 to 4, It is characterized in that The teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.
6. The method according to any one of claims 1 to 5, It is characterized in that After the student network training is completed, the method further includes: Inputting the feature of the multimedia data into the trained student network to obtain a third feature; Inputting a difference value between the first feature and the third feature into a classification network to obtain a classification result output by the classification network, wherein the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; Based on the classification result output by the classification network and the real classification result corresponding to the multimedia data, the classification network is trained, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the real classification result.
7. The method according to any one of claims 1 to 6, It is characterized in that The multimedia data is any one of the following data: image, video, text or voice.
8. A classification method for multimedia data, It is characterized in that include: Acquire multimedia data and extract features of the multimedia data; Inputting the features of the multimedia data into the teacher network and the student network respectively, obtaining a first feature output by the teacher network and a second feature output by the student network; Determine a difference value between the first feature and the second feature, and input the difference value into a classification network to obtain a classification result output by the classification network, wherein the classification result is used to indicate whether the multimedia data is generated by AI.
9. The method according to claim 8, It is characterized in that The classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
10. The method according to claim 8 or 9, It is characterized in that The teacher network is a pre-trained neural network, the student network is trained under the guidance of the teacher network, and the training goal of the student network in the training stage is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
11. The method according to claim 10, It is characterized in that The training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.
12. A model training device, It is characterized in that include: An extraction module, used to acquire multimedia data and extract features of the multimedia data, wherein the multimedia data is training data in a training set; A processing module, used for inputting the features of the multimedia data into a teacher network and a student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; A training module, configured to train the student network based on a difference value between the first feature and the second feature; Among them, the difference value is used to obtain the classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI. The training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.
13. The device according to claim 12, It is characterized in that The classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
14. The device according to claim 12 or 13, It is characterized in that The extraction module is specifically used for: In the case where the multimedia data is not generated by AI, inputting the multimedia data into a feature extraction network to obtain features of the multimedia data output by the feature extraction network; Alternatively, in the case where the multimedia data is generated by AI, the multimedia data is input into the feature extraction network, and the features output by the feature extraction network are input into a feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network.
15. The device according to claim 14, It is characterized in that The training module is also used to alternately train the student network and the feature enhancement network during the training process of the student network, and the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.
16. The device according to any one of claims 12 to 15, It is characterized in that The teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.
17. The device according to any one of claims 12 to 16, It is characterized in that The training module is also used for: After the student network is trained, the feature of the multimedia data is input into the trained student network to obtain a third feature; Inputting a difference value between the first feature and the third feature into a classification network to obtain a classification result output by the classification network, wherein the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; Based on the classification result output by the classification network and the real classification result corresponding to the multimedia data, the classification network is trained, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the real classification result.
18. A multimedia data classification device, It is characterized in that include: An acquisition module, used to acquire multimedia data and extract features of the multimedia data; A processing module, used for inputting the features of the multimedia data into a teacher network and a student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; The processing module is also used to determine the difference value between the first feature and the second feature, and input the difference value into a classification network to obtain a classification result output by the classification network, and the classification result is used to indicate whether the multimedia data is generated by AI.
19. The device according to claim 18, It is characterized in that The classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.
20. A model training device, It is characterized in that The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 1 to 7.
21. A multimedia data classification device, It is characterized in that The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 8 to 11.
22. A computer storage medium, It is characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 11.
Citation Information
Cited By
Model training method, multimedia data classification method, and related apparatus
EP4797162A1
Model training method, multimedia data classification method, and related apparatus
WO2025107784A1