Model training method, multimedia data classification method, and related apparatus

By using differential training methods of teacher networks and student networks in the field of AI, the problem of low accuracy of AI-generated image recognition in the prior art is solved, and a higher recognition accuracy of AI-generated multimedia data is achieved.

WO2025107784A1PCT designated stage expired Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/114801
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-08-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the classification model used to identify images generated by AI-generated images has a low accuracy in recognition of images generated by unknown generators, making it difficult to achieve accurate recognition.

Method used

By obtaining the characteristics of multimedia data and inputting it into the teacher network and student network, the student network is trained based on the characteristic difference values ​​output by the teacher network and student network, so that it can output features similar to those of the teacher network when processing multimedia data generated by AI, and output features with greater differences when processing multimedia data generated by AI.

Benefits of technology

The recognition accuracy of AI-generated multimedia data is improved, so that the model can more effectively deal with various types of AI-generated multimedia data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024114801_30052025_PF_FP_ABST
    Figure CN2024114801_30052025_PF_FP_ABST
Patent Text Reader

Abstract

A model training method and a multimedia data classification method, which are applied to the technical field of artificial intelligence (AI). In the model training method, by means of respectively inputting extracted multimedia data into a teacher network and a student network and training the student network on the basis of the difference between a feature output from the teacher network and a feature output from the student network, the student network can output a feature that is similar to the output feature of the teacher network when processing multimedia data that is not generated by AI, and outputs a feature that is greatly different from the output feature of the teacher network when processing multimedia data that is generated by AI. That is, during a training process of the student network, the training focus is targeted to expand the output difference between the student network and the teacher network when processing real multimedia data and multimedia data that is generated by AI, such that a final student network obtained by means of training can effectively cope with various types of multimedia data that is generated by AI, thereby improving the recognition accuracy of the multimedia data that is generated by AI.
Need to check novelty before this filing date? Find Prior Art

Description

Model training method, multimedia data classification method and related device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 22, 2023, with application number 202311577190.9 and application name “Model training method, multimedia data classification method and related device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of Artificial Intelligence (AI) technology, and in particular to a model training method, a multimedia data classification method, and related devices. Background Art

[0003] With the development of AI technology, more and more AI models are playing a vital role in our lives. For example, image classification models can help people categorize photos in their smartphones, while voice recognition models enable functions such as human-computer interaction and voice remote control. However, while AI models bring many conveniences to human life, they also have the potential to negatively impact society.

[0004] Specifically, current AI generative models, such as image and text generation models, can generate realistic images or text based on human instructions, allowing these generated images or text to be used to create malicious fake news. For example, an image generation model could be used to generate images of earthquakes or tsunamis in a specific location, which could then be used to spread malicious fake news online, causing social panic.

[0005] In order to reduce the negative impact of AI-generated images on society, classification models are used in related technologies to identify AI-generated images. Classification models are trained by using real images and AI-generated images to form a training set. However, the classification models in related technologies have low accuracy in classifying images generated by unknown generators (i.e., AI-generated images that do not appear in the training set), making it difficult to accurately identify AI-generated images in some scenarios.

[0006] Summary of the Invention

[0007] This application provides a model training method that can improve the recognition accuracy of the trained model for AI-generated multimedia data.

[0008] In a first aspect, the present application provides a model training method for training a model for identifying multimedia data types in the field of AI. The model training method comprises: first, acquiring multimedia data and extracting features of the multimedia data, where the multimedia data is training data in a training set. Extracting features of the multimedia data may refer to performing feature extraction on the multimedia data through a pre-trained feature extraction network, or may refer to converting the multimedia data into a feature matrix that can be used as a neural network input.

[0009] Then, the features of the multimedia data are input into the teacher network and the student network, respectively, to obtain a first feature output by the teacher network and a second feature output by the student network. Both the teacher network and the student network can be neural networks capable of performing feature processing on the features of multimedia data, such as deep neural networks, attention networks (such as Transformer networks), or convolutional neural networks. Furthermore, the teacher network is used to guide the training of the student network. During the training process of the student network, the parameters of the teacher network remain fixed.

[0010] Secondly, the student network is trained based on the difference between the first feature output by the teacher network and the second feature output by the student network. The difference is used to obtain a classification result for the multimedia data, which is used to indicate whether the multimedia data is AI-generated.

[0011] The training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

[0012] That is, when the inputs to the teacher and student networks are features of non-AI-generated multimedia data, the student network needs to reduce the difference between its own output features and those of the teacher network during training. When the inputs to the teacher and student networks are features of AI-generated multimedia data, the student network needs to increase the difference between its own output features and those of the teacher network during training.

[0013] In this solution, the extracted multimedia data is input into the teacher network and the student network respectively, and the student network is trained based on the difference between the features output by the teacher network and the student network. This allows the student network to output features similar to those output by the teacher network when processing non-AI-generated multimedia data, and to output features that differ significantly from those output by the teacher network when processing AI-generated multimedia data. That is, the student network and the teacher network act together as detectors for multimedia data. The output of the detector is ultimately used to determine the classification result of the multimedia data, and the difference between the output of the student network and the teacher network is the output of the detector. During the training process of the detector, the training focus is specifically placed on expanding the difference in the output of the detector when processing real multimedia data and AI-generated multimedia data, so that the detector finally trained can effectively cope with various types of AI-generated multimedia data and improve the recognition accuracy of AI-generated multimedia data.

[0014] In one possible implementation, the classification result obtained based on the difference value is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

[0015] That is, subsequent classification of multimedia data is based on the difference between the features output by the teacher network and the features output by the student network. Different difference values ​​may result in different classification results. Furthermore, the greater the difference between the features output by the teacher network and the features output by the student network, the greater the probability that the multimedia data was AI-generated; the smaller the difference between the features output by the teacher network and the features output by the student network, the lower the probability that the multimedia data was AI-generated.

[0016] In this scheme, since the expression range of non-AI generated multimedia data is usually limited (for example, non-AI generated images are usually consistent with physical common sense), and the expression range of AI generated multimedia data is relatively infinite (for example, AI can generate all kinds of imaginative images), therefore, by setting the probability that the multimedia data is generated by AI to have a positive correlation with the output difference value between the student network and the teacher network, it can be ensured that the student network can learn to output features similar to the output features of the teacher network when facing various non-AI generated multimedia data, ensuring that the multimedia data can be classified based on the output difference between the student network and the teacher network in the future, thereby ensuring the feasibility of the scheme.

[0017] In a possible implementation, there may be multiple different ways to extract features of multimedia data.

[0018] When the multimedia data is not AI-generated, the multimedia data is input into the feature extraction network to obtain the features of the multimedia data output by the feature extraction network. In other words, the non-AI-generated multimedia data is input into the feature extraction network to extract features, and then input into the teacher network and the student network. This is equivalent to the input of the teacher network and the student network being the output of the feature extraction network.

[0019] Alternatively, if the multimedia data is AI-generated, the multimedia data is input into a feature extraction network, and the features output by the feature extraction network are input into a feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network. That is, the AI-generated multimedia data is input into the feature extraction network to extract features, then input into the feature enhancement network to perform feature enhancement, and then input into the teacher network and the student network. This is equivalent to the input of the teacher network and the student network being the output of the feature enhancement network.

[0020] In this scheme, for AI-generated multimedia data, by setting up a feature enhancement network after the feature extraction network, feature enhancement can be performed on the features of the AI-generated multimedia data, which is equivalent to performing AI generation processing again on the basis of the original features, thereby obtaining features of a wider range of AI-generated multimedia data, ensuring that the student network can recognize as many features of AI-generated multimedia data as possible during the training stage, thereby improving the generalization of the student network.

[0021] In one possible implementation, during the student network training process, the student network and the feature enhancement network are trained alternately. Furthermore, the feature enhancement network is trained to reduce the difference between the features output by the student network and those output by the teacher network. For example, after a round of iterative training of the student network based on some or all of the training data, another round of iterative training of the feature enhancement network based on some or all of the training data is performed. After a round of iterative training of the feature enhancement network is completed, the student network is trained again, and so on, until multiple rounds of training of the student network and the feature enhancement network are completed.

[0022] That is to say, by training the feature enhancement network, the feature enhancement network can generate various features based on the features of the original AI-generated multimedia data, and when these features are used as inputs to the student network and the teacher network, it is difficult for the student network to output features that are sufficiently different from the output features of the teacher network, which will help the subsequent training of the student network to output features that are significantly different from the output features of the teacher network when facing various types of AI-generated multimedia data, ensuring that the student network can still have high performance when facing multimedia data generated by various unknown generators in the inference stage.

[0023] In one possible implementation, the teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.

[0024] That is to say, during the training phase of the teacher network, the features output by the teacher network will be further used to perform classification (for example, classification is performed on the features output by the teacher network based on the classifier) ​​and obtain corresponding classification results. In this way, when training the teacher network, by setting the training goal of the training phase to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results, the teacher network can output features with large differences for non-AI generated multimedia data and AI generated multimedia data, ensuring that there is sufficient difference between the features of the non-AI generated multimedia data output by the teacher network and the features of the AI ​​generated multimedia data, thereby improving the accuracy of the student network obtained by subsequent training.

[0025] In one possible implementation, after the student network is trained, the features of the multimedia data can be input into the trained student network to obtain a third feature; then, the difference between the first feature and the third feature is input into the classification network to obtain the classification result output by the classification network. The classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; finally, based on the classification result output by the classification network and the actual classification result corresponding to the multimedia data, the classification network is trained. The training goal of the classification network is to reduce the difference between the classification result output by the classification network and the actual classification result.

[0026] In other words, the input to the classification network is the difference between the features output by the teacher network and the features output by the trained student network. The output of the classification network is the classification result of the multimedia data. In other words, the classification network is used to determine the classification result based on the difference value. After the teacher network and the student network are both trained, the classification network can be trained based on the trained teacher network and the student network, so that the classification network can output accurate classification results based on the difference value.

[0027] In a possible implementation, the multimedia data is any one of the following data: image, video, text, or voice.

[0028] The second aspect of the present application provides a method for classifying multimedia data, including: obtaining multimedia data and extracting features of the multimedia data; inputting the features of the multimedia data into a teacher network and a student network respectively to obtain a first feature output by the teacher network and a second feature output by the student network; determining a difference value between the first feature and the second feature, and inputting the difference value into a classification network to obtain a classification result output by the classification network, wherein the classification result is used to indicate whether the multimedia data is generated by AI.

[0029] In this solution, by inputting the extracted multimedia data into the teacher network and the student network respectively, and performing classification of the multimedia data based on the difference between the features output by the teacher network and the student network, the classification network no longer performs classification directly based on the features of the multimedia data, but instead performs classification based on feature difference values ​​that are easier to perform classification, thereby improving the recognition accuracy of AI-generated multimedia data.

[0030] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

[0031] In one possible implementation, the teacher network is a pre-trained neural network, and the student network is trained under the guidance of the teacher network. The training goal of the student network in the training phase is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

[0032] In one possible implementation, the training goal of the teacher network during the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.

[0033] In a possible implementation, the multimedia data is any one of the following data: image, video, text, or voice.

[0034] The third aspect of the present application provides a model training device, including: an extraction module, used to obtain multimedia data and extract features of the multimedia data, the multimedia data being training data in a training set; a processing module, used to input the features of the multimedia data into a teacher network and a student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; a training module, used to train the student network based on the difference value between the first feature and the second feature; wherein the difference value is used to obtain a classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI, and the training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

[0035] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

[0036] In one possible implementation, the extraction module is specifically used to: when the multimedia data is not generated by AI, input the multimedia data into a feature extraction network to obtain the features of the multimedia data output by the feature extraction network; or, when the multimedia data is generated by AI, input the multimedia data into a feature extraction network, and input the features output by the feature extraction network into a feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network.

[0037] In one possible implementation, the training module is further used to alternately train the student network and the feature enhancement network during the training process of the student network. The training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.

[0038] In one possible implementation, the teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.

[0039] In one possible implementation, the training module is further used to: after the student network is trained, input the features of the multimedia data into the trained student network to obtain a third feature; input the difference value between the first feature and the third feature into the classification network to obtain a classification result output by the classification network, and the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; based on the classification result output by the classification network and the actual classification result corresponding to the multimedia data, train the classification network, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the actual classification result.

[0040] In a possible implementation, the multimedia data is any one of the following data: image, video, text, or voice.

[0041] In a fourth aspect, the present application provides a multimedia data classification device, comprising: an acquisition module for acquiring multimedia data and extracting features of the multimedia data; a processing module for inputting the features of the multimedia data into a teacher network and a student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; the processing module is also used to determine the difference between the first feature and the second feature, and input the difference into a classification network to obtain a classification result output by the classification network, wherein the classification result is used to indicate whether the multimedia data is generated by AI.

[0042] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

[0043] In one possible implementation, the teacher network is a pre-trained neural network, and the student network is trained under the guidance of the teacher network. The training goal of the student network in the training phase is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

[0044] In one possible implementation, the training goal of the teacher network during the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.

[0045] In a possible implementation, the multimedia data is any one of the following data: image, video, text, or voice.

[0046] In a fifth aspect, the present application provides a model training device, which may include a processor coupled to a memory, wherein the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method of the first aspect or any implementation of the first aspect is implemented. For the steps in each possible implementation of the first aspect executed by the processor, please refer to the first aspect for details, and no further description is given here.

[0047] In a sixth aspect, the present application provides a multimedia data classification device, which may include a processor coupled to a memory, wherein the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method of the second aspect or any implementation of the second aspect is implemented. For details of the steps in each possible implementation of the second aspect executed by the processor, please refer to the second aspect and will not be repeated here.

[0048] In a seventh aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes a method implemented in any one of the first and second aspects.

[0049] In an eighth aspect, the present application provides a circuit system, which includes a processing circuit, and the processing circuit is configured to execute a method implemented in any one of the first or second aspects.

[0050] In a ninth aspect, the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method implemented in any one of the first or second aspects.

[0051] In a tenth aspect, the present application provides a chip system, which includes a processor for supporting an electronic device to implement the functions involved in any implementation of the first or second aspect above, for example, processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory for storing program instructions and data necessary for the electronic device. The chip system can be composed of a chip or can include a chip and other discrete devices.

[0052] The beneficial effects of the second to tenth aspects mentioned above can be referred to the introduction of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] FIG1 is a schematic diagram of a system architecture 100 provided in an embodiment of the present application;

[0054] FIG2 is a flow chart of a model training method provided in an embodiment of the present application;

[0055] FIG3 is a schematic diagram of a student network training method according to an embodiment of the present application;

[0056] FIG4 is a schematic diagram of determining a training target for a student network according to an embodiment of the present application;

[0057] FIG5 is a schematic diagram of a training process of a student network when input is different types of data, provided by an embodiment of the present application;

[0058] FIG6 is a schematic diagram of a training process of a student network provided in an embodiment of the present application;

[0059] FIG7 is a schematic diagram of a training method of a classification network provided in an embodiment of the present application;

[0060] FIG8 is a diagram illustrating a training method of a teacher network according to an embodiment of the present application;

[0061] FIG9 is a flow chart of a method for classifying multimedia data provided in an embodiment of the present application;

[0062] FIG10 is a schematic diagram of performing classification on multimedia data during an inference process provided by an embodiment of the present application;

[0063] FIG11 is a schematic structural diagram of a model training device provided in an embodiment of the present application;

[0064] FIG12 is a schematic structural diagram of a multimedia data classification device provided in an embodiment of the present application;

[0065] FIG13 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0066] FIG14 is a schematic structural diagram of a chip provided in an embodiment of the present application;

[0067] FIG15 is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. DETAILED DESCRIPTION

[0068] In order to make the purpose, technical solutions and advantages of this application more clear, the embodiments of this application are described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only embodiments of a part of this application, rather than all embodiments. It is known to those skilled in the art that with the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0069] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the descriptions used in this way can be interchangeable where appropriate so that the embodiments can be implemented in a sequence other than that illustrated or described in this application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The named or numbered process steps can change the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved. The division of units in this application is a logical division. In actual application, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, and the indirect coupling or communication connection between units can be electrical or other similar forms, which are not limited in this application. Moreover, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed into multiple circuit units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this application.

[0070] To facilitate understanding, some technical terms involved in the embodiments of this application are first introduced below.

[0071] (1) Teacher Network

[0072] In this embodiment, the teacher network is a pre-trained neural network used to guide the training process of the student network, thereby achieving the training of the student network.

[0073] (2) Student Network

[0074] In this embodiment, the student network is also a neural network, and the training process of the student network needs to be completed under the guidance of the teacher network, that is, the training of the student network depends on the teacher network.

[0075] (3) Neural Network

[0076] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:

[0077] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0078] (4) Deep Neural Network (DNN)

[0079] Deep neural networks, also known as multi-layer neural networks, can be understood as neural networks with many hidden layers. There is no special metric for "many" here. Based on the position of different layers in DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take the coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscript corresponds to the output of the third layer index 2 and the input of the second layer index 4. In summary, the coefficient from the kth neuron in the L-1th layer to the jth neuron in the Lth layer is defined as It's important to note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters has higher complexity and greater "capacity," meaning it can handle more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrices for all layers of a trained deep neural network (a weight matrix formed by the vectors W across many layers).

[0080] (5) Convolutional Neural Network (CNN)

[0081] A convolutional neural network is a deep neural network with a convolutional structure. It consists of a feature extractor consisting of a convolutional layer and a subsampling layer. This feature extractor can be thought of as a filter, and the convolution process can be thought of as convolving a trainable filter with an input feature map. A convolutional layer refers to the layer of neural units in a convolutional neural network that performs convolution processing on the input signal. In a convolutional layer of a convolutional neural network, a neural unit can only be connected to some of the neural units in adjacent layers. A convolutional layer typically contains several feature planes, each of which can be composed of a number of neural units arranged in a rectangular pattern. Neural units in the same feature plane share weights, which are referred to as convolution kernels.

[0082] Convolution kernels can be initialized as matrices of random size, and during the training process of the convolutional neural network, the convolution kernels can be learned to obtain reasonable weights. In addition, the direct benefit of shared weights is that they reduce the number of connections between the layers of the convolutional neural network, while also reducing the risk of overfitting.

[0083] (6) Recurrent Neural Network (RNN)

[0084] A recurrent neural network is a type of recursive neural network that takes sequence data as input, performs recursion in the direction of sequence evolution, and all nodes (recurrent units) are connected in a chain-like manner.

[0085] Recurrent neural networks (RNNs) possess memory, parameter sharing, and Turing completeness, giving them advantages in learning nonlinear features of sequences. They are used in natural language processing (NLP) fields such as speech recognition, language modeling, and machine translation, and are also used for various time series forecasting applications.

[0086] (7) Attention Network

[0087] Attention networks are network models that utilize the attention mechanism to accelerate model training. Currently, typical attention networks include Transformer networks. Models using the attention mechanism assign different weights to each part of the input sequence, thereby extracting more important features from the input sequence and ultimately achieving more accurate output.

[0088] (8) Loss function

[0089] During neural network training, because we want the output of the neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vectors of each layer of the neural network based on the difference between the two. (Of course, before the first update, there is usually an initialization process, which pre-configures the parameters of each layer in the neural network.) For example, if the network's prediction is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so neural network training becomes a process of minimizing this loss.

[0090] (9) Backpropagation algorithm

[0091] Neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial prediction model during training, reducing the error loss of the prediction model. Specifically, the forward propagation of the input signal to the output generates error loss. This error loss information is then backpropagated to update the parameters in the initial prediction model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by error loss, aiming to obtain the optimal prediction model parameters, such as the weight matrix.

[0092] Specifically, during model training, the backpropagation algorithm is typically used to calculate the gradient of each node in the model. This allows the node weight parameters to be adjusted based on the gradient of each node, thereby minimizing the model's loss function. The gradient represents the rate of change of a function at a specific point. Furthermore, the gradient of each node in the model can be determined by taking partial derivatives.

[0093] (10) Gradient descent

[0094] Gradient descent is a first-order optimization algorithm commonly used in machine learning to recursively approximate minimum-deviation prediction models. To find a local minimum of a function using gradient descent, an iterative search must be performed toward a point on the function at a specified step distance in the opposite direction of the gradient (or approximate gradient) corresponding to the current point. Gradient descent is one of the most commonly used methods for solving unconstrained optimization problems, such as predictive model parameters in machine learning algorithms.

[0095] Specifically, when solving for the minimum value of the loss function, we can use the gradient descent method to iterate step by step to obtain the minimized loss function and prediction model parameter values. Conversely, if we need to solve for the maximum value of the loss function, we need to use the gradient ascent method to iterate.

[0096] The applicant's research has found that the current method for identifying AI-generated images is usually based on a simple cross-entropy loss function to train a neural network, thereby obtaining an image classification model capable of performing binary classification. However, for this conventional image classification model, if the image generated by a certain AI generator has already appeared in the training set, the image classification model can obtain relatively accurate classification results when identifying other images generated by the AI ​​generator; if the image generated by a certain AI generator has never appeared in the training set, the image classification model will find it difficult to obtain accurate classification results when identifying the image generated by the AI ​​generator. In other words, the current image classification model has difficulty in obtaining good classification performance when processing images generated by unknown AI generators.

[0097] Based on this, an embodiment of the present application provides a model training method, which inputs the extracted multimedia data into the teacher network and the student network respectively, and trains the student network based on the difference between the features output by the teacher network and the student network, so that the student network can output features similar to the output features of the teacher network when processing non-AI generated multimedia data, and output features that are significantly different from the output features of the teacher network when processing AI generated multimedia data. That is, the student network and the teacher network jointly serve as detectors of multimedia data, and the difference between the outputs of the student network and the teacher network is the output of the detector. During the training process of the detector, the training focus is targeted on expanding the output difference of the detector when processing real multimedia data and AI generated multimedia data, so that the detector finally trained can effectively deal with various types of AI generated multimedia data and improve the recognition accuracy of AI generated multimedia data.

[0098] The model training method provided in the embodiments of the present application can be applied to train models for classifying various multimedia data, such as images, videos, text, or voice, thereby classifying multimedia data as AI-generated or non-AI-generated. The multimedia data classification method provided in the embodiments of the present application can be applied to classify various multimedia data.

[0099] Please refer to Figure 1, which is a schematic diagram of a system architecture 100 provided in an embodiment of the present application. As shown in Figure 1, in the system architecture 100, the execution device 110 can be implemented by one or more servers. Optionally, the execution device 110 cooperates with other computing devices, such as data storage, routers, load balancers and other devices; the execution device 110 can be arranged at one physical site, or distributed across multiple physical sites. The execution device 110 can use the data in the data storage system 120, or call the program code in the data storage system 120 to implement the model training method and / or multimedia data classification method provided in the embodiment of the present application.

[0100] Users can operate their respective user devices (such as local device 101 and local device 102) to interact with execution device 110. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a laptop computer, and a smart car.

[0101] Each user's local device can interact with the execution device 110 through a communication network of any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, etc., or any combination thereof.

[0102] In one implementation, execution device 110 is configured to implement the model training method and multimedia data classification method provided in the embodiments of the present application, thereby obtaining a model for implementing multimedia data classification. Furthermore, when local device 101 or local device 102 needs to classify multimedia data, execution device 110 classifies the multimedia data provided by the user based on the trained model and returns the corresponding classification results to local device 101 or local device 102.

[0103] In another implementation, the execution device 110 is used to implement the model training method provided in the embodiment of the present application, and send the obtained model for implementing multimedia data classification to the local device 101 and the local device 102. In this way, when the local device 101 and the local device 102 need to classify multimedia data, the local device 101 and the local device 102 can classify the multimedia data provided by the user based on the trained model, and then obtain corresponding classification results.

[0104] In another implementation, one or more aspects of the execution device 110 can be implemented by each local device. For example, the local device 101 can provide local data or feedback calculation results to the execution device 110, or execute the model training method and multimedia data classification method provided in the embodiments of the present application.

[0105] It should be noted that all functions of the execution device 110 may also be implemented by a local device. For example, the local device 101 implements the functions of the execution device 110 and provides services to its own user, or provides services to the user of the local device 102.

[0106] In general, the model training method and / or multimedia data classification method provided in the embodiments of the present application can be applied to electronic devices, such as the above-mentioned execution device 110, local device 101 or local device 102.

[0107] Please refer to Figure 2, which is a flow chart of a model training method provided in an embodiment of the present application. As shown in Figure 2, the model training method provided in an embodiment of the present application includes the following steps 201-203.

[0108] Step 201 : Acquire multimedia data and extract features of the multimedia data, where the multimedia data is training data in a training set.

[0109] In this embodiment, the multimedia data is, for example, image, video, text, or voice data, and this embodiment does not specifically limit this.

[0110] Among them, extracting features of multimedia data can refer to performing feature extraction on multimedia data through a pre-trained feature extraction network, or it can refer to converting multimedia data into a feature matrix that can be used as a neural network input. This embodiment does not make specific limitations on this.

[0111] In step 202, the features of the multimedia data are input into the teacher network and the student network respectively to obtain a first feature output by the teacher network and a second feature output by the student network.

[0112] After obtaining the features of the multimedia data, feature processing is performed on the features of the multimedia data through the teacher network and the student network, respectively, to obtain the first feature and the second feature. Wherein, the teacher network and the student network can both be neural networks that can perform feature processing on the features of the multimedia data, such as deep neural networks, attention networks (such as Transformer networks) or convolutional neural networks. This embodiment does not limit the structure of the teacher network and the student network. In addition, the structure of the teacher network and the structure of the student network can be the same or different.

[0113] Optionally, the teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the true classification results. That is to say, in the training phase of the teacher network, the features output by the teacher network will be further used to perform classification (for example, classification is performed on the features output by the teacher network based on the classifier) ​​and obtain corresponding classification results. In this way, when training the teacher network, by setting the training goal of the training phase to reduce the difference between the classification results obtained based on the features output by the teacher network and the true classification results, the teacher network can output features with large differences for non-AI generated multimedia data and AI generated multimedia data, ensuring that there is sufficient difference between the features of the non-AI generated multimedia data output by the teacher network and the features of the AI ​​generated multimedia data, thereby improving the accuracy of the student network obtained by subsequent training.

[0114] Step 203: training the student network based on the difference between the first feature and the second feature.

[0115] In this embodiment, after obtaining the first feature output by the teacher network and the second feature output by the student network, the difference value between the first feature and the second feature is first obtained. Moreover, the difference value between the first feature and the second feature can be used to obtain the classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI. Specifically, the classification result can be used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value. That is, the subsequent classification of the multimedia data is performed based on the difference value between the feature output by the teacher network and the feature output by the student network, and different difference values ​​may result in different classification results. Moreover, the greater the difference value between the feature output by the teacher network and the feature output by the student network, the greater the probability that the multimedia data is generated by AI; the smaller the difference value between the feature output by the teacher network and the feature output by the student network, the smaller the probability that the multimedia data is generated by AI.

[0116] Since the difference between the features output by the teacher network and the features output by the student network is used to obtain the classification results, during the student network training phase, for different types of input data, it is necessary to train the student network to output features that differ from the features output by the teacher network. Specifically, the training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

[0117] That is to say, when the input of the teacher network and the student network is the features of non-AI generated multimedia data, the student network needs to reduce the difference between the features output by itself and the features output by the teacher network during training. That is, when targeting non-AI generated multimedia data, ensure that the features output by the student network are as close as possible to the features output by the teacher network. When the input of the teacher network and the student network is the features of AI generated multimedia data, the student network needs to increase the difference between the features output by itself and the features output by the teacher network during training. That is, when targeting AI generated multimedia data, ensure that the features output by the student network are as different as possible from the features output by the teacher network.

[0118] For example, please refer to Figure 3, which is a training diagram of a student network provided in an embodiment of the present application. As shown in Figure 3, after the multimedia data is input into the feature extraction network, the features of the multimedia data output by the feature extraction network are obtained. Then, the features of the multimedia data are input into the teacher network and the student network respectively to obtain the first feature output by the teacher network and the second feature output by the student network. Secondly, the difference value of the first feature and the second feature is calculated to obtain the difference value between the first feature and the second feature, and the difference value is used to construct a loss function to train the student network.

[0119] Please refer to Figure 4, which is a schematic diagram of determining the training target of the student network provided in an embodiment of the present application. As shown in Figure 4, after obtaining the multimedia data, it can be determined whether the multimedia data is generated by AI. In the case where the multimedia data is generated by AI, the training target of the student network is determined to be to reduce the difference between the features output by the student network and the features output by the teacher network; in the case where the multimedia data is not generated by AI, the training target of the student network is determined to be to increase the difference between the features output by the student network and the features output by the teacher network.

[0120] In general, by training the student network using the above-mentioned training method, we can focus on widening the gap between the features output by the student network and the teacher network when facing real multimedia data and AI-generated multimedia data. Ultimately, the student network and the teacher network can output features that are as close as possible when facing real multimedia data, and can output features that are as different as possible when facing AI-generated multimedia data. This ensures that the difference values ​​used to obtain the classification results have a sufficiently large gap when corresponding to real multimedia data and AI-generated multimedia data, ensuring that the subsequent student network also outputs features that are as different as possible from the teacher network when facing AI-generated multimedia data generated by an unknown generator, thereby improving the recognition accuracy of AI-generated multimedia data.

[0121] Optionally, in order to enable the student network to recognize the characteristics of more types of AI-generated data during the training phase, different training processes are designed for non-AI-generated multimedia data and AI-generated multimedia data respectively in this embodiment.

[0122] Exemplarily, in step 201, extracting features from multimedia data specifically includes: if the multimedia data is AI-generated, inputting the multimedia data into a feature extraction network to obtain features of the multimedia data output by the feature extraction network. Specifically, if the multimedia data is not AI-generated, after features are extracted from the feature extraction network, the data is then input into the teacher network and the student network. This means that the inputs to both the teacher network and the student network are the outputs of the feature extraction network.

[0123] Alternatively, if the multimedia data is not AI-generated, the multimedia data is input into a feature extraction network, and the features output by the feature extraction network are input into a feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network. In other words, the AI-generated multimedia data is input into the feature extraction network to extract features, then input into the feature enhancement network for feature enhancement, and then input into the teacher network and the student network. This is equivalent to the input of the teacher network and the student network being the output of the feature enhancement network.

[0124] For example, please refer to Figure 5, which is a schematic diagram of a student network training method provided by an embodiment of the present application when the input is different types of data. As shown in (a) of Figure 5, when the input data is non-AI generated multimedia data, the non-AI generated multimedia data is first input into the feature extraction network. After obtaining the features of the multimedia data output by the feature extraction network, the features of the multimedia data are input into the teacher network and the student network respectively, and the difference between the features output by the teacher network and the student network is obtained to perform student network training based on the difference.

[0125] As shown in (b) in Figure 5, when the input data is multimedia data generated by AI, the multimedia data generated by AI is first input into the feature extraction network. After obtaining the features output by the feature extraction network, the features output by the feature extraction network are input into the feature enhancement network. The feature enhancement network performs feature enhancement on the input features. Finally, the features of the multimedia data output by the feature enhancement network are input into the teacher network and the student network respectively, and the difference value between the features output by the teacher network and the student network is obtained to perform training of the student network based on the difference value.

[0126] In this scheme, for AI-generated multimedia data, by setting up a feature enhancement network after the feature extraction network, feature enhancement can be performed on the features of the AI-generated multimedia data, which is equivalent to performing AI generation processing again on the basis of the original features, thereby obtaining features of a wider range of AI-generated multimedia data, ensuring that the student network can recognize as many features of AI-generated multimedia data as possible during the training stage, thereby improving the generalization of the student network.

[0127] Optionally, during the training of the student network, the student network and the feature enhancement network may be trained alternately. For example, after a round of iterative training of the student network based on part or all of the training data, another round of iterative training of the feature enhancement network based on part or all of the training data is performed; after a round of iterative training of the feature enhancement network is completed, the student network training is continued, and so on, until multiple rounds of training of the student network and the feature enhancement network are completed.

[0128] Moreover, during the training process of the feature enhancement network, the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.

[0129] That is to say, by training the feature enhancement network, the feature enhancement network can generate various features based on the features of the original AI-generated multimedia data, and when these features are used as inputs to the student network and the teacher network, it is difficult for the student network to output features that are sufficiently different from the output features of the teacher network, which will help the subsequent training of the student network to output features that are significantly different from the output features of the teacher network when facing various types of AI-generated multimedia data, ensuring that the student network can still have high performance when facing multimedia data generated by various unknown generators in the inference stage.

[0130] For example, please refer to Figure 6, which is a schematic diagram of a student network training process provided in an embodiment of the present application. As shown in Figure 6, the student network training process actually includes three alternating phases. The first phase is to train the student network based on non-AI-generated multimedia data; the second phase is to train the student network based on AI-generated multimedia data; and the third phase is to train the feature enhancement network based on AI-generated multimedia data.

[0131] In the first stage, the teacher network and the student network are connected after the feature extraction network. After the non-AI generated multimedia data is input into the feature extraction network, the features of the multimedia data output by the feature extraction network are input into the teacher network and the student network respectively, and the difference between the features output by the teacher network and the features output by the student network is calculated. Finally, a loss function is constructed based on the obtained difference value and the student network is trained based on the loss function. In addition, the training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network. In the first stage, the student network can be trained based on some or all of the non-AI generated multimedia data in the training set.

[0132] For example, in the first stage, the loss function for training the student network can be shown as the following formula.

[0133] in, represents the loss function used to train the student network when the input is non-AI generated multimedia data; Represents the features of the teacher network output; represents the features of the student network output; N represents the batch data size; b represents the batch of training data during training; N t (x r ) represents the teacher network N t Output features; N s (x r ) represents the student network N s Output features; x r Represents non-AI-generated multimedia data.

[0134] In the second stage, the feature enhancement network is connected after the feature extraction network, and the teacher network and the student network are connected after the feature enhancement network. After the multimedia data generated by AI is input into the feature extraction network, the features output by the feature extraction network will continue to be input into the feature enhancement network, and the features of the multimedia data output by the feature enhancement network will be input into the teacher network and the student network respectively. Then, the difference value of the features output by the teacher network and the features output by the student network is calculated, and the student network is trained based on the obtained difference value. In addition, the training goal of the student network is to increase the difference value between the features output by the student network and the features output by the teacher network. In the second stage, the student network can be trained based on part or all of the AI-generated multimedia data in the training set.

[0135] For example, in the second stage, the loss function for training the student network can be shown as the following formula.

[0136] in, represents the loss function used to train the student network when the input is AI-generated multimedia data; Represents the features of the teacher network output; represents the characteristics of the student network output; N t (G(x f )) represents the teacher network N t Output features; N s (G(x f )) represents the student network N s Output features; x f Represents AI-generated multimedia data; [x, 0] + represents max(x, 0); M is a hyperparameter that can be set to 1 based on experience. In simple terms, the training process of the student network is to minimize the loss function The value of , that is, to increase the difference between the features output by the student network and the features output by the teacher network.

[0137] In the third stage, the connection relationship and input data of the neural network are the same as those in the second stage. The difference is that after obtaining the difference between the output features of the teacher network and the output features of the student network, the feature enhancement network is trained based on the difference value. In addition, the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network. In the second stage, the student network can be trained based on some or all of the AI-generated multimedia data in the training set.

[0138] Exemplarily, in the third stage, the loss function for training the feature enhancement network can be shown as the following formula.

[0139] in, Represents the loss function used to train the feature enhancement network when the input is non-AI generated multimedia data; Represents the features of the teacher network output; represents the characteristics of the student network output; N t (G(x f )) represents the teacher network N t Output features; N s (G(x f )) represents the student network N s Output features; x f represents multimedia data generated by AI; N represents the batch data size; b represents the batch of training data during training.

[0140] In the actual training process, the first stage of training can be performed first, then the second stage of training, and then the third stage of training. After the third stage of training is completed, the first stage, the second stage, and the third stage are repeated until each stage has executed the preset rounds, or the output of the student network has reached the preset conditions.

[0141] It should be noted that the above description shows the execution of the first, second, and third stages in sequence. In practice, the three stages may be executed in another order, for example, the second stage may be executed first, then the first stage, and finally the third stage. This embodiment does not specifically limit the execution order of the three stages.

[0142] Optionally, since the difference value between the output features of the student network and the output features of the teacher network is used to obtain the classification result, after completing the training of the student network, the classification network with the input as the difference value can be further trained to ensure that the classification network can obtain accurate classification results based on the input difference value.

[0143] For example, after the student network is trained, the features of the multimedia data used to train the student network can be input into the trained student network to obtain a third feature. Then, the difference between the first and third features is input into the classification network to obtain a classification result output by the classification network. The classification result output by the classification network is used to indicate whether the multimedia data is AI-generated. Finally, based on the classification result output by the classification network and the actual classification result corresponding to the multimedia data, the classification network is trained. The training goal of the classification network is to reduce the difference between the classification result output by the classification network and the actual classification result.

[0144] In other words, the input to the classification network is the difference between the features output by the teacher network and the features output by the trained student network. The output of the classification network is the classification result of the multimedia data. In other words, the classification network is used to determine the classification result based on the difference value. After the teacher network and the student network are both trained, the classification network can be trained based on the trained teacher network and the student network, so that the classification network can output accurate classification results based on the difference value.

[0145] Among them, the classification network can be a neural network that can perform binary classification, such as a deep neural network, an attention network (such as a Transformer network) or a convolutional neural network. This embodiment does not limit the structure of the classification network.

[0146] For example, please refer to Figure 7, which is a training diagram of a classification network provided in an embodiment of the present application. As shown in Figure 7, after the multimedia data is input into the feature extraction network, the features of the multimedia data output by the feature extraction network can be input into the teacher network and the trained student network respectively to obtain the first feature and the third feature, respectively. Then, by calculating the difference between the first feature and the third feature and inputting the difference into the classification network, the classification result output by the classification network is obtained. Finally, a loss function is constructed based on the classification result output by the classification network and the actual classification result of the multimedia data to realize the training of the classification network.

[0147] The above describes the training process of the student network and the classification network. The following describes the training process of the teacher network. The teacher network training process is before the training of the student network and the classification network.

[0148] For example, please refer to Figure 8, which is a training diagram of a teacher network provided in an embodiment of the present application. As shown in Figure 8, the teacher network is connected between the feature extraction network and the classifier. Among them, the feature extraction network is used to perform feature extraction on the multimedia data and output the extracted features; then, the teacher network processes the features output by the feature extraction network and outputs the processed features. The classifier is a binary classification network, which is used to process the features output by the teacher network, thereby outputting a classification result. Based on the classification result output by the classifier and the actual classification result corresponding to the multimedia data, a loss function (such as a cross entropy loss function) can be constructed to train the teacher network so that the classifier outputs a classification result that is as close to the actual classification result as possible.

[0149] The above describes the model training method provided by the embodiment of the present application, and the model training method based on the above embodiment can realize the training of the teacher network, the student network and the classification network. The following will introduce the classification method of multimedia data provided by the embodiment of the present application.

[0150] For example, please refer to Figures 9 and 10. Figure 9 is a flowchart illustrating a multimedia data classification method provided in an embodiment of the present application; Figure 10 is a schematic diagram illustrating the classification of multimedia data during an inference process provided in an embodiment of the present application. As shown in Figure 9, the multimedia data classification method includes the following steps 901-903.

[0151] Step 901: Acquire multimedia data and extract features of the multimedia data.

[0152] In this embodiment, the acquired multimedia data is multimedia data that needs to be identified as AI-generated during the inference process, such as images, videos, text, or voice data. Extracting features from the multimedia data can refer to performing feature extraction on the multimedia data using a pre-trained feature extraction network, or can refer to converting the multimedia data into a feature matrix that can serve as a neural network input, which is not specifically limited in this embodiment.

[0153] Step 902: Input the features of the multimedia data into the teacher network and the student network respectively to obtain a first feature output by the teacher network and a second feature output by the student network.

[0154] In this embodiment, step 902 is similar to step 202 described above. For details, please refer to step 202 described above and will not be repeated here. Furthermore, the teacher network and student network in this embodiment are both trained neural networks, such as the teacher network and student network trained using the model training method described in the above embodiment.

[0155] Step 903: Determine the difference between the first feature and the second feature, and input the difference into a classification network to obtain a classification result output by the classification network. The classification result is used to indicate whether the multimedia data is generated by AI.

[0156] In other words, the features output by the teacher network and the features output by the student network will first calculate a difference value, and then the difference value will be processed based on the classification network to obtain the classification result of the multimedia data.

[0157] In this solution, by inputting the extracted multimedia data into the teacher network and the student network respectively, and performing classification of the multimedia data based on the difference between the features output by the teacher network and the student network, the classification network no longer performs classification directly based on the features of the multimedia data, but instead performs classification based on feature difference values ​​that are easier to perform classification, thereby improving the recognition accuracy of AI-generated multimedia data.

[0158] Optionally, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

[0159] Optionally, the teacher network is a pre-trained neural network, and the student network is trained under the guidance of the teacher network. The training goal of the student network in the training phase is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

[0160] Optionally, a training goal of the teacher network during the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the true classification results.

[0161] The above describes in detail the method provided by the embodiment of the present application. Next, the device provided by the embodiment of the present application for executing the above method will be introduced.

[0162] Please refer to Figure 11, which is a structural diagram of a model training device provided by an embodiment of the present application. As shown in Figure 11, the model training device provided by an embodiment of the present application includes: an extraction module 1101, which is used to obtain multimedia data and extract features of the multimedia data, and the multimedia data is training data in the training set; a processing module 1102, which is used to input the features of the multimedia data into the teacher network and the student network respectively, and obtain the first feature output by the teacher network and the second feature output by the student network; a training module 1103, which is used to train the student network based on the difference value between the first feature and the second feature; wherein the difference value is used to obtain the classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI. The training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

[0163] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

[0164] In one possible implementation, the extraction module 1101 is specifically used to: when the multimedia data is not generated by AI, input the multimedia data into a feature extraction network to obtain the features of the multimedia data output by the feature extraction network; or, when the multimedia data is generated by AI, input the multimedia data into a feature extraction network, and input the features output by the feature extraction network into a feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network.

[0165] In one possible implementation, the training module 1103 is further used to alternately train the student network and the feature enhancement network during the training process of the student network. The training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.

[0166] In one possible implementation, the teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.

[0167] In one possible implementation, the training module 1103 is further used to: after the student network is trained, input the features of the multimedia data into the trained student network to obtain a third feature; input the difference value between the first feature and the third feature into the classification network to obtain a classification result output by the classification network, and the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; based on the classification result output by the classification network and the actual classification result corresponding to the multimedia data, train the classification network, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the actual classification result.

[0168] In a possible implementation, the multimedia data is any one of the following data: image, video, text, or voice.

[0169] Please refer to Figure 12, which is a schematic diagram of the structure of a multimedia data classification device provided in an embodiment of the present application. As shown in Figure 12, the multimedia data classification device provided in an embodiment of the present application includes: an acquisition module 1201, which is used to acquire multimedia data and extract features of the multimedia data; a processing module 1202, which is used to input the features of the multimedia data into the teacher network and the student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; the processing module 1202 is also used to determine the difference between the first feature and the second feature, and input the difference into the classification network to obtain a classification result output by the classification network, and the classification result is used to indicate whether the multimedia data is generated by AI.

[0170] In one possible implementation, the classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

[0171] In one possible implementation, the teacher network is a pre-trained neural network, and the student network is trained under the guidance of the teacher network. The training goal of the student network in the training phase is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

[0172] In one possible implementation, the training goal of the teacher network during the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.

[0173] In a possible implementation, the multimedia data is any one of the following data: image, video, text, or voice.

[0174] Please refer to Figure 13, which is a structural diagram of an electronic device provided in an embodiment of the present application. The electronic device 1300 can be specifically manifested as a mobile phone, a tablet, a laptop computer, a smart wearable device, a server, etc., which is not limited here. Specifically, the electronic device 1300 includes: a receiving module 1301, a sending module 1302, a processor 1303 and a memory 1304 (wherein the number of processors 1303 in the electronic device 1300 can be one or more, and Figure 13 takes one processor as an example), wherein the processor 1303 may include an application processor 13031 and a communication processor 13032. In some embodiments of the present application, the receiving module 1301, the sending module 1302, the processor 1303 and the memory 1304 may be connected via a bus or other means.

[0175] Memory 1304 may include read-only memory and random access memory, and provides instructions and data to processor 1303. A portion of memory 1304 may also include non-volatile random access memory (NVRAM). Memory 1304 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0176] Processor 1303 controls the operation of the electronic device. In specific applications, the various components of the electronic device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0177] The method disclosed in the above embodiment of the present application can be applied to the processor 1303, or implemented by the processor 1303. The processor 1303 can be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 1303 or an instruction in the form of software. The above-mentioned processor 1303 can be a general-purpose processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.

[0178] The processor 1303 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. A general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1304, and the processor 1303 reads the information in the memory 1304 and completes the steps of the above method in combination with its hardware.

[0179] Receiving module 1301 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the electronic device. Transmitting module 1302 can be used to output digital or character information through the first interface. Transmitting module 1302 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group. Transmitting module 1302 can also include a display device such as a display screen.

[0180] The electronic device provided in the embodiment of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the electronic device executes the classification method of multimedia data described in the above embodiment, or so that the chip in the training device executes the model training method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device end, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0181] Specifically, see Figure 14 , which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1400. NPU 1400 is mounted on the host CPU (host CPU) as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1403, which is controlled by controller 1404 to extract matrix data from memory and perform multiplication operations.

[0182] In some implementations, the arithmetic circuit 1403 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1403 is a two-dimensional systolic array. The arithmetic circuit 1403 can also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1403 is a general-purpose matrix processor.

[0183] For example, assume there are input matrix A, weight matrix B, and output matrix C. The computation circuit retrieves the corresponding data of matrix B from weight memory 1402 and caches it on each PE in the computation circuit. The computation circuit then retrieves the data of matrix A from input memory 1401 and performs a matrix operation on it with matrix B. The partial or final matrix result is stored in accumulator 1408.

[0184] Unified memory 1406 is used to store input and output data. Weight data is directly transferred to weight memory 1402 through the Direct Memory Access Controller (DMAC) 1405. Input data is also transferred to unified memory 1406 through the DMAC.

[0185] BIU stands for Bus Interface Unit, i.e., bus interface unit 1410 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1409 .

[0186] The bus interface unit 1410 (BIU) is used for the instruction fetch memory 1409 to obtain instructions from the external memory, and is also used for the storage unit access controller 1405 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0187] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1406 or transfer weight data to the weight memory 1402 or transfer input data to the input memory 1401.

[0188] The vector calculation unit 1407 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1403, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0189] In some implementations, the vector calculation unit 1407 can store the processed output vector to the unified memory 1406. For example, the vector calculation unit 1407 can apply a linear function or a nonlinear function to the output of the operation circuit 1403, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1407 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1403, for example, for use in subsequent layers in a neural network.

[0190] An instruction fetch buffer 1409 connected to the controller 1404 is used to store instructions used by the controller 1404;

[0191] Unified memory 1406, input memory 1401, weight memory 1402, and instruction fetch memory 1409 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0192] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0193] Please refer to Figure 15, which is a schematic diagram of the structure of a computer-readable storage medium provided in an embodiment of the present application. The present application also provides a computer-readable storage medium. In some embodiments, the method disclosed in Figures 2 or 9 above can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or products.

[0194] 15 schematically illustrates a conceptual partial view of an example computer-readable storage medium including a computer program for executing a computer process on a computing device, arranged in accordance with at least some embodiments presented herein.

[0195] In one embodiment, computer readable storage medium 1500 is provided using signal bearing medium 1501. Signal bearing medium 1501 may include one or more program instructions 1502 that, when executed by one or more processors, may provide the functionality or portions of the functionality described above with respect to FIG. 2 or FIG.

[0196] In some examples, the signal bearing medium 1501 may include a computer readable medium 1503 such as, but not limited to, a hard drive, a compact disk (CD), a digital video disk (DVD), a digital tape, a memory, a ROM or RAM, and the like.

[0197] In some embodiments, the signal-bearing medium 1501 may include a computer-recordable medium 1504, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, or the like. In some embodiments, the signal-bearing medium 1501 may include a communication medium 1505, such as, but not limited to, a digital and / or analog communication medium (e.g., a fiber optic cable, a waveguide, a wired communication link, a wireless communication link, or the like). Thus, for example, the signal-bearing medium 1501 may be communicated via a wireless form of the communication medium 1505 (e.g., a wireless communication medium conforming to the IEEE 802.X standard or other transmission protocol).

[0198] The one or more program instructions 1502 may be, for example, computer-executable instructions or logic-implemented instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to the program instructions 1502 communicated to the computing device via one or more of computer-readable media 1503, computer-recordable media 1504, and / or communication media 1505.

[0199] It should also be noted that the device embodiments described above are merely illustrative, in which the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0200] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods of each embodiment of the present application.

[0201] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0202] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training equipment or data center to another website, computer, training equipment or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training equipment, data center, etc. that includes one or more available media integrations. Available media can be magnetic media, (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)), etc.

Claims

1. A model training method, characterized in that: include: Acquire multimedia data and extract features of the multimedia data, wherein the multimedia data is training data in a training set; Inputting the features of the multimedia data into the teacher network and the student network respectively, obtaining a first feature output by the teacher network and a second feature output by the student network; Training the student network based on the difference between the first feature and the second feature; The difference value is used to obtain a classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI. The training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

2. The method according to claim 1, characterized in that The classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

3. The method according to claim 1 or 2, characterized in that: The extracting and obtaining the features of the multimedia data includes: In the case where the multimedia data is not generated by AI, inputting the multimedia data into a feature extraction network to obtain features of the multimedia data output by the feature extraction network; Alternatively, in the case where the multimedia data is generated by AI, the multimedia data is input into the feature extraction network, and the features output by the feature extraction network are input into a feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network.

4. The method according to claim 3, characterized in that The method further comprises: During the training process of the student network, the student network and the feature enhancement network are trained alternately, and the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.

5. The method according to any one of claims 1 to 4, characterized in that: The teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.

6. The method according to any one of claims 1 to 5, characterized in that: After the student network training is completed, the method further includes: Inputting the feature of the multimedia data into the trained student network to obtain a third feature; Inputting a difference value between the first feature and the third feature into a classification network to obtain a classification result output by the classification network, wherein the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; Based on the classification result output by the classification network and the real classification result corresponding to the multimedia data, the classification network is trained, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the real classification result.

7. The method according to any one of claims 1 to 6, characterized in that: The multimedia data is any one of the following data: image, video, text or voice.

8. A method for classifying multimedia data, characterized in that: include: Acquire multimedia data and extract features of the multimedia data; Inputting the features of the multimedia data into the teacher network and the student network respectively, obtaining a first feature output by the teacher network and a second feature output by the student network; Determine a difference value between the first feature and the second feature, and input the difference value into a classification network to obtain a classification result output by the classification network, wherein the classification result is used to indicate whether the multimedia data is generated by AI.

9. The method according to claim 8, characterized in that The classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

10. The method according to claim 8 or 9, characterized in that: The teacher network is a pre-trained neural network, the student network is trained under the guidance of the teacher network, and the training goal of the student network in the training stage is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

11. The method according to claim 10, characterized in that The training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.

12. A model training device, characterized in that: include: An extraction module, used to acquire multimedia data and extract features of the multimedia data, wherein the multimedia data is training data in a training set; A processing module, used for inputting the features of the multimedia data into a teacher network and a student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; A training module, configured to train the student network based on a difference value between the first feature and the second feature; The difference value is used to obtain a classification result of the multimedia data, and the classification result is used to indicate whether the multimedia data is generated by AI. The training goal of the student network is to reduce the difference between the features output by the student network and the features output by the teacher network when processing non-AI generated multimedia data, and to increase the difference between the features output by the student network and the features output by the teacher network when processing AI generated multimedia data.

13. The device according to claim 12, characterized in that The classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

14. The device according to claim 12 or 13, characterized in that The extraction module is specifically used for: In the case where the multimedia data is not generated by AI, inputting the multimedia data into a feature extraction network to obtain features of the multimedia data output by the feature extraction network; Alternatively, in the case where the multimedia data is generated by AI, the multimedia data is input into the feature extraction network, and the features output by the feature extraction network are input into a feature enhancement network to obtain the features of the multimedia data output by the feature enhancement network.

15. The device according to claim 14, characterized in that The training module is also used to alternately train the student network and the feature enhancement network during the training process of the student network, and the training goal of the feature enhancement network is to reduce the difference between the features output by the student network and the features output by the teacher network.

16. The device according to any one of claims 12 to 15, characterized in that: The teacher network is a pre-trained neural network, and the training goal of the teacher network in the training phase is to reduce the difference between the classification results obtained based on the features output by the teacher network and the actual classification results.

17. The device according to any one of claims 12 to 16, characterized in that: The training module is also used for: After the student network is trained, the feature of the multimedia data is input into the trained student network to obtain a third feature; Inputting a difference value between the first feature and the third feature into a classification network to obtain a classification result output by the classification network, wherein the classification result output by the classification network is used to indicate whether the multimedia data is generated by AI; Based on the classification result output by the classification network and the real classification result corresponding to the multimedia data, the classification network is trained, and the training goal of the classification network is to reduce the difference between the classification result output by the classification network and the real classification result.

18. A multimedia data classification device, characterized in that: include: An acquisition module, used to acquire multimedia data and extract features of the multimedia data; A processing module, used for inputting the features of the multimedia data into a teacher network and a student network respectively, to obtain a first feature output by the teacher network and a second feature output by the student network; The processing module is also used to determine the difference value between the first feature and the second feature, and input the difference value into a classification network to obtain a classification result output by the classification network, and the classification result is used to indicate whether the multimedia data is generated by AI.

19. The device according to claim 18, characterized in that The classification result is specifically used to indicate the probability that the multimedia data is generated by AI and the probability that the multimedia data is not generated by AI, and the probability that the multimedia data is generated by AI has a positive correlation with the difference value.

20. A model training device, characterized in that: The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 1 to 7.

21. A multimedia data classification device, characterized in that: The device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the device executes the method according to any one of claims 8 to 11.

22. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 11.

Citation Information

Patent Citations

  • Model training method, multimedia data classification method and related device

    CN120031102A

  • Neural network training method, video frame processing method and related equipment

    CN111401406A

  • Industrial part surface defect detection method and system based on anomaly detection algorithm

    CN113902710A

  • Method and device for training and identifying living body model, equipment and medium

    CN117037294A

  • Fast anomaly detection method and system based on contrastive representation distillation

    US20230368372A1