Multimedia data authenticity identification method and device
By integrating feature extraction and cross-attention mechanisms for multimedia data, the problem of how to identify forged content generated based on AI technology is solved, and the authenticity of multimedia data is identified and social risks are reduced.
Patent Information
- Application Number
- CN202510073842.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
As the content generated by content generation technology based on AI technology becomes more and more realistic, it is difficult to distinguish its authenticity, criminals use this technology to generate fake content, causing negative social impact. How to provide an identification method to identify content generated by the aforementioned content generated by content generation technology based on AI technology has become an urgent problem.
By obtaining the multimedia data to be identified, including video and audio, extracting them, and using a cross-attention mechanism to fuse the video and audio features, the authenticity of the multimedia data is determined.
The authenticity of multimedia data is realized, and the forged content generated based on AI technology can be effectively identified, thereby reducing the risk of criminals using this technology to engage in malicious behavior.
Smart Images

Figure CN119989050A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a method and device for authenticating multimedia data. Background Art
[0002] With the development of artificial intelligence (AI) technology, content generation technology based on AI technology is becoming more and more mature, and the content it generates is becoming more and more realistic, and it is becoming increasingly difficult to distinguish its authenticity. For example, Deepfake technology is a technology that uses AI technology to generate seemingly real fake videos and / or fake audio; another example is AIGC (Artificial Intelligence Generated Content) technology, which is based on AI technology and generates new data (i.e., new content, such as text, images, audio and video) with generalization capabilities through learning and identifying existing data.
[0003] As the content generated by AI-based content generation technology becomes more and more realistic and more and more difficult to distinguish its authenticity, in some scenarios, it is inevitable that criminals will use such technology to generate fake content for profit, such as generating fake ID images or fake videos containing faces to forge other users' identities, steal other users' information and property, etc.; for example, based on Deepfake technology, realistic false content is generated to create fake news and / or to carry out malicious acts such as defamation and fraud, all of which will have a negative impact on society. Therefore, how to provide an identification method to identify the aforementioned generated content as content generated by AI-based content generation technology has become an urgent problem to be solved. Summary of the invention
[0004] One or more embodiments of the present specification provide a method and device for authenticating multimedia data, so as to achieve authenticity authentication of multimedia data.
[0005] According to a first aspect, a method for authenticating multimedia data is provided, comprising:
[0006] Acquire first multimedia data to be authenticated, wherein the first multimedia data includes a first video and a first audio corresponding thereto;
[0007] Extracting features from the first video to obtain first video features;
[0008] Extracting features from the first audio to obtain first audio features;
[0009] Using a cross attention mechanism, based on the first video feature and the first audio feature, obtain a video fusion feature and an audio fusion feature;
[0010] Based on the video fusion feature and the audio fusion feature, the authenticity of the first multimedia data is determined.
[0011] According to a second aspect, a method for authenticating multimedia data is provided, comprising:
[0012] An acquisition module, configured to acquire first multimedia data to be authenticated, wherein the first multimedia data includes a first video and a first audio corresponding thereto;
[0013] A first obtaining module is configured to extract features from the first video to obtain first video features;
[0014] A second obtaining module is configured to extract features from the first audio to obtain first audio features;
[0015] A third obtaining module is configured to adopt a cross attention mechanism to obtain a video fusion feature and an audio fusion feature based on the first video feature and the first audio feature;
[0016] The determination module is configured to determine the authenticity of the first multimedia data based on the video fusion feature and the audio fusion feature.
[0017] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in the first aspect.
[0018] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in the first aspect is implemented.
[0019] According to the method and device for identifying the authenticity of multimedia data provided by the embodiments of this specification, after obtaining the first multimedia data including the first video and the first audio, feature extraction is performed on the first video and the first audio respectively to obtain their respective first video features and first audio features, and then a cross-attention mechanism is adopted to fuse part of the features of the first audio with the first video features, and to fuse part of the features of the first video with the first audio features, so as to obtain corresponding video fusion features and audio fusion features; then, based on the video fusion features and the audio fusion features, the authenticity of the first multimedia data is determined to achieve authentication of the authenticity of the multimedia data. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0021] Figure 1 A schematic diagram of an implementation framework of an embodiment disclosed in this specification;
[0022] Figure 2 A schematic diagram of a flow chart of a method for identifying the authenticity of multimedia data provided by an embodiment;
[0023] Figure 3A A schematic diagram of a model structure of a model for determining the authenticity identification result of multimedia data provided by an embodiment;
[0024] Figure 3B A schematic diagram of a process for determining the authenticity identification result of multimedia data provided by an embodiment;
[0025] Figure 4 A schematic diagram of the deformation process of the Mel-spectrogram;
[0026] Figure 5 A schematic block diagram of a multimedia data authenticity identification device provided in an embodiment. DETAILED DESCRIPTION
[0027] The technical solutions of the embodiments of this specification will be described in detail below with reference to the accompanying drawings.
[0028] The embodiments of this specification disclose a method and device for authenticating multimedia data. The application scenarios and technical concepts of the method are first introduced as follows:
[0029] As mentioned above, as the content generated by AI-based content generation technology becomes more and more realistic and more and more difficult to distinguish its authenticity, in some scenarios, it is inevitable that criminals will use such technology to generate fake content for profit, such as generating fake ID images or fake videos containing faces to forge other users' identities, steal other users' information and property, etc.; for example, based on Deepfake technology, realistic false content is generated to create fake news and / or to carry out malicious acts such as defamation and fraud. All of the above situations will have a negative impact on society. So, how to provide an identification method to identify the aforementioned generated content as content generated by AI-based content generation technology has become an urgent problem to be solved.
[0030] In view of this, the inventor proposes a method for authenticating multimedia data. Figure 1A schematic diagram of an implementation scenario according to an embodiment disclosed in this specification is shown. In this implementation scenario, first multimedia data to be authenticated is obtained, wherein the first multimedia data includes a first video and its corresponding first audio; feature extraction is performed on the first video to obtain first video features; feature extraction is performed on the first audio to obtain first audio features; a cross-attention mechanism is used to obtain a video fusion feature that integrates audio synchronization features and an audio fusion feature that integrates video synchronization features based on the first video features and the first audio features; and then the authenticity of the first multimedia data is determined based on the video fusion feature and the audio fusion feature.
[0031] In the above process, through the cross-attention mechanism, the synchronized features of the first audio are fused with the first video features to obtain video fusion features, and the synchronized features of the first video are fused with the first audio features to obtain audio fusion features. Then, based on such video fusion features and audio fusion features, a consistency judgment is made between them to determine the authenticity of the first multimedia data, thereby realizing the authentication of the authenticity of the multimedia data.
[0032] The authenticity identification method of multimedia data provided in this specification is described in detail below in conjunction with specific embodiments.
[0033] Figure 2 The flowchart of a method for authenticating multimedia data in one embodiment of the present specification is shown. The method is executed by an electronic device, and the electronic device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities. In the process of authenticating multimedia data, if Figure 2 As shown, the method includes the following steps S210-S250:
[0034] In step S210, first multimedia data to be authenticated is obtained, wherein the first multimedia data includes a first video and a first audio corresponding thereto.
[0035] The first multimedia data may be any multimedia data that needs to be identified. The first multimedia data includes a first video and its corresponding first audio, and may be, for example, but not limited to: a short video containing audio, a news video, a video containing audio generated during a video call, and a film or TV series or its clips. In some examples, the electronic device may obtain multimedia data to be identified that is input by a user or sent by other devices as the first multimedia data.
[0036] Afterwards, in step S220, feature extraction is performed on the first video to obtain first video features.
[0037] In some possible examples, a traditional feature extraction algorithm may be used to extract features from the first video to obtain first video features. In some other possible examples, a coding network based on deep learning may be used to extract features from the first video to obtain first video features. The deep learning-based coding network, i.e., the first coding network, may be implemented as a coding network based on RNN (Recurrent Neural Network) or a coding network based on a Transformer structure, etc., which may be used to process sequence data.
[0038] In some possible implementations, step S220 may include step 11:
[0039] In step 11, a first encoding network is used to extract features of the first video to obtain first video features, wherein the first encoding network processes the first video based on a pooling attention mechanism and a relative position encoding method.
[0040] In this step, the electronic device can input the first video, that is, the video frame sequence composed of multiple video frames, into the first encoding network to extract features of the first video, that is, the video frame sequence, through the first encoding network to obtain first video features, wherein the first encoding network processes the first video based on the pooling attention mechanism and the relative position encoding method, such as Figure 3A shown.
[0041] In some possible implementations, the first coding network is implemented as a coding network based on a Transformer structure as an example for explanation, wherein the first coding network may include multiple network layers, each network layer may process its input based on a pooled attention mechanism and a relative position encoding method, wherein the input of the first network layer may be the first video, and the input of other network layers may be the output of the previous network layer. Based on the pooled attention mechanism, the resource consumption of the encoding process of the first video may be reduced, and the encoding efficiency may be improved. In some other possible implementations, each network layer of the first coding network may also process its input based on other self-attention mechanisms, etc.
[0042] Exemplarily, the data processing process in a single network layer (taking network layer i as an example) is described below: the output i-1 of the previous network layer i-1 of network layer i is obtained; the query Q weight matrix, key K weight matrix and value V weight matrix of network layer i are used to process the output i-1 respectively to obtain the corresponding third Q matrix, third K matrix and third V matrix; then the downsampling factor P is used to Q , P K and P v, respectively, the third Q matrix, the third K matrix and the third V matrix are pooled to obtain the fourth Q matrix, the fourth K matrix and the fourth V matrix, so as to realize the processing of the input based on the pooling attention mechanism, where the downsampling factor P Q , P K and P v They can be the same, or downsampled by a factor of P K and P v Possibly with P Q different.
[0043] The fourth Q matrix, the fourth K matrix and the fourth V matrix can be expressed by the following formula (1):
[0044]
[0045] Wherein, Q represents the fourth Q matrix, K represents the fourth K matrix, V represents the fourth V matrix, and W Q , W K and W V They represent the Q weight matrix, K weight matrix and V weight matrix of network layer i respectively, and X represents output i-1 or the normalized output i-1 mentioned later.
[0046] Next, the fourth Q matrix, the fourth K matrix and the fourth V matrix are used to obtain the initial output of the network layer i through the first attention formula of the network layer i and the relative position encoding method. The first attention formula can be expressed as the following formula (2):
[0047]
[0048] in, O represents the initial output of network layer i, d k represents the dimension of the second K matrix, and softmax(.) represents the activation function. p(m),p(n) represents the relative position encoding between the mth element (patch) and the nth element in output i-1, Q m Represents the query value corresponding to the mth element in the fourth Q matrix.
[0049] in, Understandable, R. h , R w and R tIt is the position code corresponding to each element (e.g., the mth element and the nth element) in the output i-1 along the height, width, and time axis, h(m), w(m), and t(m) respectively represent the position of the mth element in the output i-1 in the height (vertical), width (horizontal), and time axis, and h(n), w(n), and t(n) respectively represent the position of the nth element in the output i-1 in the height (vertical), width (horizontal), and time axis. It can be understood that each video frame in the first video corresponds to a position on the time axis.
[0050] In some other implementations, in order to better improve the performance and expression ability of the first coding network, network layer i also includes a residual connection sublayer. Accordingly, after obtaining the above-mentioned initial output O of network layer i, the above-mentioned initial output O of network layer i and the input of network layer i, that is, the aforementioned output i-1, can be superimposed through the residual connection sublayer, and the superposition result is used as the output of network layer i, and then as the input of the next network layer i+1 of network layer i.
[0051] It can be understood that before being input into the first encoding network, each video frame in the first video can be pre-divided into multiple image blocks of a first size, namely patches, as elements in the input, namely as processing units of the first encoding network, and the first size is a preset size.
[0052] In some cases, after obtaining the output i-1 of the previous network layer i-1 of network layer i, the output i-1 can be normalized first, and then the query Q weight matrix, key K weight matrix and value V weight matrix of network layer i are used to process the output i-1 respectively to obtain the corresponding third Q matrix, third K matrix and third V matrix.
[0053] The data processing process of other network layers of the first coding network can refer to the data processing process of the aforementioned network layer i, which will not be described in detail here. By analogy, the first video feature can be obtained.
[0054] In some possible examples, the output of the last network layer can be used as the first video feature; or, in some other possible examples, the electronic device can obtain the output of at least two specified network layers of the first coding network, and based on the output of the at least two specified network layers, jointly determine to obtain the first video feature, wherein the at least two specified network layers can be all network layers of the first coding network. The process of jointly determining to obtain the first video feature can be: using a specific algorithm, upsampling or downsampling the output of the at least two specified network layers to unify the size of the output of the at least two specified network layers, and obtaining the output of the specified size corresponding to each specified network layer, and then, performing aggregation processing on the output of the specified size corresponding to each specified network layer to obtain the first video feature. The aggregation processing is, for example, superposition processing, or aggregation processing based on an attention mechanism, etc. In this way, a first video feature that aggregates the multi-scale features of the first video and makes the features richer can be obtained, providing a basis for subsequently improving the accuracy of the authenticity identification results of multimedia data.
[0055] In step S230, feature extraction is performed on the first audio to obtain first audio features. In this step, the first audio may be preprocessed to obtain a Mel-spectrogram corresponding to the audio.
[0056] In some other possible implementations, step S230 may include steps 21-22:
[0057] In step 21, the first audio is preprocessed to obtain a Mel-spectrogram corresponding to the first audio. Exemplarily, the preprocessing may include, for example: pre-emphasis processing, frame processing, short-time Fourier transform (STFT), filtering processing using a Mel filter bank, and logarithmic compression processing. Accordingly, by performing a series of preprocessing on the first audio, a Mel-spectrogram corresponding to the first audio is obtained. Among them, the Mel-spectrogram can be understood as a two-dimensional matrix. For example, the shape of the Mel-spectrogram can be expressed as [T, F], where T represents the number of time frames and F represents the Mel frequency. In some other examples, the preprocessing may also include discrete cosine transform (DCT) processing while including the aforementioned processing. Accordingly, by performing a series of preprocessing on the first audio, a Mel-spectrogram [T, F] corresponding to the first audio is obtained. At this time, T represents the number of time frames and F can represent the length of the Mel-frequency cepstrum feature.
[0058] In step 22, a second encoding network is used to perform feature extraction on the Mel-spectrogram to obtain a first audio feature.
[0059] In some possible examples, in order to ensure synchronization between audio and video, so as to better determine the consistency between them and to better improve the accuracy of the authenticity identification result of the first multimedia data, step 22 may include the following steps 221-222:
[0060] In step 221, based on the specified deformation operation, the Mel-spectrogram is deformed to obtain multiple spectrograms corresponding to multiple video frames in the first video. This is to achieve synchronization between the video and the audio. Exemplarily, the specified deformation operation may include a reshape operation. For example, the shape of the Mel-spectrogram is [T, F], where Figure 4 As shown, assuming T=3, multiple video frames in the first video correspond to multiple time frames (i.e., the positions of the aforementioned time axis), that is, the first video includes T frames (i.e., 3 frames) of video frames, and the Mel-spectrogram can be reshaped into 3 spectrograms of shape [F, 1]. Figure 4 As shown, the T spectrograms correspond to T video frames respectively, so as to achieve synchronization between the audio spectrogram and the first video in the time dimension.
[0061] In step 222, the plurality of spectrograms are input into a second coding network, and the second coding network is used to extract features from the plurality of spectrograms to obtain a first audio feature. In this step, a spectrogram sequence consisting of a plurality of spectrograms is input into a second coding network, and the second coding network is used to extract features from the plurality of spectrograms to obtain a first audio feature. Figure 3A shown.
[0062] Exemplarily, the second coding network can be implemented as a coding network based on RNN (Recurrent Neural Network, recurrent neural network), or as a coding network based on a Transformer structure, etc., which can be used to process sequence data. Taking the second coding network as a coding network based on a Transformer structure as an example, the second coding network can include multiple processing layers, each processing layer can process its input based on a self-attention mechanism or the aforementioned pooling attention mechanism, wherein the input of the first processing layer can be a first video, and the input of the other processing layers is the output of the previous processing layer. Afterwards, the output of the last processing layer can be used as the first audio feature, or the first audio feature can be determined based on the output of at least two specified processing layers, wherein the process of determining the first audio feature based on the output of at least two specified processing layers can refer to the aforementioned process of determining the first video feature based on the output of at least two specified network layers, which will not be repeated here. Exemplarily, the second coding network can be implemented as a ViT (Vision Transformer) network.
[0063] Next, after obtaining the first video feature and the first audio feature, in step S240, a cross-attention mechanism is used to obtain a video fusion feature and an audio fusion feature based on the first video feature and the first audio feature. In this step, a cross-attention mechanism is used to integrate the synchronized first audio feature into the first video feature to obtain a video fusion feature, and integrate the synchronized first video feature into the first audio feature to obtain an audio fusion feature. Exemplarily, the electronic device can obtain a video fusion feature and an audio fusion feature based on the first video feature and the first audio feature through an aggregation network, wherein the aggregation network is a network using a cross-attention mechanism, such as Figure 3A shown.
[0064] In some possible implementations, step S240 may include steps 31-33:
[0065] In step 31, a cross attention mechanism is used to determine the first query matrix, the first key matrix and the first value matrix corresponding to the first video feature. In this step, an aggregation network (such as Figure 3A As shown), based on the first video feature, determine the first query matrix, the first key matrix and the first value matrix corresponding to it, such as Figure 3B As shown. The implementation principle of determining the first query matrix, the first key matrix and the first value matrix is similar to the implementation principle of determining the third query matrix, the third key matrix and the third value matrix. The implementation process can refer to the implementation process of determining the third query matrix, the third key matrix and the third value matrix, which will not be described in detail here.
[0066] In step 32, a cross attention mechanism is used to determine the second query matrix, the second key matrix and the second value matrix corresponding to the first audio feature. In this step, an aggregation network (such as Figure 3A As shown), based on the first audio feature, determine the corresponding second query matrix, second key matrix and second value matrix, such as Figure 3B As shown. The implementation principle of determining the second query matrix, the second key matrix and the second value matrix is similar to the implementation principle of determining the third query matrix, the third key matrix and the third value matrix. The implementation process can refer to the implementation process of determining the third query matrix, the third key matrix and the third value matrix, which will not be described in detail here.
[0067] Exemplarily, the query weight matrix, key weight matrix and value weight matrix used to determine the first query matrix, the first key matrix and the first value matrix may be the same as or different from the query weight matrix, key weight matrix and value weight matrix used to determine the second query matrix, the second key matrix and the second value matrix, and both may be determined based on the training process of the aggregation network.
[0068] In step 33, a cross attention mechanism is adopted to combine the first query matrix, the second key matrix, the second value matrix, the second query matrix, the first key matrix and the first value matrix to determine the video fusion feature and the audio fusion feature.
[0069] In some possible examples, step 33 may include 331-332:
[0070] In step 331, based on the first query matrix, the second key matrix and the second value matrix, the video fusion feature is determined. In this step, the video fusion feature is determined based on the first query matrix, the second key matrix and the second value matrix through the second attention formula of the aforementioned aggregation network. In one implementation, the result (hereinafter referred to as the first result) calculated based on the first query matrix, the second key matrix and the second value matrix through the second attention formula can be directly used as the video fusion feature. In another implementation, the result of superimposing the aforementioned first result and the first video feature can be used as the video fusion feature, as shown in Figure 3. In this step, the feature of synchronizing the first audio in the first video feature can be realized. Among them, the second attention formula can refer to the aforementioned first attention formula, and the position encoding therein can be implemented as the aforementioned relative position encoding method, or it can be implemented as other position encoding methods.
[0071] In step 332, the audio fusion feature is determined based on the second query matrix, the first key matrix and the first value matrix. In this step, the implementation principle of determining the audio fusion feature based on the second query matrix, the first key matrix and the first value matrix is similar to the implementation principle of determining the video fusion feature based on the first query matrix, the second key matrix and the second value matrix. The implementation process can refer to the implementation process of determining the video fusion feature based on the first query matrix, the second key matrix and the second value matrix, which will not be described in detail here. In this step, the synchronous feature of the first video can be integrated into the first audio feature. As shown in Figure 3.
[0072] In some other possible examples, a cross-attention mechanism can also be used to combine the first query matrix, the first key matrix, the first value matrix, the second key matrix and the second value matrix to determine the video fusion feature; a cross-attention mechanism can be used to combine the second query matrix, the first key matrix, the first value matrix, the second key matrix and the second value matrix to determine the audio fusion feature. Taking the combination of the first query matrix, the first key matrix and the first value matrix, the second key matrix and the second value matrix to determine the video fusion feature as an example, the following is explained: the first key matrix and the second key matrix are aggregated to obtain an aggregated key matrix; the first value matrix and the second value matrix are aggregated to obtain an aggregated value matrix; then, based on the first query matrix, the aggregated key matrix and the aggregated value matrix, the video fusion feature is determined. In order to achieve the synchronization feature of the first audio integrated into the first video feature. Similarly, the audio fusion feature is obtained. Among them, the aggregation can be superposition or element multiplication calculation.
[0073] In step S250, the authenticity of the first multimedia data is determined based on the video fusion feature and the audio fusion feature. In this step, the authenticity of the first multimedia data is determined based on the video fusion feature and the audio fusion feature through the trained classification network. Figure 3A shown.
[0074] Exemplarily, the classification network can be implemented as a fully connected layer, or other types of classifiers.
[0075] In some possible examples, the classification network is trained based on several sample multimedia data and their corresponding label data, wherein the label data includes a first label indicating that the corresponding sample multimedia data is real video and real audio, a second label indicating that the corresponding sample multimedia data is real video and false audio, a third label indicating that the corresponding sample multimedia data is false video and true audio, and a fourth label indicating that the corresponding sample multimedia data is false video and false audio.
[0076] Among them, fake video (and fake audio) may refer to all or part of the video clips (and all or part of the audio clips) being generated by content generation technology based on artificial intelligence technology. That is, for sample multimedia data, if all or part of the video clips included therein are generated by content generation technology based on artificial intelligence technology, then the video needs to be identified as a fake video (then the sample multimedia data corresponds to the third label or the fourth label); similarly, if all or part of the audio clips included therein are generated by content generation technology based on artificial intelligence technology, then the audio needs to be identified as fake audio (then the sample multimedia data corresponds to the fourth label or the second label).
[0077] True video (and true audio) may refer to the entire video (and the entire audio) not containing content generated by content generation technology based on artificial intelligence technology. For example, the entire video is collected by real image acquisition equipment, and the entire audio is collected by real recording equipment.
[0078] By dividing the label data more finely into true and false for video and audio in multimedia data, that is, separately modeling the video and audio modal data, the learning accuracy of the classification network can be better improved, and the accuracy and precision of the classification results of the classification network can be improved.
[0079] In some other possible examples, the aforementioned first encoding network, second encoding network, aggregation network, and classification network can be jointly trained based on the aforementioned plurality of sample multimedia data and their corresponding label data.
[0080] In some possible implementations, step S250 may include steps 41-42:
[0081] In step 41, the video fusion feature and the audio fusion feature are concatenated to obtain a concatenated fusion feature. In this step, the video fusion feature and the audio fusion feature can be concatenated by a concat operation to obtain a concatenated fusion feature, such as Figure 3B shown.
[0082] In step 42, the authenticity of the first multimedia data is determined based on the splicing and fusion features through the trained classification network. In this step, the splicing and fusion features are input into the classification network to process the splicing and fusion features through the classification network to obtain the output of the classification network, and the authenticity of the first multimedia data is determined based on the output of the classification network. In the case where the label data for training the classification network includes the aforementioned first label, second label, third label and fourth label, if the output of the classification network indicates that the first video in the first multimedia data is true and the first audio is true, then the first multimedia data is determined to be true, that is, the first video and the first audio are not generated by the content generation technology based on artificial intelligence technology; on the contrary, if the output of the classification network indicates that the first video in the first multimedia data is false, and / or the first audio is false, that is, the output of the classification network indicates that all or part of the first video segments in the first multimedia data are generated by the content generation technology based on artificial intelligence technology, and / or indicates that all or part of the first audio segments in the first multimedia data are generated by the content generation technology based on artificial intelligence technology, then the first multimedia data is determined to be false.
[0083] Accordingly, in some possible examples, step 42 may include steps 421-422:
[0084] In step 421, based on the splicing and fusion features, the first authenticity result corresponding to the first video and the second authenticity result of the first audio are determined through the trained classification network. In this step, the electronic device inputs the splicing and fusion features into the classification network, and processes the splicing and fusion features through the classification network to obtain the first authenticity result corresponding to the first video and the second authenticity result of the first audio. The first authenticity result can indicate whether the first video is true or false; the second authenticity result can indicate whether the first audio is true or false.
[0085] Then, in step 422, the authenticity of the first multimedia data is determined based on the first authenticity result and the second authenticity result. In this step, if the first authenticity result indicates that the first video is authentic and the second authenticity result indicates that the first audio is authentic, the first multimedia data is determined to be authentic; conversely, if the first authenticity result indicates that the first video is false and / or the second authenticity result indicates that the first audio is false, the first multimedia data is determined to be false, that is, if one of the first authenticity result and the second authenticity result indicates that the corresponding video or audio is false, the first multimedia data is determined to be false.
[0086] In the above process, the first video feature and the first audio feature are processed by the cross-attention mechanism to obtain the video fusion feature that learns the features of the synchronized audio and the audio fusion feature that learns the features of the synchronized video, and then the authenticity of the first multimedia data is determined by using such video fusion features and audio fusion features. The judgment of the consistency between the video and the audio can be better realized, and then the authenticity of the multimedia data in which part of the content in the included video and / or audio is generated (i.e., generated based on artificial intelligence technology) and is forged, and the other part of the content is real can be recognized. In addition, the first coding network processes the first video based on the pooled attention mechanism and the relative position coding method, which can reduce the computing resource consumption of the authenticity identification process of the multimedia data, and based on the output of at least two specified network layers of the first coding network, the first video feature is jointly determined, and the first video feature with richer features can be obtained, which can better improve the accuracy of the authenticity identification results of the multimedia data.
[0087] The foregoing describes certain embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps described in the claims may be performed in an order different from that in the embodiments, and the desired results may still be achieved. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0088] Corresponding to the above method embodiment, the present specification embodiment provides a multimedia data authenticity identification device 500, whose schematic block diagram is as follows: Figure 5 As shown, including:
[0089] An acquisition module 510 is configured to acquire first multimedia data to be identified, wherein the first multimedia data includes a first video and a first audio corresponding thereto;
[0090] A first obtaining module 520 is configured to extract features from the first video to obtain first video features;
[0091] A second obtaining module 530 is configured to extract features from the first audio to obtain first audio features;
[0092] A third obtaining module 540 is configured to adopt a cross attention mechanism to obtain a video fusion feature and an audio fusion feature based on the first video feature and the first audio feature;
[0093] The determination module 550 is configured to determine the authenticity of the first multimedia data based on the video fusion feature and the audio fusion feature.
[0094] In some possible implementations, the first obtaining module 520 is specifically configured to use a first encoding network to extract features from the first video to obtain features of the first video, wherein the first encoding network processes the first video based on a pooled attention mechanism and a relative position encoding method.
[0095] In some possible implementations, the second obtaining module 530 includes:
[0096] A first obtaining unit (not shown in the figure) is configured to preprocess the first audio to obtain a Mel-spectrogram corresponding to the first audio;
[0097] The second obtaining unit (not shown in the figure) is configured to use a second encoding network to perform feature extraction on the Mel-spectrogram to obtain the first audio feature.
[0098] In some possible implementations, the second obtaining unit is specifically configured to deform the Mel-spectrogram based on a specified deformation operation to obtain a plurality of spectrograms corresponding to a plurality of video frames in the first video;
[0099] The multiple spectrograms are input into the second encoding network, so as to use the second encoding network to perform feature extraction on the multiple spectrograms to obtain the first audio feature.
[0100] In some possible implementations, the third obtaining module 540 includes:
[0101] A first determining unit (not shown in the figure) is configured to adopt the cross attention mechanism to determine the first query matrix, the first key matrix and the first value matrix corresponding to the first video feature based on the first video feature;
[0102] A second determination unit (not shown in the figure) is configured to adopt the cross attention mechanism to determine the second query matrix, the second key matrix and the second value matrix corresponding to the first audio feature based on the first audio feature;
[0103] A third determination unit (not shown in the figure) is configured to adopt the cross-attention mechanism to jointly determine the video fusion features and the audio fusion features with the first query matrix, the second key matrix, the second value matrix, the second query matrix, the first key matrix and the first value matrix.
[0104] In some possible implementations, the third determining unit is specifically configured to determine the video fusion feature based on the first query matrix, the second key matrix, and the second value matrix;
[0105] The audio fusion feature is determined based on the second query matrix, the first key matrix and the first value matrix.
[0106] In some possible implementations, the determining module 550 is specifically configured to splice the video fusion feature and the audio fusion feature to obtain a spliced fusion feature;
[0107] Based on the splicing and fusion features, the authenticity of the first multimedia data is determined through a trained classification network.
[0108] In some possible implementations, the determination module 550 is specifically configured to determine, based on the splicing and fusion features, a first authenticity result corresponding to the first video and a second authenticity result corresponding to the first audio through a trained classification network;
[0109] The authenticity of the first multimedia data is determined based on the first authenticity result and the second authenticity result.
[0110] In some possible implementations, the classification network is trained based on a number of sample multimedia data and their corresponding label data, wherein the label data includes a first label indicating that the corresponding sample multimedia data is real video and real audio, a second label indicating that the corresponding sample multimedia data is real video and false audio, a third label indicating that the corresponding sample multimedia data is false video and true audio, and a fourth label indicating that the corresponding sample multimedia data is false video and false audio.
[0111] The above device embodiments correspond to the method embodiments. For specific descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, please refer to the corresponding method embodiments.
[0112] The embodiments of the present specification also provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the authenticity identification method of multimedia data provided in the present specification.
[0113] The embodiment of the present specification also provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the authenticity identification method of the multimedia data provided in the present specification is implemented.
[0114] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the storage medium and computing device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0115] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the embodiments of the present invention may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0116] The specific implementation methods described above further describe the purpose, technical solutions and beneficial effects of the embodiments of the present invention in detail. It should be understood that the above description is only a specific implementation method of the embodiments of the present invention and is not intended to limit the scope of protection of the present invention. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solution of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A method for identifying the authenticity of multimedia data, comprising: Acquire first multimedia data to be authenticated, wherein the first multimedia data includes a first video and a first audio corresponding thereto; Extracting features from the first video to obtain first video features; Extracting features from the first audio to obtain first audio features; Using a cross attention mechanism, based on the first video feature and the first audio feature, obtain a video fusion feature and an audio fusion feature; Based on the video fusion feature and the audio fusion feature, the authenticity of the first multimedia data is determined.
2. The method of claim 1, wherein: The extracting features of the first video to obtain first video features includes: A first encoding network is used to extract features of the first video to obtain features of the first video, wherein the first encoding network processes the first video based on a pooled attention mechanism and a relative position encoding method.
3. The method of claim 1, wherein: The extracting features of the first audio to obtain first audio features includes: Preprocessing the first audio to obtain a Mel-spectrogram corresponding to the first audio; A second encoding network is used to perform feature extraction on the Mel-spectrogram to obtain the first audio feature.
4. The method of claim 3, wherein: The step of extracting features from the Mel-spectrogram using the second encoding network to obtain the first audio feature includes: Based on a specified deformation operation, deforming the Mel-spectrogram to obtain a plurality of spectrograms corresponding to a plurality of video frames in the first video; The multiple spectrograms are input into the second encoding network, so as to use the second encoding network to perform feature extraction on the multiple spectrograms to obtain the first audio feature.
5. The method of claim 1, wherein: The obtaining of the video fusion feature and the audio fusion feature includes: Using the cross attention mechanism, based on the first video feature, determine a first query matrix, a first key matrix and a first value matrix corresponding to the first video feature; Using the cross attention mechanism, based on the first audio feature, determine a second query matrix, a second key matrix, and a second value matrix corresponding to the first audio feature; The cross-attention mechanism is adopted to combine the first query matrix, the second key matrix, the second value matrix, the second query matrix, the first key matrix and the first value matrix to determine the video fusion feature and the audio fusion feature.
6. The method of claim 5, wherein: The determining the video fusion feature and the audio fusion feature includes: Determining the video fusion feature based on the first query matrix, the second key matrix and the second value matrix; The audio fusion feature is determined based on the second query matrix, the first key matrix and the first value matrix.
7. The method according to any one of claims 1 to 6, wherein: The determining the authenticity of the first multimedia data includes: Splicing the video fusion feature and the audio fusion feature to obtain a spliced fusion feature; Based on the splicing and fusion features, the authenticity of the first multimedia data is determined through a trained classification network.
8. The method of claim 7, wherein: The determining the authenticity of the first multimedia data includes: Based on the splicing and fusion features, determining a first authenticity result corresponding to the first video and a second authenticity result corresponding to the first audio through a trained classification network; The authenticity of the first multimedia data is determined based on the first authenticity result and the second authenticity result.
9. The method of claim 7, wherein: The classification network is trained based on a number of sample multimedia data and their corresponding label data, wherein the label data includes a first label indicating that the corresponding sample multimedia data is real video and real audio, a second label indicating that the corresponding sample multimedia data is real video and fake audio, a third label indicating that the corresponding sample multimedia data is fake video and real audio, and a fourth label indicating that the corresponding sample multimedia data is fake video and fake audio.
10. A method for identifying the authenticity of multimedia data, comprising: An acquisition module, configured to acquire first multimedia data to be authenticated, wherein the first multimedia data includes a first video and a first audio corresponding thereto; A first obtaining module is configured to extract features from the first video to obtain first video features; A second obtaining module is configured to extract features from the first audio to obtain first audio features; A third obtaining module is configured to adopt a cross attention mechanism to obtain a video fusion feature and an audio fusion feature based on the first video feature and the first audio feature; The determination module is configured to determine the authenticity of the first multimedia data based on the video fusion feature and the audio fusion feature.
11. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 9 is implemented.