Multimedia resource forgery detection method and device
By fusing images and audio information to build a multimedia resource forgery detection method, the problem of low detection accuracy in the prior art is solved, higher detection accuracy and interpretability are achieved, and detailed forgery detection reports are generated.
Patent Information
- Application Number
- CN202510597705.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-19
AI Technical Summary
The existing deep forgery detection technology has low detection accuracy and is difficult to provide intuitive and interpretable detection basis, making it difficult for users to trust the detection results.
By fusing image information and audio information, a fusion feature set of multimedia resources is constructed, a target forged features is calculated, a fine-tuned data set is constructed to fine-tune the resource detection model, and a fine-tuned model is used for forged detection to generate a detailed forged detection report.
It improves the accuracy of multimedia resource forgery detection, enhances the interpretability and credibility of detection, and provides a detailed basis for forgery detection.
Smart Images

Figure CN120508978A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and more particularly to a method and apparatus for detecting forgery of multimedia resources. Background Art
[0002] Deepfake technology is a technique that automatically edits or synthesizes fake content based on artificial intelligence methods such as deep learning. In recent years, the development of deep learning information technology in the field of computer vision has made deepfake technology more intelligent and streamlined, significantly reducing the cost and barrier to entry for fabrication. Furthermore, relying on powerful intelligent algorithms and continuously improving deepfake models, the generated fake visuals have achieved realistic scenes and are difficult to distinguish between real and fake. However, malicious deepfake visuals, particularly deepfake facial videos targeting public figures, have spread rapidly on social media and content sharing platforms in recent years, attracting widespread public attention. The generation and dissemination of these audio and video data has seriously eroded social trust and disrupted work and life. Current deepfake detection technology has limitations, resulting in inaccurate detection results. Therefore, improving the accuracy of deepfake detection is an urgent issue. Summary of the Invention
[0003] In view of this, embodiments of this specification provide a method for detecting forgery of multimedia resources. One or more embodiments of this specification also relate to a device for detecting forgery of multimedia resources, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.
[0004] According to a first aspect of an embodiment of this specification, a method for detecting forgery of multimedia resources is provided, comprising:
[0005] Determining a fusion feature set of multimedia resources through a resource detection model, wherein each fusion feature in the fusion feature set includes image information and audio information;
[0006] Calculating forgery weight information corresponding to each fused feature in the fused feature set, and screening target forgery features from the fused features according to the forgery weight information;
[0007] Constructing a fine-tuning dataset based on the target forgery features, and using the fine-tuning dataset to fine-tune the resource detection model;
[0008] Each fusion feature in the fusion feature set is subjected to forgery detection by using the fine-tuned resource detection model to obtain a forgery detection report corresponding to the multimedia resource.
[0009] According to a second aspect of the embodiments of this specification, a device for detecting forgery of multimedia resources is provided, comprising:
[0010] a determination module configured to determine a fusion feature set of a multimedia resource through a resource detection model, wherein each fusion feature in the fusion feature set includes image information and audio information;
[0011] a calculation module configured to calculate forgery weight information corresponding to each fused feature in the fused feature set, and to filter a target forgery feature from the fused features according to the forgery weight information;
[0012] a fine-tuning module configured to construct a fine-tuning dataset according to the target forgery features, and to fine-tune the resource detection model using the fine-tuning dataset;
[0013] The detection module is configured to perform forgery detection on each fusion feature in the fusion feature set by using the fine-tuned resource detection model to obtain a forgery detection report corresponding to the multimedia resource.
[0014] According to a third aspect of an embodiment of this specification, a computing device is provided, including:
[0015] memory and processor;
[0016] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the forgery detection method of multimedia resources are implemented.
[0017] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the forgery detection method of multimedia resources are implemented.
[0018] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, comprising a computer program or instructions, which, when executed by a processor, implement the steps of the forgery detection method for multimedia resources.
[0019] One embodiment of the present specification realizes the extraction of fusion features of multimedia resources through a resource detection model, and each fusion feature includes image information and audio information, so that when the multimedia resources are subsequently subjected to forgery detection based on the fusion features, the image information and audio information can be fully utilized to improve the accuracy of forgery detection. After determining the fusion feature set, the target forgery features are screened out using the calculated forgery weight information of each fusion feature, and a fine-tuning data set is constructed for fine-tuning the resource detection model, so as to improve the sensitivity and judgment accuracy of the fine-tuned resource detection model to high-weight forgery features. The fine-tuned resource detection model is used to perform forgery detection on each fusion feature in the fusion feature set to obtain a forgery detection report for the multimedia resources, and the interpretability and credibility of the detection basis are improved based on the forgery detection report. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flowchart of a method for detecting forgery of multimedia resources provided by one embodiment of this specification;
[0021] Figure 2 This is a flowchart of a processing process of a method for detecting forgery of multimedia resources provided by one embodiment of this specification;
[0022] Figure 3 This is a schematic structural diagram of a multimedia resource forgery detection device provided by one embodiment of this specification;
[0023] Figure 4 This is a structural block diagram of a computing device provided by one embodiment of this specification. DETAILED DESCRIPTION
[0024] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0025] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0026] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0027] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0028] First, the terms involved in one or more embodiments of this specification are explained.
[0029] Deepfake technology: Deepfake technology is a method that uses artificial intelligence, particularly deep learning technology, to create highly realistic fake images, videos, or audio content. This technology can generate fake content that looks very real, making it difficult for ordinary viewers to distinguish the authenticity.
[0030] Multimodal Large Language Models: Multimodal Large Language Models (MLLMs) are large deep learning models that can process and understand data from multiple modalities, such as text, images, and audio. By integrating multiple types of data, these models aim to provide more comprehensive and accurate understanding and generation capabilities.
[0031] With the rapid development of deep fake technology, images and audio generated by AI (Artificial Intelligence) are becoming more and more realistic, posing serious challenges to information security, privacy protection, and social ethics. Current deep fake detection methods can usually only give a true or false judgment, but lack intuitive and explainable detection basis, making it difficult for users to trust the detection results. In addition, there are still technical difficulties in the fusion processing and interpretation of multimodal data (images, audio). Some methods attempt to improve detection accuracy by fusing image and audio information. However, due to the complexity of multimodal information processing, the judgment basis of such methods is often unclear, and it is difficult to provide a user-understandable description of the fake characteristics.
[0032] Based on this, this specification provides a method for detecting forgery of multimedia resources. This specification also involves a device for detecting forgery of multimedia resources, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0033] See also Figure 1 , Figure 1 A flowchart of a method for detecting forgery of multimedia resources provided according to an embodiment of this specification is shown, which specifically includes the following steps.
[0034] Step 102: Determine a fusion feature set of the multimedia resource through a resource detection model, wherein each fusion feature in the fusion feature set includes image information and audio information.
[0035] The resource detection model can be understood as a large multimodal model used to detect deepfakes in multimedia resources. Because multimedia resources include data resources such as images and audio, such as a video, a large multimodal model is required to detect deepfakes in these resources.
[0036] In practical applications, when a large multimodal model processes complex forged content such as images and audio, in order to improve the model's processing capabilities and enable the model to process image and audio data simultaneously, it is necessary to perform feature fusion on the image data and audio data included in the multimedia resources to obtain a fusion feature set of the multimedia resources. The fusion feature set includes fusion features corresponding to the multimedia resources, and each fusion feature contains image information and audio information. That is, the subsequent visual and auditory characteristics can be fully utilized in deep forgery detection to improve the accuracy of subsequent detection.
[0037] In specific implementations, the fused features in the fused feature set can be obtained by fusing image and audio features from multimedia resources. These features are then extracted from the multimedia resources using a resource detection model. After obtaining the fused feature set, the resource detection model can then be used to perform deepfake detection on the multimedia resources based on the fused features in the fused feature set. This can determine whether the multimedia resources contain forgeries and provide corresponding confidence scores, explanations, and other test results.
[0038] Furthermore, since the resource detection model needs to extract the fusion features of multimedia resources, data preprocessing and data separation operations need to be performed on the multimedia resources before extraction. Specifically, before determining the fusion feature set of multimedia resources through the resource detection model, the method also includes: performing data separation on the multimedia resources to obtain image data and audio data corresponding to the multimedia resources; and inputting the image data and the audio data into the resource detection model.
[0039] The image data may be understood as image data separated from the multimedia resource, and the audio data may be understood as audio data separated from the multimedia resource. By performing data separation on the multimedia resource, the image data and audio data of the multimedia resource may be obtained.
[0040] In practical applications, the modal separation module can be used to separate multimedia resources into image data and audio data. The image data can then be processed through the image channel, and the audio data through the audio channel. The main task of the modal separation module is to separate the different modal data within a multimedia resource, such as a video file, for subsequent processing. The modal separation module can achieve data separation by extracting image frame data from a video file frame by frame using open-source tools for image separation, and extracting audio track data from a video file using open-source tools for audio separation.
[0041] During specific implementation, after extracting image data and audio data from multimedia resources, data preprocessing can be performed on the image data and audio data to ensure data validity. Data preprocessing can include data enhancement and noise filtering. Data enhancement can include cropping, flipping, rotation, scaling, color jittering, and adding noise. Generating additional training samples through data enhancement can increase the generalization ability of the model and prevent overfitting. Noise filtering can include Gaussian blurring, median filtering, bilateral filtering, and other methods. Noise filtering can remove noise from the image to improve the quality of feature extraction. After obtaining image data and audio data, the image data can be input into the resource detection model through the image channel, and the audio data can be input into the resource detection model through the audio channel. After obtaining the image data and audio data, the resource detection model will extract image features from the image data and audio features from the audio data, respectively, for the subsequent generation of fusion features.
[0042] In a specific embodiment of this specification, the multimedia resource is a video file. The video file is subjected to data separation to obtain image data, i.e., image frames, and audio data, i.e., audio tracks. The image data and audio data are then preprocessed and input into the resource detection model.
[0043] Based on this, by performing data separation on multimedia resources, we can effectively separate image data and audio data from them, making full use of the characteristics of each modality and improving subsequent detection performance and accuracy. During the separation process, data is preprocessed to ensure its validity and reliability, thereby enhancing the model's detection and interpretation capabilities.
[0044] Furthermore, in order to generate a fusion feature set of multimedia resources, it is necessary to extract features from image data and audio data respectively. Specifically, the fusion feature set of multimedia resources is determined through a resource detection model, including: extracting features from the image data and the audio data respectively through a resource detection model to obtain image features corresponding to the image data and audio features corresponding to the audio data; fusing the image features and the audio features to obtain multiple fusion features; and constructing the fusion feature set of the multimedia resources based on the multiple fusion features.
[0045] The resource detection model has the ability to extract features, including image features from image data and audio features from audio data. Image and audio features are key information extracted from these data and used for subsequent analysis, fusion, and ultimately forgery detection. Specifically, image features refer to information extracted from image data that aids in identification and classification. These features can describe image content, structure, texture, and other characteristics. Audio features refer to information extracted from audio data that aids in identification and classification. These features can describe characteristics such as the pitch, rhythm, and background noise of a sound.
[0046] During specific implementation, convolutional neural networks and other technologies can be used to extract feature representations containing idle information, namely image features, from image data. Mel-frequency cepstral coefficients, spectrograms and other technologies can be used to extract time-frequency domain features, namely audio features, from audio data. It should be noted that in this embodiment, there is no specific restriction on the method of extracting image features and audio features. It is only necessary to be able to extract image features from image data and audio features from audio data. After the image features and audio features are extracted, since the image features and audio features are feature representations extracted from different dimensions, in order to fully utilize the image features and audio features in the future, the image features and audio features can be fused to form a unified multimodal representation, namely fused features. The fused features contain the image information of the image features and the audio information of the audio features, so that when performing forgery detection based on the fused features, the relevant information of the visual dimension and the auditory dimension can be fully considered.
[0047] In practical applications, different feature fusion strategies can be used to fuse image features and audio features, such as feature splicing, common representation learning, attention mechanism and other strategies. In the embodiments of this specification, common representation learning is used as an example to illustrate the fusion of image features and audio features. The core idea of common representation learning is to find a shared representation space so that features of different modalities can be directly compared, aligned and fused in the shared space. The general steps of common representation learning can be divided into mapping image features and audio features to a shared embedding space through contrastive learning, generative adversarial networks and other methods, and then using splicing, weighted averaging and other methods to fuse the image features and audio features in the same space to generate corresponding fusion features. It should be noted that there can be multiple image features and audio features, and there is a corresponding relationship between the two, so that multiple fusion features can be generated.
[0048] In a specific embodiment of the present specification, feature extraction is performed on image data and audio data respectively through a resource detection model to obtain image features corresponding to the image data and audio features corresponding to the audio data, and the image data and audio features are fused to obtain multiple fused features. The multiple fused features are combined into a fused feature set of multimedia resources, so that each fused feature in the fused feature set contains image information and audio information.
[0049] Based on this, by fusing image features with audio features, the fused features have information in both visual and auditory dimensions. In forgery detection, relying solely on images or audio may not capture all forgery clues, while combining the two can more accurately identify forgery behavior. Therefore, forgery detection based on fused features can greatly improve the detection accuracy; and when detecting through multimodal fusion features, if one modality is damaged, detection can rely on another modality, thereby improving the robustness of the overall system.
[0050] Step 104: Calculate forgery weight information corresponding to each fused feature in the fused feature set, and select a target forgery feature from the fused features according to the forgery weight information.
[0051] The forgery weight information corresponding to the fused feature can be understood as the forgery correlation corresponding to the fused feature. The forgery weight information can be a forgery weight value. A higher forgery weight value indicates that the fused feature is closer to the forgery feature. The forgery weight information can serve as an important basis for the model to determine whether a multimedia resource is forged.
[0052] In practical applications, forgery weight information refers to the measure of forgery relevance corresponding to each fused feature. Specifically, it is a numerical value, namely the forgery weight value, which reflects the degree of similarity between the feature and the known forgery pattern. The higher the forgery weight value, the closer the feature is to the typical forgery feature, and vice versa, the closer it is to the real content. Based on the forgery weight information, the target forgery feature can be screened out from the fused features. The target forgery feature is the fused feature with a high weight in the fused features, that is, the fused features with higher forgery weight values are screened out from the fused features and used as the target forgery feature. The target forgery feature is used to subsequently construct a fine-tuning dataset to fine-tune the resource detection model, improve the model's sensitivity and judgment accuracy to high-weight forgery features, and optimize the quality of the generated explanatory text.
[0053] Furthermore, in order to accurately calculate the forgery weight information of each fusion feature, it is necessary to determine the forgery clue information. Specifically, the forgery weight information corresponding to each fusion feature in the fusion feature set is calculated, including: determining the forgery clue information corresponding to each fusion feature based on a preset feature selection algorithm; calculating the forgery weight value corresponding to each forgery clue information, and determining the forgery weight information corresponding to each fusion feature according to the forgery weight value corresponding to each forgery clue information.
[0054] Among them, the preset feature selection algorithm can be understood as an algorithm for performing forgery correlation analysis on fused features. The preset feature selection algorithm is used to determine forgery clue information associated with forgery behavior in the fused features. The preset feature selection algorithm can include filtering, packaging, embedding, and domain knowledge-driven algorithms. Through the preset feature selection algorithm, the parts of the fused features associated with forgery behavior can be screened and quantified, and which features are key forgery clue information can be determined. Forgery clue information can be understood as information in the fused features that can reflect patterns or anomalies in forgery behavior. Forgery clue information can come from a single modality such as images or audio, or from cross-modal interactions such as the alignment of vision and hearing. Through the preset feature selection algorithm, forgery clue information can be effectively identified, and the forgery weight value of each fused feature can be calculated based on the forgery clue information, thereby providing key support for subsequent forgery detection.
[0055] In practical applications, a preset feature selection algorithm can be used to determine the forgery clue information corresponding to each fused feature. The forgery clue information can reflect whether the fused feature is associated with forgery behavior, such as clues such as inconsistent lighting and changes in background noise. The forgery weight value of each fused feature can be calculated through the forgery clue information. The higher the forgery weight value, the closer the feature is to the typical forgery feature, and vice versa, the closer it is to the real content.
[0056] During specific implementation, for the forgery clues of the fusion features, appropriate statistical methods or machine learning models can be used to calculate their forgery weight values. For example, for inconsistent lighting, the brightness differences in different areas can be calculated; for changes in background noise, the energy changes in the audio spectrum can be analyzed. For each fusion feature, the forgery weight information of the fusion feature can be calculated based on its corresponding forgery clues and its corresponding forgery values. When a fusion feature contains multiple forgery clues, the forgery value corresponding to each forgery clue can be calculated first, and then the weighted average or attention mechanism can be used to calculate the forgery weight value of the fusion feature. Specifically, it can be judged whether the fusion feature has forgery behavior based on the forgery weight value. For example, a forgery threshold is set. When the forgery weight value is greater than the forgery threshold, it means that the fusion feature is forged, otherwise the fusion feature is real.
[0057] Based on this, by calculating the forgery weight value of each fused feature, not only can the accuracy of forgery detection be improved, but also corresponding explanatory text can be generated based on the forgery clue information to explain why the fused feature is analyzed as a forgery feature, thereby enhancing the system's explanatory ability and making it easier for users to understand and trust the detection results.
[0058] Furthermore, in order to be able to filter out the target forged features in the fused features based on the forged weight information, it is also necessary to sort the fused features based on the forged weight information. Specifically, the target forged features are filtered out in the fused features according to the forged weight information, including: sorting the importance of each fused feature according to the forged weight information corresponding to each fused feature to obtain a forged feature sorting list; and selecting the target forged feature in the forged feature sorting list.
[0059] Among them, the forgery weight value of the fused feature can be determined based on the forgery weight value of the forgery clue information corresponding to each fused feature. Therefore, the forgery weight information corresponding to each fused feature can be used to sort the importance of each fused feature in the fused feature set. Importance sorting is to sort the fused features according to the forgery weight value. Since the forgery weight value reflects whether the fused feature is a forgery feature, the forgery weight value of each fused feature can reflect the importance of determining whether the multimedia resource is forged. After sorting the importance of the fused features, a forgery feature sorting list can be obtained. The forgery feature sorting category is a list arranged in descending order according to the forgery feature value of the fused feature. According to the forgery feature sorting list, the fused features with higher forgery feature values can be clearly screened out as target forgery features.
[0060] In practical applications, the target forged features can be selected from the forged feature ranking list using a preset forged threshold. When the forged weight value of the fused feature is greater than the forged threshold, it can be used as the target forged feature. Alternatively, a preset number of fused features can be selected from the forged feature ranking list as target forged features. For example, if the forged feature ranking list includes 10 fused features and the preset number of selections is 3, the first 3 fused features from the forged feature ranking list are selected as target forged features. The screened target forged features are identified as features in multimedia resources that are closer to forged behavior. Subsequently, a fine-tuning dataset can be constructed based on the target forged features, and the fine-tuning dataset can be used to fine-tune the resource detection model. Through reinforcement learning and supervisory signal injection, the model's sensitivity to high-weight forged features and judgment accuracy can be improved, while optimizing the quality of the generated explanatory text.
[0061] In a specific embodiment of the present specification, after calculating the forgery weight value of each fused feature, the forgery weight value of each fused feature is used to sort the importance, and a forgery feature ranking list is generated in descending order. A target forgery feature is selected from the forgery feature ranking list according to a preset selection method.
[0062] Based on this, by sorting the fused features according to the forgery weight value, the target forgery features that are most likely to indicate forgery behavior can be effectively identified. Subsequently, the target forgery features are used to fine-tune the resource detection model, further improving the model's sensitivity and judgment ability to forgery behavior.
[0063] Step 106: construct a fine-tuning dataset based on the target forgery features, and use the fine-tuning dataset to fine-tune the resource detection model.
[0064] The fine-tuning dataset can be understood as a dataset used to fine-tune the resource detection model. The fine-tuning dataset can be constructed from high-weighted forged feature samples. In practical applications, these high-weighted forged feature samples can include both target forged features and historical forged features. This means that the fine-tuning dataset is constructed using both historical forged features from historical data and the currently selected target forged features. The resource detection model is then fine-tuned based on the fine-tuning dataset, further improving the model's sensitivity to high-weighted forged features and its accuracy.
[0065] In practice, a comprehensive fine-tuning dataset can be constructed by combining target forged features with historical forged features. This dataset contains sufficient positive samples (i.e., genuine forged features) and negative samples (i.e., genuine non-forged features) for the model to learn the difference between the two. When fine-tuning the resource detection model using the fine-tuning dataset, transfer learning, incremental learning, or hyperparameter adjustment can be used to fine-tune the model.
[0066] Based on this, by constructing a fine-tuning dataset based on the target forgery features and using the fine-tuning dataset to fine-tune the resource detection model, the sensitivity and judgment accuracy of the resource detection model to high-weight forgery features can be effectively improved.
[0067] Furthermore, constructing a fine-tuning dataset based on the target forgery features includes: determining historical forgery features and historical forgery clue information corresponding to the historical forgery features; and constructing a fine-tuning dataset based on the historical forgery features, the historical forgery clue information, the target forgery features, and the target forgery clue information corresponding to the target forgery features.
[0068] Among them, historical forgery features can be understood as the features of known forgery samples extracted from past accumulated data, and historical forgery clue information can be understood as the information of specific forgery patterns or anomalies corresponding to historical forgery features; target forgery features are high-weight forgery features screened out from the current batch of data, and target forgery clue information is the information of specific forgery models or anomalies corresponding to the target forgery features.
[0069] In practice, the constructed fine-tuning dataset can be divided into a training set and a validation set. The training set is used for model fine-tuning, while the validation set is used to evaluate model performance and avoid overfitting. In practical applications, historical forgery features and their corresponding historical forgery clues can be extracted from the historical dataset, and target forgery features and their corresponding target forgery clues can be screened from the current batch of data. The fine-tuning dataset is constructed based on the historical forgery features, historical forgery clues, target forgery features, and the corresponding target forgery clues.
[0070] In a specific embodiment of this specification, the historical forgery features include "forgery feature A1," and the historical forgery clue information is "lighting consistency index, inconsistent lighting, forgery." The target forgery features include "forgery feature b1," and the target forgery clue information is "normal lighting distribution, no anomalies, authentic." A fine-tuning dataset is constructed based on the historical forgery features, historical forgery clue information, target forgery features, and the target forgery clue information corresponding to the target forgery features.
[0071] Based on this, a fine-tuning dataset is constructed by effectively integrating historical and target forgery features. This effectively leverages historical experience and the latest forgery patterns in current data, enabling the resource detection model to continuously evolve and address increasingly complex forgery challenges. The fine-tuned model not only improves sensitivity and accuracy for high-weighted forgery features but also generates more detailed explanations to help users understand why a piece of content is considered forged. This process significantly enhances the robustness and interpretability of the system, providing strong support for practical applications.
[0072] Step 108: Perform forgery detection on each fusion feature in the fusion feature set using the fine-tuned resource detection model to obtain a forgery detection report corresponding to the multimedia resource.
[0073] After fine-tuning the model, a fine-tuned resource detection model can be obtained. The fine-tuned resource detection model can then be used to detect forgeries of multimedia resources based on fused features, thereby obtaining a forgery detection report corresponding to the multimedia resources. The forgery detection report can include the detection result, i.e., whether the multimedia resource is forged; the multimedia resource forgery basis, i.e., how the model determined that the multimedia resource was forged; and the contribution of each fused feature in forgery detection, i.e., an analysis based on the forgery weight value of each fused feature.
[0074] Furthermore, in order to avoid the situation where the model ignores low-weight features, after the model outputs the detection results, it can also be combined with external detection results for fusion output. Specifically, each fused feature in the fused feature set is detected for forgery by the fine-tuned resource detection model to obtain a forgery detection report corresponding to the multimedia resource, including: performing forgery detection on each fused feature in the fused feature set by the fine-tuned resource detection model to obtain a first detection result corresponding to the multimedia resource; performing forgery detection on the multimedia resource by a deep forgery detector to obtain a second detection result corresponding to the multimedia resource; and generating a forgery detection report corresponding to the multimedia resource based on the first detection result and the second detection result.
[0075] The first detection result can be understood as the detection result output by the fine-tuned resource detection model after performing forgery detection on the multimedia resource based on the fused features. The first detection result can include the forgery probability or forgery weight value of each fused feature, as well as the forgery judgment of the multimedia resource as a whole (such as "forgery" or "real"). Since the fine-tuned model pays more attention to high-weight forgery features, it may not be sensitive enough to low-weight features. For this reason, it is necessary to introduce external detection results, namely the second detection result, for fusion. The second detection result can be understood as the detection result output after the multimedia resource is forged by the deep forgery detector. The deep forgery detector is a tool or model specifically used to detect and identify deep forgery content. The deep forgery detector can detect specific modalities, such as only detecting images or only detecting audio. The second detection result can include the forgery judgment of the multimedia resource as a whole (such as "forgery" or "real"), as well as some specific forgery clue information (such as abnormal areas in video frames, unnatural clips in audio, etc.). The deep forgery detector can usually capture some global forgery patterns, but it may not be able to analyze individual fused features as deeply as the fine-tuned resource detection model.
[0076] In practical applications, a fine-tuned resource detection model performs forgery detection based on fused features, yielding a first detection result for multimedia resources. Deepfake detection then performs forgery detection on multimedia resources, yielding a second detection result for the multimedia resource. By fusing the first and second detection results, a forgery detection report for the multimedia resource is generated.
[0077] In specific implementations, the first and second detection results can be input into a fusion module, which then analyzes the detection results and detailed explanations of the multimedia resource based on the two detection results. This analysis involves performing a comprehensive analysis based on the corresponding detection data in the first and second detection results, such as forgery weights and forgery clues, to produce a more accurate and comprehensive forgery detection report.
[0078] Based on this, by integrating an external, dedicated deepfake detector and combining its results with the output of a fine-tuned resource detection model, we can effectively improve overall detection and explanation capabilities, especially for low-ranked or under-detectable forgery features. This approach not only improves the accuracy and robustness of detection results, but also enhances the interpretability of the system, providing users with a detailed and trustworthy forgery detection report.
[0079] Furthermore, a forgery detection report corresponding to the multimedia resource is generated based on the first detection result and the second detection result, including: determining the feature contribution information, detection conclusion information and detection annotation information of the multimedia resource based on the first detection result and the second detection result; and generating a forgery detection report corresponding to the multimedia resource based on the feature contribution information, the detection conclusion information and the detection annotation information.
[0080] Among them, the feature contribution can be understood as the contribution of each fused feature in the multimedia resource, that is, the contribution of the forgery analysis based on the forgery weight value of each fused feature. The detection result information can be understood as the forgery judgment result of the multimedia resource as a whole (such as "forgery" or "real"). The detection annotation information can be understood as the annotation information of forgery patterns, abnormal points, forgery clues, etc. in the multimedia resource, such as the lighting anomaly of a certain frame image in the video, the background noise pattern change in the video audio, and other annotation information. The forgery annotation information may include the information generated by the first detection result and the second detection result, or it may be manually annotated. According to the feature contribution information, the detection conclusion information and the detection annotation information, a forgery detection report corresponding to the multimedia resource can be generated.
[0081] In practical applications, the contribution of a fused feature to forgery analysis can be determined based on its forgery weight, for example, "Feature A: Forgery Weight = 0.95, Contribution = High (Inconsistent Lighting)." Detection conclusion information can be derived from the first and second detection results, using methods such as weighted averaging, voting, or an attention mechanism. This information is then compared to a preset forgery probability threshold to determine whether the multimedia resource is "forged" or "authentic." Detection annotation information is specific annotation information for forgery patterns, anomalies, and forgery clues within the multimedia resource. This information can include lighting anomalies in video frames and background noise changes in audio. Detection annotation information can include local forgery clues and their locations provided by the first detection result, and global forgery patterns and their locations provided by the second detection result. By integrating feature contribution, detection conclusion information, and detection annotation information, a forgery detection report corresponding to the multimedia resource is generated. This report allows users to gain a detailed understanding of the detection process and basis, enhancing their confidence in the detection results.
[0082] Based on this, by combining the first and second detection results, a comprehensive and accurate forgery detection report for multimedia resources can be generated. This report not only provides detailed feature contribution information, detection conclusion information, and detection annotation information, but also enhances the system's interpretability and credibility, helping users better understand and respond to forged content.
[0083] Furthermore, after obtaining the forgery detection report corresponding to the multimedia resource, the method also includes: determining the marking information of the multimedia resource in response to an update operation on the forgery detection report; updating the forgery detection report based on the marking information to obtain the updated forgery detection report corresponding to the multimedia resource.
[0084] The marking information for multimedia resources can be understood as information manually marked for outliers within a multimedia resource. For example, if a user marks a video clip as forged, and believes that the scene in the current video clip contains forgery, the user marks the location and generates the corresponding marking information. This marking information can then be updated in the forgery detection report to obtain an updated forgery detection report.
[0085] In practical applications, a user interface can be provided to easily review the forgery detection report and mark suspicious areas. Marking methods include selecting a specific time point or frame to add a forgery mark, entering a detailed description of the forgery clue, and uploading related supporting evidence. Subsequently, based on the user's selection, the marked information can be directly added to the forgery detection report. Alternatively, the user-added marking information can be used to perform a specific forgery detection verification to verify that the user's markings are correct.
[0086] Based on this, updating the forgery detection report based on manually labeled information can significantly improve the accuracy and credibility of the system. This approach not only makes full use of user feedback, but also enhances the interactivity and transparency of the system.
[0087] This specification provides a method for detecting forgery of multimedia resources, including determining a fused feature set of the multimedia resource through a resource detection model, wherein each fused feature in the fused feature set includes image information and audio information; calculating forgery weight information corresponding to each fused feature in the fused feature set, and screening target forgery features from the fused features based on the forgery weight information; constructing a fine-tuning dataset based on the target forgery features, and using the fine-tuning dataset to fine-tune the resource detection model; performing forgery detection on each fused feature in the fused feature set using the fine-tuned resource detection model, and obtaining a forgery detection report corresponding to the multimedia resource. The method achieves the extraction of fused features of the multimedia resource through the resource detection model, wherein each fused feature includes image information and audio information, so that when the multimedia resource is subsequently forged based on the fused features, the image and audio information can be fully utilized, thereby improving the accuracy of forgery detection. After determining the fused feature set, the calculated forgery weight information of each fused feature is used to screen out the target forgery features, and a fine-tuning dataset is constructed for fine-tuning the resource detection model, thereby improving the sensitivity and judgment accuracy of the fine-tuned resource detection model to high-weight forgery features. The fine-tuned resource detection model is used to perform forgery detection on each fusion feature in the fusion feature set to obtain a forgery detection report for multimedia resources. The interpretability and credibility of the detection basis are improved based on the forgery detection report.
[0088] The following combined Figure 2 , taking the application of the multimedia resource forgery detection method provided in this specification in video detection as an example, the multimedia resource forgery detection method is further described. Figure 2 A flowchart of a processing process of a method for detecting forgery of multimedia resources provided by an embodiment of this specification is shown, which specifically includes the following steps.
[0089] Step 202: Separate the multimedia resources to obtain image data and audio data corresponding to the multimedia resources, and input the image data and audio data into a resource detection model.
[0090] In one implementation, a multimedia resource forgery detection method is applied to social platforms. Specifically, the social platform wishes to perform deepfake detection on user-uploaded multimedia content, such as videos and audio, automatically detecting and flagging such forgeries to protect the interests of platform users. The multimedia resource can be a video uploaded by a user to the social platform. By performing data separation on the video, the image frame sequence (image data) and the audio track (audio data) are extracted. The image and audio data are then fed into a resource detection model.
[0091] Step 204: Perform feature extraction on the image data and the audio data respectively using the resource detection model to obtain image features corresponding to the image data and audio features corresponding to the audio data.
[0092] In one achievable manner, a convolutional neural network is used to extract image features from an image frame sequence, and an audio extraction model is used to extract audio features from an audio track.
[0093] Step 206: Fusing the image features and the audio features to obtain multiple fusion features, and constructing a fusion feature set of the multimedia resource based on the multiple fusion features.
[0094] In one achievable approach, the image features and audio features are mapped into the same embedding space and fused to generate multiple fused features. A fused feature set corresponding to the video segment is constructed based on the multiple fused features.
[0095] Step 208: Based on a preset feature selection algorithm, determine the forgery clue information corresponding to each fusion feature.
[0096] In one achievable manner, a preset feature selection algorithm is used to determine the forgery clue information corresponding to each fused feature, that is, the information associated with the forgery behavior in each fused feature.
[0097] Step 210: Calculate the forgery weight value corresponding to each forgery clue information, and determine the forgery weight information corresponding to each fusion feature according to the forgery weight value corresponding to each forgery clue information.
[0098] In one achievable method, the forgery weight value corresponding to each forgery clue information is calculated, and the forgery weight information is determined according to the forgery weight value of the forgery clue corresponding to each fusion feature, for example, the forgery weight information of each fusion feature is determined by a weighted average method.
[0099] Step 212: sorting the importance of each fused feature according to the forged weight information corresponding to each fused feature, obtaining a forged feature sorting list, and selecting a target forged feature from the forged feature sorting list.
[0100] In one feasible method, the importance of the fused features is sorted in descending order according to the forgery weight value of each fused feature to obtain a forged feature sorting list, and the target forged feature is selected from the forged feature sorting list according to a preset selection method such as threshold comparison.
[0101] Step 214: Construct a fine-tuning dataset based on the target forgery features, and use the fine-tuning dataset to fine-tune the resource detection model.
[0102] In one achievable manner, historical forgery features and historical forgery clue information corresponding to the historical forgery features are determined, and a fine-tuning dataset is constructed based on the historical forgery features, the historical forgery clue information, the target forgery features, and the target forgery clue information corresponding to the target forgery features.
[0103] Step 216: Perform forgery detection on each fused feature in the fused feature set using the fine-tuned resource detection model to obtain a forgery detection report corresponding to the multimedia resource.
[0104] In one feasible method, each fusion feature in the fusion feature set is detected for forgery through a fine-tuned resource detection model to obtain a first detection result corresponding to the multimedia resource, and the multimedia resource is detected for forgery through a deep forgery detector to obtain a second detection result corresponding to the multimedia resource, and the feature contribution information, detection conclusion information and detection annotation information of the multimedia resource are determined based on the first detection result and the second detection result, and a forgery detection report corresponding to the multimedia resource is generated based on the feature contribution information, detection conclusion information and detection annotation information, and if it is determined that there is forgery in the video based on the forgery detection report, the social platform can remind the user to delete the video or remove the video, and can also add an AI synthetic picture of the video in the video browsing page to remind other users to identify it.
[0105] The method for detecting forgery of multimedia resources provided in this specification realizes the extraction of fusion features of multimedia resources through a resource detection model, and each fusion feature includes image information and audio information, so that when the multimedia resources are subsequently subjected to forgery detection based on the fusion features, the image information and audio information can be fully utilized to improve the accuracy of forgery detection. After determining the fusion feature set, the target forgery features are screened out using the calculated forgery weight information of each fusion feature, and a fine-tuning data set is constructed for fine-tuning the resource detection model, so as to improve the sensitivity and judgment accuracy of the fine-tuned resource detection model to high-weight forgery features. The fine-tuned resource detection model is used to perform forgery detection on each fusion feature in the fusion feature set to obtain a forgery detection report for the multimedia resource, and the interpretability and credibility of the detection basis are improved based on the forgery detection report.
[0106] Corresponding to the above method embodiment, this specification also provides an embodiment of a device for detecting forgery of multimedia resources. Figure 3 FIG. 1 shows a schematic diagram of a device for detecting forgery of multimedia resources provided by an embodiment of this specification. Figure 3 As shown, the device includes:
[0107] The determination module 302 is configured to determine a fusion feature set of a multimedia resource through a resource detection model, wherein each fusion feature in the fusion feature set includes image information and audio information;
[0108] The calculation module 304 is configured to calculate forgery weight information corresponding to each fused feature in the fused feature set, and select a target forgery feature from the fused features according to the forgery weight information;
[0109] A fine-tuning module 306 is configured to construct a fine-tuning dataset according to the target forgery features, and use the fine-tuning dataset to fine-tune the resource detection model;
[0110] The detection module 308 is configured to perform forgery detection on each fused feature in the fused feature set using the fine-tuned resource detection model to obtain a forgery detection report corresponding to the multimedia resource.
[0111] Optionally, the device further includes a separation module configured to perform data separation on the multimedia resource to obtain image data and audio data corresponding to the multimedia resource; and input the image data and the audio data into a resource detection model.
[0112] Optionally, the determination module 302 is further configured to perform feature extraction on the image data and the audio data respectively through a resource detection model to obtain image features corresponding to the image data and audio features corresponding to the audio data; fuse the image features and the audio features to obtain multiple fusion features; and construct a fusion feature set of the multimedia resource based on the multiple fusion features.
[0113] Optionally, the calculation module 304 is further configured to determine the forged clue information corresponding to each fused feature based on a preset feature selection algorithm; calculate the forged weight value corresponding to each forged clue information, and determine the forged weight information corresponding to each fused feature according to the forged weight value corresponding to each forged clue information.
[0114] Optionally, the calculation module 304 is further configured to sort the importance of each fused feature according to the forgery weight information corresponding to each fused feature to obtain a forgery feature sorting list; and select a target forgery feature from the forgery feature sorting list.
[0115] Optionally, the fine-tuning module 306 is further configured to determine historical forgery features and historical forgery clue information corresponding to the historical forgery features; and construct a fine-tuning dataset based on the historical forgery features, the historical forgery clue information, the target forgery features, and the target forgery clue information corresponding to the target forgery features.
[0116] Optionally, the detection module 308 is further configured to perform forgery detection on each fusion feature in the fusion feature set through the fine-tuned resource detection model to obtain a first detection result corresponding to the multimedia resource; perform forgery detection on the multimedia resource through a deep forgery detector to obtain a second detection result corresponding to the multimedia resource; and generate a forgery detection report corresponding to the multimedia resource based on the first detection result and the second detection result.
[0117] Optionally, the detection module 308 is further configured to determine the feature contribution information, detection conclusion information and detection annotation information of the multimedia resource based on the first detection result and the second detection result; and generate a counterfeit detection report corresponding to the multimedia resource based on the feature contribution information, the detection conclusion information and the detection annotation information.
[0118] Optionally, the device further includes an update module configured to determine tag information of the multimedia resource in response to an update operation on the forgery detection report; update the forgery detection report based on the tag information, and obtain an updated forgery detection report corresponding to the multimedia resource.
[0119] The above is a schematic diagram of a multimedia resource forgery detection device according to this embodiment. It should be noted that the technical solution of this multimedia resource forgery detection device and the technical solution of the multimedia resource forgery detection method described above are based on the same concept. For details not described in detail in the technical solution of the multimedia resource forgery detection device, please refer to the description of the technical solution of the multimedia resource forgery detection method described above.
[0120] Figure 4 The block diagram of a computing device 400 according to one embodiment of the present disclosure is shown. Components of the computing device 400 include, but are not limited to, a memory 410 and a processor 420. The processor 420 is connected to the memory 410 via a bus 430, and a database 450 is used to store data.
[0121] The computing device 400 also includes an access device 440 that enables the computing device 400 to communicate via one or more networks 460. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 440 may include one or more of any type of network interface (e.g., a network interface card (NIC)) whether wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.
[0122] In one embodiment of the present specification, the above components of the computing device 400 and Figure 4 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 4 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.
[0123] Computing device 400 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). Computing device 400 may also be a mobile or stationary server.
[0124] The processor 420 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the forgery detection method for multimedia resources.
[0125] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the aforementioned method for detecting forgery of multimedia resources are based on the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned method for detecting forgery of multimedia resources.
[0126] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the forgery detection method for multimedia resources are implemented.
[0127] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solution of the aforementioned method for detecting forgery of multimedia resources. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned method for detecting forgery of multimedia resources.
[0128] An embodiment of the present specification further provides a computer program product, including a computer program or instructions, which implements the steps of the forgery detection method for multimedia resources when executed by a processor.
[0129] The above is a schematic diagram of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product is based on the same concept as the technical solution of the aforementioned method for detecting forgery of multimedia resources. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the aforementioned method for detecting forgery of multimedia resources.
[0130] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0131] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0132] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0133] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0134] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of the embodiments described herein. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification.
Claims
1. A method for detecting forgery of multimedia resources, characterized in that: include: Determining a fusion feature set of multimedia resources through a resource detection model, wherein each fusion feature in the fusion feature set includes image information and audio information; Calculating forgery weight information corresponding to each fused feature in the fused feature set, and screening target forgery features from the fused features according to the forgery weight information; Constructing a fine-tuning dataset based on the target forgery features, and using the fine-tuning dataset to fine-tune the resource detection model; Each fusion feature in the fusion feature set is subjected to forgery detection by using the fine-tuned resource detection model to obtain a forgery detection report corresponding to the multimedia resource.
2. The method according to claim 1, characterized in that Before determining the fusion feature set of the multimedia resource by the resource detection model, the method further includes: Separating the multimedia resources to obtain image data and audio data corresponding to the multimedia resources; The image data and the audio data are input to a resource detection model.
3. The method according to claim 2, characterized in that The resource detection model is used to determine the fusion feature set of multimedia resources, including: Performing feature extraction on the image data and the audio data respectively through a resource detection model to obtain image features corresponding to the image data and audio features corresponding to the audio data; fusing the image features and the audio features to obtain a plurality of fused features; A fusion feature set of the multimedia resource is constructed based on the multiple fusion features.
4. The method according to claim 1, wherein Calculating forged weight information corresponding to each fused feature in the fused feature set, including: Based on the preset feature selection algorithm, determine the forgery clue information corresponding to each fusion feature; Calculate the forged weight value corresponding to each forged clue information, and determine the forged weight information corresponding to each fusion feature according to the forged weight value corresponding to each forged clue information.
5. The method according to claim 4, characterized in that Screening target forged features from the fused features according to the forged weight information includes: sorting the importance of each fused feature according to the forged weight information corresponding to each fused feature to obtain a forged feature sorting list; A target forged feature is selected from the forged feature sorting list.
6. The method according to claim 1, characterized in that Construct a fine-tuning dataset based on the target forged features, including: determining historical forgery features and historical forgery clue information corresponding to the historical forgery features; A fine-tuning dataset is constructed according to the historical forgery features, the historical forgery clue information, the target forgery features, and the target forgery clue information corresponding to the target forgery features.
7. The method according to claim 1, characterized in that Performing forgery detection on each fusion feature in the fusion feature set by using the fine-tuned resource detection model to obtain a forgery detection report corresponding to the multimedia resource includes: Performing forgery detection on each fusion feature in the fusion feature set using the fine-tuned resource detection model to obtain a first detection result corresponding to the multimedia resource; Performing a forgery detection on the multimedia resource using a deep forgery detector to obtain a second detection result corresponding to the multimedia resource; A forgery detection report corresponding to the multimedia resource is generated according to the first detection result and the second detection result.
8. The method according to claim 7, characterized in that Generating a forgery detection report corresponding to the multimedia resource according to the first detection result and the second detection result includes: Determining feature contribution information, detection conclusion information, and detection annotation information of the multimedia resource according to the first detection result and the second detection result; A forgery detection report corresponding to the multimedia resource is generated according to the feature contribution information, the detection conclusion information and the detection annotation information.
9. The method according to claim 1, characterized in that After obtaining the forgery detection report corresponding to the multimedia resource, the method further includes: In response to an update operation on the forgery detection report, determining tag information of the multimedia resource; The forgery detection report is updated based on the marking information to obtain an updated forgery detection report corresponding to the multimedia resource.
10. A device for detecting forgery of multimedia resources, characterized in that: include: a determination module configured to determine a fusion feature set of a multimedia resource through a resource detection model, wherein each fusion feature in the fusion feature set includes image information and audio information; a calculation module configured to calculate forgery weight information corresponding to each fused feature in the fused feature set, and to filter a target forgery feature from the fused features according to the forgery weight information; a fine-tuning module configured to construct a fine-tuning dataset according to the target forgery features, and to fine-tune the resource detection model using the fine-tuning dataset; The detection module is configured to perform forgery detection on each fusion feature in the fusion feature set by using the fine-tuned resource detection model to obtain a forgery detection report corresponding to the multimedia resource.
11. A computing device, characterized in that include: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method for detecting forgery of multimedia resources according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium, characterized in that It stores computer-executable instructions, which, when executed by a processor, implement the steps of the method for detecting forgery of multimedia resources according to any one of claims 1 to 9.
13. A computer program product, characterized in that The method comprises a computer program or an instruction, which, when executed by a processor, implements the steps of the method for detecting counterfeit multimedia resources according to any one of claims 1 to 9.