Multimedia resource processing method and device
By extracting audio and video features and aligning the timing of multimedia resources, and using the abnormal detection model after weighted fusion, the problem of local forged fragments is solved, and efficient abnormal area detection is achieved.
Patent Information
- Application Number
- CN202510463718.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to accurately locate local forged segments mixed into real videos, and single modal feature detection is prone to missed detection, with large calculations and slow response.
By extracting audio and video features of multimedia resources, performing time-series alignment and weighted fusion, using anomaly detection model to perform abnormality detection to determine the target abnormality area.
It realizes the accurate distinction between local abnormal clips and normal audio and video clips, reduces the false alarm rate and missed detection rate, and improves the detection accuracy of abnormal areas.
Smart Images

Figure CN120337075A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and particularly to a method and apparatus for processing multimedia resources. Background Art
[0002] With the rapid development of deepfake technology, the generated fake audio and video content is becoming increasingly realistic. Although traditional forgery detection methods have achieved certain results in overall detection, it is difficult for existing methods to accurately locate local forged segments mixed in real videos.
[0003] Current research on forged video recognition mainly focuses on methods such as deep neural networks, frequency domain analysis, and generative adversarial networks. However, most of these methods focus on global feature learning and are difficult to capture local subtle forgery traces. For multi-modal data (such as synchronous tampering of audio and video), using only single-modal features is likely to result in missed detections. Some methods have the disadvantages of large computational complexity and slow response when performing high-precision positioning. Therefore, there is an urgent need for a more effective method for processing multimedia resources to solve the above problems. Summary of the Invention
[0004] In view of this, embodiments of this specification provide a method for processing multimedia resources. This specification also relates to a multimedia resource processing apparatus, a computing device, a computer-readable storage medium, and a computer program product to solve the above problems existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, a method for processing multimedia resources is provided, including:
[0006] Determine the audio resource and video resource included in the multimedia resource, perform feature extraction on the audio resource to obtain target audio features, and perform feature extraction on the video resource to obtain target video features;
[0007] After performing temporal alignment on the target audio features and the target video features, perform weighted fusion to obtain audio-visual features;
[0008] Use an anomaly detection model to perform anomaly detection on the audio-visual features, and determine a target anomaly region in the multimedia resource according to the detection result.
[0009] Optionally, the determining the audio resource and video resource included in the multimedia resource includes:
[0010] Perform audio extraction on the multimedia resource to obtain waveform data, and perform noise reduction on the waveform data to obtain the audio resource;
[0011] Perform video frame segmentation on the multimedia resource to obtain target video frames, and perform noise reduction on the target video frames to obtain the video resource.
[0012] Optionally, the feature extraction of the audio resource to obtain the target audio feature includes:
[0013] Performing audio feature extraction on the audio resource to obtain time-frequency features, and determining the spectral feature and time-domain feature of the audio resource;
[0014] Taking the video feature, the spectral feature, and the time-domain feature as the target audio feature.
[0015] Optionally, the feature extraction of the video resource to obtain the target video feature includes:
[0016] Performing video feature extraction on the video resource to obtain key visual features including edge information, light and shadow information, and texture information;
[0017] Identifying the inter-frame motion feature in the video resource by using an optical flow algorithm, and taking the key visual feature and the inter-frame motion feature as the target video feature.
[0018] Optionally, after the temporal alignment of the target audio feature and the target video feature, the weighted fusion into an audio-visual feature includes:
[0019] Performing temporal alignment on the target audio feature and the target video feature to obtain a target aligned audio feature and a target aligned video feature;
[0020] Using an attention mechanism to perform weighted fusion on the target aligned audio feature and the target aligned video feature to obtain the audio-visual feature.
[0021] Optionally, the use of an anomaly detection model to perform anomaly detection on the audio-visual feature and determine a target anomaly region in the multimedia resource includes:
[0022] Using the anomaly detection model to perform anomaly detection on the audio-visual feature, and determining at least one candidate anomaly region in the multimedia resource according to the detection result;
[0023] Determining the target anomaly region among the at least one candidate anomaly region.
[0024] Optionally, the determining of the target anomaly region among the at least one candidate anomaly region includes:
[0025] Determining the anomaly data corresponding to each of the at least one candidate anomaly region according to the detection result;
[0026] Determine the target abnormal area among the at least one candidate abnormal area according to the abnormal threshold and the abnormal data corresponding to the at least one candidate abnormal area respectively.
[0027] Optionally, after using the anomaly detection model to perform anomaly detection on the audio-visual features, the method further includes:
[0028] Determine at least one abnormal area including a bounding box in the multimedia resource according to the detection result;
[0029] Perform bounding box display detection on the at least one abnormal area including a bounding box, and determine the target area including a bounding box among the at least one abnormal area including a bounding box according to the detection result;
[0030] Generate detection data based on the target area, and display the detection data on the target page.
[0031] According to a second aspect of the embodiments of the present specification, there is provided a multimedia resource processing apparatus, including:
[0032] A determination module, configured to determine the audio resource and the video resource included in the multimedia resource, extract features from the audio resource to obtain target audio features, and extract features from the video resource to obtain target video features;
[0033] An alignment module, configured to perform temporal alignment on the target audio features and the target video features, and then perform weighted fusion to obtain audio-visual features;
[0034] A detection module, configured to use an anomaly detection model to perform anomaly detection on the audio-visual features, and determine a target abnormal area in the multimedia resource according to the detection result.
[0035] According to a third aspect of the embodiments of the present specification, there is provided a computing device, including a memory, a processor, and a computer program or instruction stored in the memory and executable on the processor. When the processor executes the computer program or instruction, the steps of the multimedia resource processing method are implemented.
[0036] According to a fourth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, which stores a computer program or instruction. When the computer program or instruction is executed by a processor, the steps of the multimedia resource processing method are implemented.
[0037] According to a fifth aspect of the embodiments of the present specification, there is provided a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, the steps of the above-mentioned multimedia resource processing method are implemented.
[0038] The multimedia resource processing method provided in this specification determines the audio resources and video resources included in the multimedia resources, extracts features from the audio resources to obtain target audio features, and extracts features from the video resources to obtain target video features; after temporally aligning the target audio features and the target video features, they are weighted and fused into audio-visual features; an anomaly detection model is used to perform anomaly detection on the audio-visual features, and according to the detection results, target anomaly regions are determined in the multimedia resources. By combining the two modalities of audio and video, more refined feature fusion is achieved through temporal alignment. Using the feature extraction and fusion strategy, the problem of confusion between local abnormal segments and normal audio-visual segments is solved, the false alarm rate and the missed detection rate are reduced, and the accurate detection of target anomaly regions is realized. Description of the Drawings
[0039] Figure 1 is a flowchart of a multimedia resource processing method provided by an embodiment of this specification;
[0040] Figure 2 is a processing flowchart of a multimedia resource processing method applied to false audio-visual localization detection provided by an embodiment of this specification;
[0041] Figure 3 is a schematic structural diagram of a multimedia resource processing device provided by an embodiment of this specification;
[0042] Figure 4 is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed Embodiments
[0043] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.
[0044] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0045] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0046] First, the noun terms involved in one or more embodiments of this specification are explained.
[0047] Dynamic Time Warping (DTW): An algorithm for measuring the similarity or distance between two time series, especially suitable for time series with non-linear changes on the time axis.
[0048] Librosa: A library designed for the Python programming language, specifically for analyzing music and audio. It provides a set of simple and easy-to-use interfaces to load, process, and analyze audio signals. With Librosa, audio feature extraction, calculation of rhythm, pitch, and audio visualization can be easily performed.
[0049] FFmpeg: A general multimedia framework that can decode, encode, transcode, multiplex, demultiplex, stream, filter, and play multimedia data in almost all formats. It supports a wide range of audio, video, and subtitle formats and can run on most operating systems. It is suitable for tasks that require complex processing of audio and video, such as format conversion, editing, merging, resolution or bitrate adjustment, etc.
[0050] Wiener filtering: A noise reduction technique used in signal processing. Based on the minimum mean square error criterion, it dynamically adjusts the filter parameters by estimating the statistical characteristics of the noise and the signal (such as the power spectrum) to achieve the optimal denoising effect in different environments.
[0051] Z-score normalization: Also known as standard deviation normalization, a technique for converting data to have zero mean and unit variance.
[0052] OpenCV: An open-source computer vision and machine learning software library. It is mainly used for image processing and computer vision tasks.
[0053] Short-Time Fourier Transform (STFT): A method for analyzing the spectrum of a signal that varies with time.
[0054] Mel Spectrogram: A spectral representation method based on the characteristics of the human auditory system's different sensitivities to different frequencies. It converts linearly spaced frequencies into frequencies on the Mel scale, which is closer to the way the human ear perceives sound.
[0055] MFCC (Mel Frequency Cepstral Coefficients): Features obtained by further processing on the basis of the Mel spectrogram, usually used in speech recognition and other audio processing tasks. By taking the logarithm and discrete cosine transform, a set of low-dimensional feature vectors are extracted from the Mel spectrogram.
[0056] Transformer: A deep learning model architecture, originally proposed for the field of natural language processing, which can effectively capture long-range dependencies in sequential data due to its self-attention mechanism. It is now also applied to the processing of other types of data such as audio and video.
[0057] BiLSTM (Bidirectional Long Short-Term Memory Network): A variant of LSTM that not only considers past information in the time series but also future information simultaneously. This is particularly useful for tasks that require a comprehensive understanding of the context, such as speech recognition and time series prediction.
[0058] Convolutional Neural Network (CNN): A deep neural network mainly used to process data with a similar grid structure (such as images). It automatically learns hierarchical feature representations of the input data through local connections and weight sharing.
[0059] ResNet-50: A residual network consisting of 50 layers, which solves the problem of gradient vanishing in deep networks by introducing skip connections, greatly improving the stability and performance of model training.
[0060] EfficientNet: A series of models that balance the width, depth, and resolution of the network through a compound scaling method, achieving a better balance between performance and efficiency.
[0061] Xception: Based on extreme Inception modules, it uses depthwise separable convolutions instead of traditional convolution operations, improving computational efficiency and model performance.
[0062] Optical Flow: A technique for estimating object motion based on the pixel movement between consecutive frames. It is widely used in video analysis, such as action recognition, video compression, and motion estimation.
[0063] Farneback: A dense optical flow estimation algorithm based on polynomial expansion, suitable for real-time applications.
[0064] TV-L1 optical flow algorithm: Combines the advantages of Total Variation (TV) regularization and L1 norm minimization, providing a robust optical flow estimation method, especially suitable for handling optical flow estimation in cases of large displacements and occlusions.
[0065] Attention mechanism (Self-Attention): Self-Attention is one of the core components of the Transformer architecture. In it, the vector at each position is compared with the vectors at all other positions to calculate its correlation score, and then these scores are used to weighted sum the values at other positions to generate a new representation for that position.
[0066] Gated Recurrent Unit (GRU): A simplified version of LSTM (Long Short-Term Memory network), designed to address the problem of vanishing or exploding gradients encountered by traditional RNNs when dealing with long time series. GRU controls the flow of information through two gates (update gate and reset gate), enabling the model to learn which information should be retained and which should be discarded.
[0067] Non-Maximum Suppression (NMS): A post-processing technique used in object detection algorithms to filter out redundant overlapping prediction bounding boxes and only retain the bounding box that is most likely to correspond to the actual object.
[0068] Heatmap: A data visualization technique that uses different colors to represent the magnitude or density of values, in order to visually display the distribution of data. In the field of computer vision, heatmaps are often used to show the degree of importance of a specific feature in an image or the confidence level of a prediction result.
[0069] In this specification, a method for processing multimedia resources is provided. This specification also relates to a multimedia resource processing device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0070] Figure 1 The flowchart of a method for processing multimedia resources according to an embodiment of this specification is shown, which specifically includes the following steps:
[0071] Step 102: Determine the audio resources and video resources included in the multimedia resource, perform feature extraction on the audio resources to obtain target audio features, and perform feature extraction on the video resources to obtain target video features.
[0072] Specifically, the multimedia resource is a media resource that includes audio and video and plays the audio and video synchronously. The multimedia resource contains false audio and video segments, and the false audio and video segments are audio and video segments that are forged (artificially edited and deviate from the truth). The audio resource contained in the multimedia resource is the resource of the sound part, and the video resource contained in the multimedia resource is the resource of the picture part. The target audio feature refers to the audio spectrum and time-domain features obtained by extracting features from the audio resource, and the target audio feature contains time-domain, frequency-domain, and time-frequency joint information. The target video feature can reflect the local detail features of the video resource. The target video feature contains key features such as the edges, light and shadow, and texture of the video resource, and also contains the inter-frame motion feature, that is, the inconsistency of the inter-frame motion.
[0073] Based on this, the multimedia resource is split to obtain the audio resource and video resource contained in the multimedia resource. The short-time Fourier transform is used to extract features from the audio resource to obtain the target audio feature, and the convolutional neural network is used to extract features from the local details of the video resource to obtain the target video feature. Subsequently, anomaly detection can be performed by fusing the target video feature and the target audio feature to determine the forged audio and video segments in the multimedia resource.
[0074] Further, when processing the audio content in the multimedia resource, it is necessary to first extract the audio part from the multimedia resource, and then process the extracted audio part to obtain the audio resource. Similarly, for the video part in the multimedia resource, video frame extraction needs to be performed to obtain the video resource. The specific implementation is as follows:
[0075] Audio extraction is performed on the multimedia resource to obtain waveform data, and noise reduction is performed on the waveform data to obtain the audio resource; video frame segmentation is performed on the multimedia resource to obtain target video frames, and noise reduction is performed on the target video frames to obtain the video resource.
[0076] Specifically, the waveform data refers to the waveform data corresponding to the audio in the multimedia resource, which is data extracted by parsing the audio in the multimedia resource. Performing noise reduction on the waveform data means performing noise suppression and normalization processing on the waveform data to achieve zero mean and standard deviation normalization of the waveform data. The audio resource is the audio that has undergone waveform data extraction and noise reduction processing. The target video frame refers to the key frame in the multimedia resource, which is a video frame obtained by extracting frame images from the multimedia resource at a fixed frame rate. Noise reduction of the target video frame can be achieved by using Gaussian filtering or bilateral filtering. The video resource is the video that has undergone video frame segmentation and video frame noise reduction.
[0077] Based on this, audio extraction is performed on multimedia resources, which are parsed through the Librosa or Librosa multimedia framework to achieve standardized sampling and obtain waveform data. Adaptive Wiener filtering is used to denoise the waveform data to obtain audio resources, realizing noise suppression of the waveform data. The multimedia resources are segmented into video frames through the Librosa multimedia framework, and frame images are extracted at a fixed frame rate to obtain target video frames. The target video frames are denoised and normalized to obtain video resources.
[0078] For example, multimedia resources can be audio-visual contents such as animations, TV dramas, short videos, surveillance videos, etc. that contain audio content and video content. In the case where the multimedia resource is a short video, the abnormal segments contained in the short video can be determined by detecting the short video. The short video is segmented into frames and audio is extracted. For video frame segmentation, FFmpeg or OpenCV is used for decoding, and frame images are extracted at a fixed frame rate (such as 30 FPS). At the same time, key frames are extracted in combination with the H.264 coding standard, and I frames are preferentially selected to reduce redundant calculations and improve calculation efficiency. For the audio part, it is parsed through Librosa or FFmpeg, and audio waveform data is extracted. The sampling rate is standardized to 16 kHz to ensure the consistency of feature calculation. In terms of data quality optimization, the video frames are denoised through Gaussian filtering or bilateral filtering, and the audio signal is generally subjected to noise suppression using adaptive Wiener filtering, and at the same time, z-score normalization is performed to make its mean zero and standard deviation normalized.
[0079] In summary, denoised and normalized audio resources and video resources are extracted from multimedia resources, providing standardized processable resources for subsequent abnormal segment detection.
[0080] Furthermore, when extracting target audio features from the audio resources, it is necessary to generate audio feature vectors in combination with the characteristics of the audio resources to obtain target audio features containing time-domain, frequency-domain, and time-frequency joint information. The specific implementation is as follows:
[0081] Audio feature extraction is performed on the audio resources to obtain time-frequency features, and the spectral features and time-domain features of the audio resources are determined; the video features, the spectral features, and the time-domain features are used as the target audio features.
[0082] Specifically, time-frequency features combine information in the time domain and the frequency domain, aiming to represent the variations of an audio signal in both the time and frequency dimensions. Spectral features show the distribution of frequency components of an audio signal at different time points. Time-domain features are characteristics directly extracted from the original audio signal (i.e., time series) without any frequency conversion. Specifically, they include envelope, zero-crossing rate, energy, etc. These features describe the properties of an audio signal on the time axis, such as the intensity change of the sound and the fluctuation of the waveform.
[0083] Based on this, the short-time Fourier transform (STFT) is used to extract audio features from the audio resource to obtain time-frequency features. The formants and spectral characteristics of the audio are further analyzed through the Mel spectrum and MFCC (Mel-frequency cepstral coefficients) to determine the spectral features and time-domain features of the audio resource. The video features, spectral features, and time-domain features are used as the target audio features.
[0084] Following the above example, for the extracted audio resource, in-depth feature analysis is performed on the audio resource. The short-time Fourier transform (STFT) is used to extract the time-frequency features of the audio resource, and then the formants and spectral characteristics of the audio are further analyzed through the Mel spectrum and MFCC (Mel-frequency cepstral coefficients). At the same time, Transformer or BiLSTM is combined for audio feature time series modeling to detect possible abnormal patterns in deepfake audio.
[0085] In summary, the video features, spectral features, and time-domain features are used as the target audio features. Subsequently, in-depth mining of abnormal segments can be achieved by detecting the target audio features, laying a foundation for subsequent multi-modal fusion.
[0086] Furthermore, when extracting target video features from the video resource, feature extraction can be performed on information such as the light and shadow, edges, and texture of the images in the video. The specific implementation is as follows:
[0087] Video feature extraction is performed on the video resource to obtain key visual features including edge information, light and shadow information, and texture information; the optical flow algorithm is used to identify the inter-frame motion features in the video resource, and the key visual features and the inter-frame motion features are used as the target video features.
[0088] Specifically, edge information refers to the boundaries between different regions in an image, usually manifested as a sharp change in color or brightness.
[0089] Light and shadow information refers to the light and dark distribution in an image, which reflects the position and intensity of the light source as well as the reflection properties of the object surface. Texture information refers to a certain repetitive pattern or structure on the image surface, which can describe the microscopic geometric properties of the material and its arrangement. Texture features contain information about surface roughness, consistency, etc. Inter-frame motion features can represent the movement of objects in a video scene and are used to detect abnormal behaviors.
[0090] Based on this, video feature extraction is performed on the video resource to obtain key visual features including edge information, light and shadow information, and texture information. The optical flow algorithm is used to identify the inter-frame motion features in the video resource to identify the inconsistencies in inter-frame motion. The key visual features and inter-frame motion features are used as target video features for subsequent feature fusion and abnormal segment detection.
[0091] Following the above example, for the extracted video resource, the video part uses a convolutional neural network (CNN) to extract local detail features, and uses ResNet-50, EfficientNet, or Xception to extract key features such as edges, light and shadow, and texture. At the same time, optical flow is calculated, and the inconsistencies in inter-frame motion are identified based on the Farneback method or the TV-L1 optical flow algorithm to discover possible forgery operations.
[0092] In summary, the key visual features and inter-frame motion features are used as target video features. Furthermore, based on the key visual features and inter-frame motion features, the abnormal segments in the video resource can be determined, laying a foundation for subsequent multi-modal fusion.
[0093] Step 104: After temporally aligning the target audio features and the target video features, they are weighted and fused into audio-visual features.
[0094] Specifically, after determining the audio resources and video resources included in the multimedia resource, extracting features from the audio resources to obtain target audio features, and extracting features from the video resources to obtain target video features, the target audio features and target video features can be temporally aligned and then weighted and fused into audio-visual features. Among them, temporally aligning the target audio features and target video features means that the target audio features and target video features need to be synchronously matched. Using the dynamic time warping method, the target audio features and target video features are temporally aligned to ensure the time synchronization of the target audio features and target video features. Weighted fusion means that by dynamically adjusting the weighted weights of the target audio features and dynamically adjusting the weighted weights of the target video features, the signal-to-noise ratios corresponding to the audio resources and video resources are optimized. If the forgery traces of the audio resources are relatively obvious, the weighted weights will tend to the target audio features; if the tampering traces of the video resources are relatively obvious, the weighted weights of the target video features will be enhanced.
[0095] Based on this, after determining the audio resources and video resources included in the multimedia resource, extracting features from the audio resources to obtain target audio features, and extracting features from the video resources to obtain target video features, the target audio features and target video features are synchronously matched to achieve the temporal alignment of the target audio features and target video features, and then the target audio features and target video features are weighted and fused to obtain audio-visual features.
[0096] Further, after temporally aligning the target audio features and target video features, the attention mechanism can be adopted in the feature fusion stage, and the specific implementation is as follows:
[0097] Temporally align the target audio features and the target video features to obtain target aligned audio features and target aligned video features; use the attention mechanism to perform weighted fusion on the target aligned audio features and the target aligned video features to obtain the audio-visual features.
[0098] Specifically, the target aligned audio features and target aligned video features are the audio features and video features obtained after synchronization in the temporal dimension, and the target aligned audio features and target aligned video features have the characteristic of time synchronization. The dynamic time warping algorithm (DTW) can be used to temporally align the target audio features and target video features. The attention mechanism can be used to dynamically adjust the weighted weights of the target audio features and dynamically adjust the weighted weights of the target video features. If the forgery traces of the audio resources are relatively obvious, the weighted weights will tend to the target audio features; if the tampering traces of the video resources are relatively obvious, the weighted weights of the target video features will be enhanced.
[0099] Based on this, the dynamic time warping algorithm (DTW) is used to synchronously match the target audio features and the target video features, achieving the temporal alignment of the target audio features and the target video features, and obtaining the target aligned audio features and the target aligned video features. The attention mechanism is used to perform weighted fusion on the target aligned audio features and the target aligned video features, dynamically adjusting the weighted weights of the target aligned audio features according to the forgery traces of the target aligned audio features, and dynamically adjusting the weighted weights of the target aligned video features according to the tampering traces of the target aligned video features, to obtain the audio-visual features.
[0100] In summary, the attention mechanism is used to perform weighted fusion on the target aligned audio features and the target aligned video features to obtain the audio-visual features, providing higher information integrity and robustness for subsequent local anomaly detection.
[0101] Step 106: Use the anomaly detection model to perform anomaly detection on the audio-visual features, and determine the target anomaly region in the multimedia resource according to the detection result.
[0102] Specifically, after the temporal alignment of the target audio features and the target video features above and the weighted fusion into the audio-visual features, the anomaly detection model can be used to perform anomaly detection on the audio-visual features, and determine the target anomaly region in the multimedia resource according to the detection result. Among them, the anomaly detection model includes a convolutional neural network and a gated recurrent unit, realizing the joint modeling of spatio-temporal features. The anomaly detection model is used to detect the abnormal resource segments in the multimedia resource, that is, the forged, altered, and tampered resource segments. The detection result is the detection result of the abnormal segments in the multimedia resource, and the detection result can include the position, length, anomaly type, anomaly confidence level, anomaly feature description of the abnormal segments in the multimedia resource, and can also include the displayed bounding box for boxing the abnormal segments in the multimedia resource. The target anomaly region is the abnormal segment included in the multimedia resource, that is, the tampered, forged, or altered resource segment.
[0103] Based on this, after the temporal alignment of the target audio features and the target video features above and the weighted fusion into the audio-visual features, the anomaly detection model is used to perform anomaly detection on the audio-visual features, realizing the positioning and marking of the abnormal resource segments, and determining the target anomaly region in the multimedia resource and marking the bounding box according to the detection result.
[0104] Furthermore, when using the anomaly detection model to perform anomaly detection on the audio-visual features, the candidate anomaly region in the multimedia resource can be determined according to the detected anomaly features, and the specific implementation is as follows:
[0105] Perform anomaly detection on the audio-visual features using the anomaly detection model, and determine at least one candidate anomaly region in the multimedia resource according to the detection result; determine the target anomaly region in the at least one candidate anomaly region.
[0106] Specifically, a candidate anomaly region refers to a resource segment determined in the multimedia resource that is suspected of having an anomaly, and the anomaly type can be resource anomaly types such as alteration, forgery, and tampering. The target anomaly region refers to a candidate anomaly region with a relatively high anomaly probability determined among at least one candidate anomaly region.
[0107] Based on this, perform anomaly detection on the audio-visual features using the anomaly detection model, and determine at least one candidate anomaly region in the multimedia resource that may have resource anomalies such as alteration, forgery, and tampering according to the detection result. Determine a candidate anomaly region with a relatively high anomaly probability among the at least one candidate anomaly region as the target anomaly region.
[0108] Continuing with the above example, a local anomaly detection model is constructed based on the fused features (audio-visual features). This model combines a CNN and a gated recurrent unit (GRU) to achieve joint modeling of spatio-temporal features. The CNN is responsible for extracting spatial features such as local edge discontinuity and abnormal light and shadow, while the GRU is responsible for temporal modeling to capture forged patterns of inter-frame changes or audio-visual asynchrony. Calculate the forged probability distribution at each pixel level and perform binarization using a threshold strategy to screen out possible forged regions, that is, candidate anomaly regions, and determine the target anomaly region with a relatively high anomaly probability among the candidate anomaly regions.
[0109] In summary, determine a candidate anomaly region with a relatively high anomaly probability among the at least one candidate anomaly region as the target anomaly region, improve the accuracy of determining the target anomaly region, and achieve precise anomaly detection of the multimedia resource.
[0110] Furthermore, after obtaining at least one candidate anomaly region, the target anomaly region with a relatively high anomaly probability can be selected among the at least one candidate anomaly region according to the anomaly data corresponding to the at least one candidate anomaly region. The specific implementation is as follows:
[0111] Determine the anomaly data corresponding to each of the at least one candidate anomaly region according to the detection result; determine the target anomaly region among the at least one candidate anomaly region according to the anomaly threshold and the anomaly data corresponding to each of the at least one candidate anomaly region.
[0112] Specifically, the anomaly data may include the anomaly confidence score of the candidate anomaly region.
[0113] Based on this, determine the abnormal data corresponding to at least one candidate abnormal area according to the detection result. According to the abnormal threshold and the abnormal data corresponding to at least one candidate abnormal area respectively, determine the target abnormal area among at least one candidate abnormal area, that is, by comparing the abnormal confidence level of each candidate abnormal area with the abnormal threshold, and taking the candidate abnormal area with the abnormal confidence level greater than the abnormal threshold as the target abnormal area, so as to screen out the resource segments with a higher probability of abnormality and avoid the influence of false abnormality judgment.
[0114] To sum up, according to the abnormal threshold and the abnormal data corresponding to at least one candidate abnormal area respectively, determine the target abnormal area among at least one candidate abnormal area, so as to detect the resource segments with a higher probability of abnormality in the multimedia resource and realize the accurate detection of abnormal resource segments.
[0115] Furthermore, after using the anomaly detection model to detect anomalies in the audio-visual features, the detection data can be displayed on the target page in a visual way for subsequent manual review and viewing of the anomaly detection results. The specific implementation is as follows:
[0116] Determine at least one abnormal area containing a bounding box in the multimedia resource according to the detection result; perform bounding box display detection on the at least one abnormal area containing a bounding box, and determine the target area containing a bounding box among the at least one abnormal area containing a bounding box according to the detection result; generate detection data based on the target area and display the detection data on the target page.
[0117] Specifically, the bounding box is used to select the abnormal resource segment in the multimedia resource, and the area corresponding to the abnormal resource segment is the abnormal area. Performing bounding box display detection on at least one abnormal area containing a bounding box is used to detect whether there is an overlapping area between at least two bounding boxes. The target area is the abnormal area remaining after removing the overlapping bounding box area. The detection data includes the position and shape of the bounding box of the target area, as well as the confidence score of the abnormal degree of the target area and the description content of the abnormal features.
[0118] Based on this, determine at least one abnormal area containing a bounding box in the multimedia resource according to the detection result. Perform bounding box display detection on at least one abnormal area containing a bounding box to detect whether there are overlapping bounding boxes between at least two bounding boxes. If it is determined according to the detection result that there are overlapping bounding boxes, delete the overlapping bounding boxes, and then determine the target area containing a bounding box among at least one abnormal area containing a bounding box. The bounding box of the target area has no overlapping problem, and the confidence score of the anomaly detection of the target area is relatively high. Generate detection data based on data such as the abnormal confidence score, abnormal feature description, abnormal position, and abnormal type corresponding to the target area, and display the detection data on the target page through charts such as heat maps.
[0119] Continuing with the above example, to optimize the detection accuracy of the local forgery area, the Non-Maximum Suppression (NMS) method is used to screen overlapping bounding boxes, remove low-confidence overlapping areas, and ensure that the boundaries of the finally output forgery areas are clear and accurate. Finally, a detailed detection report is generated, including the bounding boxes, confidence scores, descriptions of abnormal features, etc. of all forgery areas, and the detection results are displayed through a visualization interface. Users can intuitively view the heatmap of the forgery areas and the timing analysis curve of audio-video desynchronization.
[0120] In summary, detection data is generated based on the target area, and the detection data is displayed through charts such as heatmaps on the target page, improving the visualization of the detection data and providing strong support for the accurate identification of deepfake videos.
[0121] The multimedia resource processing method provided in this specification determines the audio resources and video resources included in the multimedia resource, extracts features from the audio resources to obtain target audio features, and extracts features from the video resources to obtain target video features; after performing temporal alignment on the target audio features and target video features, they are weighted and fused into audio-video features; an anomaly detection model is used to perform anomaly detection on the audio-video features, and according to the detection results, target anomaly areas are determined in the multimedia resource. By combining the two modalities of audio and video, more refined feature fusion is achieved through temporal alignment. Using the feature extraction and fusion strategy, the problem of confusion between local abnormal segments and normal audio-video segments is solved, the false alarm rate and missed detection rate are reduced, and accurate detection of target anomaly areas is realized.
[0122] The following Figure 2 , taking the application of the multimedia resource processing method provided in this specification in the false audio-video localization detection as an example, further illustrates the multimedia resource processing method. Among them, Figure 2 FIG. shows the processing flow chart of a multimedia resource processing method applied to false audio-video localization detection provided in an embodiment of this specification, which specifically includes the following steps:
[0123] Step 202: Extract the waveform data from the multimedia resource for audio, and perform noise reduction on the waveform data to obtain the audio resource.
[0124] For the audio part in the multimedia resource, it is parsed through Librosa or FFmpeg, and the audio waveform data is extracted. The sampling rate is standardized to 16 kHz to ensure the consistency of feature calculation. In practical applications, the audio signal generally uses adaptive Wiener filtering for noise suppression, and at the same time, z-score normalization is performed to make its mean zero and standard deviation normalized to ensure the stability of the model input data.
[0125] Step 204: Perform video frame segmentation on the multimedia resource to obtain target video frames, and perform noise reduction on the target video frames to obtain a video resource.
[0126] For the video part in the multimedia resource, video frame segmentation is decoded using FFmpeg or OpenCV, and frame images are extracted at a fixed frame rate (such as 30 FPS). Meanwhile, key frame extraction is performed in combination with the H.264 encoding standard, and I-frames are preferentially selected to reduce redundant calculations and improve computational efficiency. Thus, a set of high-quality, denoised and normalized audio-visual data is obtained, providing a standardized input for subsequent feature extraction.
[0127] Step 206: Extract audio features from the audio resource to obtain time-frequency features, and determine the spectral features and time-domain features of the audio resource, and use the video features, spectral features, and time-domain features as target audio features.
[0128] For the audio resource, the short-time Fourier transform (STFT) is used to extract time-frequency features, and the resonance peaks and spectral characteristics of the audio are further analyzed through the Mel spectrum and MFCC (Mel-frequency cepstral coefficients). Meanwhile, Transformer or BiLSTM is combined for audio feature temporal modeling to detect possible abnormal patterns in deepfake audio.
[0129] Step 208: Extract video features from the video resource to obtain key visual features including edge information, light and shadow information, and texture information.
[0130] Step 210: Use the optical flow algorithm to identify the inter-frame motion features in the video resource, and use the key visual features and inter-frame motion features as target video features.
[0131] For the video resource, a convolutional neural network (CNN) is used for local detail feature extraction, and key features such as edges, light and shadow, and texture are extracted using ResNet-50, EfficientNet, or Xception. Meanwhile, optical flow (OpticalFlow) is calculated, and the inconsistency of inter-frame motion is identified based on the Farneback method or the TV-L1 optical flow algorithm to detect possible forgery operations. Thus, a set of high-dimensional audio-visual feature vectors is obtained, respectively containing time-domain, frequency-domain, and time-frequency joint information, laying a foundation for subsequent multi-modal fusion.
[0132] Step 212: Perform temporal alignment on the target audio features and target video features to obtain target aligned audio features and target aligned video features.
[0133] Step 214: Use the attention mechanism to perform weighted fusion on the target aligned audio features and target aligned video features to obtain audio-visual features.
[0134] Using the self-attention mechanism, the weighted weights of different modality features (target audio features and target video features) are dynamically adjusted to optimize the signal-to-noise ratio of forgery traces. In practical applications, if the forgery traces in the audio part are more obvious, the attention weights will tend to the audio features, while when the tampering traces in the video part are stronger, the weights of visual features will be correspondingly enhanced. Finally, this stage generates a set of fused feature vectors, that is, audio-visual features. The feature vectors specifically contain temporally synchronized audio-visual features, providing higher information integrity and robustness for subsequent local anomaly detection.
[0135] Step 216: Use the anomaly detection model to perform anomaly detection on the audio-visual features, and determine at least one anomaly region containing a bounding box in the multimedia resource according to the detection results.
[0136] Step 218: Perform bounding box display detection on at least one anomaly region containing a bounding box, and determine a target region containing a bounding box in at least one anomaly region containing a bounding box according to the detection results.
[0137] Step 220: Generate detection data based on the target region and display the detection data on the target interface.
[0138] Construct a local anomaly detection model based on the fused features (audio-visual features). This model combines a convolutional neural network (CNN) and a gated recurrent unit (GRU) to achieve joint modeling of spatio-temporal features. The CNN is responsible for extracting spatial features, such as local edge discontinuities and abnormal light and shadow, while the GRU is responsible for temporal modeling to capture forgery patterns of inter-frame changes or audio-visual asynchrony. Calculate the forgery probability distribution at each pixel level and perform binary processing using a threshold strategy to screen out possible forgery regions. At the same time, generate bounding boxes to mark possible local forgery regions and give a confidence score for each candidate region. Finally, this stage outputs a set of boundary information of the forgery regions, providing a basis for subsequent result optimization.
[0139] Furthermore, in order to optimize the detection accuracy of the local forgery area, the Non-Maximum Suppression (NMS) method is adopted to screen the overlapping bounding boxes, remove the overlapping areas with low confidence, and ensure that the boundaries of the finally output forgery areas are clear and accurate. Finally, a detailed detection report is generated, including the bounding boxes, confidence scores, and descriptions of abnormal features of all forgery areas, and the detection results are displayed through a visualization interface (target interface). Users can intuitively view the heatmap of the forgery areas and the timing analysis curve of audio-video asynchrony. The final output of this stage is a set of detection results for manual review or automatic forensics, providing strong support for the accurate identification of deepfake videos.
[0140] The multimedia resource processing method provided by an embodiment of this specification uses multi-modal deep feature extraction and attention mechanism to achieve local accurate positioning of forged segments. An efficient audio-video feature synchronization and fusion method is designed, which can simultaneously capture the forgery traces in audio and video. A positioning algorithm based on temporal anomaly detection is proposed, and the forgery areas are accurately labeled through local anomaly probability scoring.
[0141] Corresponding to the above method embodiment, this specification also provides an embodiment of a multimedia resource processing device. Figure 3 The structural schematic diagram of a multimedia resource processing device provided by an embodiment of this specification is shown. As Figure 3 shown, the device includes:
[0142] A determination module 302, configured to determine the audio resource and video resource included in the multimedia resource, extract features from the audio resource to obtain target audio features, and extract features from the video resource to obtain target video features;
[0143] An alignment module 304, configured to perform temporal alignment on the target audio features and the target video features, and then weighted-fuse them into audio-video features;
[0144] A detection module 306, configured to perform anomaly detection on the audio-video features by using an anomaly detection model, and determine a target anomaly area in the multimedia resource according to the detection result.
[0145] In an optional embodiment, the determination module 302 is further configured to:
[0146] Extract audio from the multimedia resource to obtain waveform data, and perform noise reduction on the waveform data to obtain the audio resource;
[0147] Perform video frame segmentation on the multimedia resource to obtain target video frames, and perform noise reduction on the target video frames to obtain the video resource.
[0148] An optional embodiment, the determining module 302 is further configured to:
[0149] Extract audio features from the audio resource to obtain time-frequency features, and determine the spectral features and time-domain features of the audio resource;
[0150] Use the video features, the spectral features, and the time-domain features as the target audio features.
[0151] An optional embodiment, the determining module 302 is further configured to:
[0152] Extract video features from the video resource to obtain key visual features including edge information, light and shadow information, and texture information;
[0153] Use the optical flow algorithm to identify the inter-frame motion features in the video resource, and use the key visual features and the inter-frame motion features as the target video features.
[0154] An optional embodiment, the alignment module 304 is further configured to:
[0155] Perform temporal alignment on the target audio features and the target video features to obtain target aligned audio features and target aligned video features;
[0156] Use the attention mechanism to perform weighted fusion on the target aligned audio features and the target aligned video features to obtain the audio-visual features.
[0157] An optional embodiment, the detection module 306 is further configured to:
[0158] Use the anomaly detection model to perform anomaly detection on the audio-visual features, and determine at least one candidate anomaly region in the multimedia resource according to the detection result;
[0159] Determine the target anomaly region in the at least one candidate anomaly region.
[0160] An optional embodiment, the detection module 306 is further configured to:
[0161] Determine the anomaly data corresponding to each of the at least one candidate anomaly region according to the detection result;
[0162] According to the anomaly threshold and the anomaly data corresponding to each of the at least one candidate anomaly region, determine the target anomaly region in the at least one candidate anomaly region.
[0163] An optional embodiment, the detection module 306 is further configured to:
[0164] Determine at least one abnormal area containing a bounding box in the multimedia resource according to the detection result;
[0165] Perform bounding box display detection on the at least one abnormal area containing a bounding box, and determine a target area containing a bounding box in the at least one abnormal area containing a bounding box according to the detection result;
[0166] Generate detection data based on the target area, and display the detection data on the target page.
[0167] The multimedia resource processing device provided in this specification determines the audio resource and video resource included in the multimedia resource, extracts features from the audio resource to obtain target audio features, and extracts features from the video resource to obtain target video features; after performing temporal alignment on the target audio features and target video features, they are weighted and fused into audio-visual features; an anomaly detection model is used to perform anomaly detection on the audio-visual features, and a target abnormal area is determined in the multimedia resource according to the detection result. By combining two modalities of audio and video, more refined feature fusion is achieved through temporal alignment. Using the feature extraction and fusion strategy, the problem of confusion between local abnormal segments and normal audio-visual segments is solved, the false alarm rate and missed detection rate are reduced, and accurate detection of the target abnormal area is realized.
[0168] The above is a schematic solution of a multimedia resource processing device according to this embodiment. It should be noted that the technical solution of this multimedia resource processing device and the technical solution of the above multimedia resource processing method belong to the same concept. For the details not described in detail in the technical solution of the multimedia resource processing device, reference can be made to the description of the technical solution of the above multimedia resource processing method.
[0169] Figure 4 The structural block diagram of a computing device 400 provided according to an embodiment of this specification is shown. The components of the computing device 400 include but are not limited to a memory 410 and a processor 420. The processor 420 is connected to the memory 410 through a bus 430, and a database 450 is used to store data.
[0170] The computing device 400 also includes an access device 440, which enables the computing device 400 to communicate via one or more networks 460. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 440 may include one or more of any type of wired or wireless network interfaces (e.g., network interface controller (NIC)), such as IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, Worldwide Interoperability for Microwave Access (Wi-MAX) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.
[0171] In one embodiment of the present specification, the above components of the computing device 400 and Figure 4 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 4 the block diagram of the computing device shown is only for illustrative purposes and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0172] The computing device 400 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptops, notebooks, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 400 can also be a mobile or stationary server.
[0173] Wherein, when the processor 420 executes the computer program or instructions, the steps of the multimedia resource processing method are implemented.
[0174] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above multimedia resource processing method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the description of the technical solution of the above multimedia resource processing method.
[0175] An embodiment of this specification also provides a computer-readable storage medium, which stores a computer program or instruction. When the computer program or instruction is executed by a processor, the steps of the multimedia resource processing method described above are implemented.
[0176] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above multimedia resource processing method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above multimedia resource processing method.
[0177] An embodiment of this specification also provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, the steps of the above multimedia resource processing method are implemented.
[0178] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above multimedia resource processing method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the description of the technical solution of the above multimedia resource processing method.
[0179] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0180] The computer program or instruction includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0181] It should be noted that, for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that this specification is not limited by the described action sequence, because according to this specification, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this specification.
[0182] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0183] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A method for processing multimedia resources, characterized in that, Including: Determine the audio resources and video resources included in the multimedia resource, extract features from the audio resources to obtain target audio features, and extract features from the video resources to obtain target video features; After temporally aligning the target audio features and the target video features, perform weighted fusion to obtain audio-visual features; Use an anomaly detection model to perform anomaly detection on the audio-visual features, and determine a target anomaly region in the multimedia resource according to the detection result.
2. The multimedia resource processing method according to claim 1, wherein The determination of the audio resources and video resources included in the multimedia resource includes: Extract waveform data from the multimedia resource for audio extraction, and perform noise reduction on the waveform data to obtain the audio resources; Perform video frame segmentation on the multimedia resource to obtain target video frames, and perform noise reduction on the target video frames to obtain the video resources.
3. The multimedia resource processing method according to claim 1, wherein The extraction of features from the audio resources to obtain target audio features includes: Extract audio features from the audio resources to obtain time-frequency features, and determine the spectral features and time-domain features of the audio resources; Use the video features, the spectral features, and the time-domain features as the target audio features.
4. The multimedia resource processing method according to claim 1, characterized in that, The extraction of features from the video resources to obtain target video features includes: Extract video features from the video resources to obtain key visual features including edge information, light and shadow information, and texture information; Use an optical flow algorithm to identify the inter-frame motion features in the video resources, and use the key visual features and the inter-frame motion features as the target video features.
5. The multimedia resource processing method according to claim 1, wherein The temporal alignment of the target audio features and the target video features and then weighted fusion to obtain audio-visual features includes: Perform temporal alignment on the target audio features and the target video features to obtain target aligned audio features and target aligned video features; Use an attention mechanism to perform weighted fusion on the target aligned audio features and the target aligned video features to obtain the audio-visual features.
6. The multimedia resource processing method according to claim 1, wherein The use of an anomaly detection model to perform anomaly detection on the audio-visual features and determine a target anomaly region in the multimedia resource according to the detection result includes: Use the anomaly detection model to perform anomaly detection on the audio-visual features, and determine at least one candidate anomaly region in the multimedia resource according to the detection result; Determine the target anomaly region among the at least one candidate anomaly region.
7. The multimedia resource processing method according to claim 6, wherein The determination of the target anomaly region among the at least one candidate anomaly region includes: Determine the anomaly data corresponding to each of the at least one candidate anomaly region according to the detection result; Determine the target anomaly region among the at least one candidate anomaly region according to an anomaly threshold and the anomaly data corresponding to each of the at least one candidate anomaly region.
8. The multimedia resource processing method according to claim 1, wherein After using the anomaly detection model to perform anomaly detection on the audio-visual features, it further includes: Determine at least one anomaly region including a bounding box in the multimedia resource according to the detection result; Perform bounding box display detection on the at least one abnormal region containing a bounding box, and determine a target region containing a bounding box in the at least one abnormal region containing a bounding box according to the detection result; Generate detection data based on the target region and display the detection data on the target page.
9. A multimedia resource processing device, characterized in that, Including: A determination module configured to determine the audio resources and video resources included in the multimedia resource, perform feature extraction on the audio resources to obtain target audio features, and perform feature extraction on the video resources to obtain target video features; An alignment module configured to perform temporal alignment on the target audio features and the target video features and then perform weighted fusion to obtain audio-visual features; A detection module configured to perform anomaly detection on the audio-visual features by using an anomaly detection model and determine a target abnormal region in the multimedia resource according to the detection result.
10. A computing device, comprising a memory, a processor, and a computer program or instructions stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program or instruction, the steps of the method according to any one of claims 1-8 are implemented.
11. A computer-readable storage medium stores a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.
12. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instruction is executed by the processor, the steps of the method according to any one of claims 1-8 are implemented.