Data processing method and electronic equipment
By identifying and processing target multimedia data, extracting features using a large visual text model and an audio encoder, and generating genuine and fake information, the problem of identifying deepfake audio and video is solved, protecting user privacy and information security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-03
AI Technical Summary
The high realism and complexity of deepfake audio and video make it difficult for users to distinguish between genuine and fake content, threatening personal privacy and information security, and affecting the social trust system.
By identifying and processing the target multimedia data, the authenticity information is output, including the authenticity type of the data frame, the time period of forgery, and spatial location. Feature extraction and fusion are performed using a large visual text model and an audio encoder to generate the recognition result.
It effectively identifies deepfake audio and video, protects user privacy, ensures information security, provides authenticity verification and traceability information, and improves the accuracy of user decision-making.
Smart Images

Figure CN121789296A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of data processing and artificial intelligence technology, and in particular to a data processing method and an electronic device. Background Technology
[0002] With the rapid development of artificial intelligence technology, especially generative artificial intelligence technology, the threshold for generating fake audio and video has been significantly lowered. Deepfake audio and video has a high degree of realism and complexity, making it difficult for people to distinguish the authenticity of deepfake audio and video by ear and naked eye.
[0003] For example, in applications such as audio and video conferencing, individuals can simulate the timbre and tone of others' voices and mimic their facial features through forged voices and video footage. The highly realistic overlay of these forged audio and video clips significantly increases the difficulty of identifying them, making users more susceptible to making incorrect decisions. This not only threatens personal privacy and property security but also poses a serious challenge to information security and social trust systems. Therefore, addressing the challenges posed by deepfake audio and video to protect user privacy and ensure user information security has become a pressing technical problem that needs to be solved in this field. Summary of the Invention
[0004] Therefore, this application discloses the following technical solution:
[0005] A data processing method, comprising:
[0006] In response to a target trigger operation, target multimedia data is identified and processed. The target multimedia data includes at least one data frame input to the target application or currently output by the target application. The target application is an application capable of performing the identification and processing or an application capable of calling a target program file to perform the identification and processing.
[0007] Output the identification result for the target multimedia data, the identification result being able to indicate at least some data frames and / or at least some data in the target multimedia data the authenticity information.
[0008] An electronic device includes at least one processor and at least one processing model capable of running on the processor, the processing model being invoked by a target application to perform the following operations:
[0009] In response to a target trigger operation, target multimedia data is identified and processed. The target multimedia data includes at least one data frame input to the target application or currently output by the target application. The target application is an application capable of performing the identification and processing or an application capable of calling a target program file to perform the identification and processing.
[0010] Output the identification result for the target multimedia data, the identification result being able to indicate at least some data frames and / or at least some data in the target multimedia data the authenticity information.
[0011] A storage medium carrying one or more computer instruction sets, which, when executed by an electronic device, enable the electronic device to perform any of the data processing methods provided above. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0013] Figure 1 This is a flowchart of the data processing method provided in this application;
[0014] Figure 2 This application provides an implementation architecture for identifying and processing target audio and video data.
[0015] Figure 3 This is the network structure diagram of the large visual text model provided in this application;
[0016] Figure 4 This is a network structure diagram of the video encoder provided in this application;
[0017] Figure 5 This is a network structure diagram of the video decoder provided in this application;
[0018] Figure 6 This is a network structure diagram of the audio encoder provided in this application;
[0019] Figure 7 This is a network structure diagram of the audio-video fusion module provided in this application;
[0020] Figure 8 This is a network structure diagram of the boundary matching layer in the audio-video fusion module provided in this application;
[0021] Figure 9 This is a network structure diagram of the boundary fusion module in the audio-visual fusion module provided in this application;
[0022] Figure 10 This is a network structure diagram of the audio classifier provided in this application;
[0023] Figure 11This application provides another implementation architecture for identifying and processing target audio and video data;
[0024] Figure 12 This is an example diagram illustrating the application of deepfake audio and video detection, spatiotemporal localization, and text interpretation provided in this application;
[0025] Figure 13 This is a schematic diagram of the input / output of the recognition processing provided in this application. Detailed Implementation
[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0027] This application provides a data processing method and an electronic device. The provided data processing method can be applied to electronic devices in a variety of general or special computing device environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor devices, etc.
[0028] See Figure 1 The flowchart shown illustrates the data processing method provided in this application embodiment, which may include the following steps 101 to 102, which are described in detail below.
[0029] Step 101: In response to the target trigger operation, identify and process the target multimedia data.
[0030] The target triggering operation may include, but is not limited to, user input of various types of target multimedia data such as target video, target image, target audio, or target audio-visual data, or user / electronic device query operation of target multimedia data such as target image (which may include, but is not limited to, the image currently displayed on the device or the image input to the target application).
[0031] The query operation for target multimedia data such as target images can be, but is not limited to, any one of the following ad:
[0032] a. Voice questioning: Ask questions using voice commands;
[0033] b. Text-based questioning: Asking questions via text.
[0034] c. The operation of selecting question options in the application interface: Asking a question is accomplished by selecting the desired option from the various question options provided in the application interface;
[0035] d. Image recognition operation based on scene switching or recognition determination: In response to scene switching or recognition, image recognition is triggered and the operation is regarded as a questioning operation on the image. The requirement or goal of the questioning operation is to recognize the specified image in order to realize scene-related processing based on image recognition, such as scene recognition and / or scene recognition-based related control.
[0036] The target multimedia data includes at least one data frame input to the target application or currently output by the target application, wherein the target application is an application capable of performing the recognition process or an application capable of calling the target program file to perform the recognition process.
[0037] The at least one data frame may be, but is not limited to, a video stream or an audio stream, or it may be an image frame.
[0038] In response to a target trigger operation, the identification processing of the target multimedia data may include, but is not limited to, identifying the source of various types of target multimedia data such as images, audio, and video based on the target application or the target program file called by the target application (e.g., which generation model it was generated from, which terminal it came from, or where it came from), and / or identifying the authenticity of the target multimedia data, and / or, if it is found that there is forged data in the target multimedia data, further identifying the forged time period of the audio or video frame and / or identifying the forged area of the video frame to achieve forged spatial positioning, etc.
[0039] In one possible implementation, the target application is capable of performing the identification process. In this implementation, the target application may be, but is not limited to, an application with the ability to identify the authenticity of multimedia data and / or to perform local authenticity identification of multimedia data (such as identifying which areas of a fake video frame in a video are fake areas). In practice, the target application can be designed and implemented based on the identification processing logic provided in this application.
[0040] In other possible implementations, the target application cannot directly implement the recognition process, but can execute it by calling a target program file. The target program file can be, but is not limited to, an application, a piece of program code, or a model file that can be used to execute the recognition process.
[0041] In other words, in this embodiment of the application, the user can input the target multimedia data into the target application that can perform the recognition process, or the target application currently displaying the output can directly perform recognition processing such as authenticity verification on the multimedia data currently displayed by the target application by calling the target program file. For example, WeChat, a browser, or a website can call the model file to perform authenticity verification on the multimedia data stream output by WeChat, a browser, or a website.
[0042] Step 102: Output the recognition result for the target multimedia data, the recognition result being able to indicate at least some data frames and / or at least some data in the target multimedia data, the authenticity information.
[0043] After the target multimedia data is processed for identification and the identification result is obtained, the identification result can be output to facilitate users to view or use the information.
[0044] The recognition results of the target multimedia data can be output through any one or more of the following methods, including but not limited to text output, image output, and voice playback.
[0045] The at least some data frames may include, but are not limited to, data frames within a specific time period in the target multimedia data, such as audio frames and / or video frames within a specific time period in the target multimedia data where the corresponding data content is forged. The authenticity information of the at least some data frames may include, but is not limited to, the authenticity type of the data frames within a specific time period in the target multimedia data, and the corresponding time period (such as the forged time period).
[0046] At least a portion of the data in the target multimedia data's data frame may include, but is not limited to, a certain image region in a video frame or image frame, such as a face region in a video frame or image frame. The authenticity information of at least a portion of the data in the target multimedia data's data frame may include, but is not limited to, the authenticity type of a certain region in a video frame or image frame, and the spatial positioning information of that region in the corresponding video frame / image frame (such as the region pixel position of the face forgery region) if that region is a forged region (such as a face forgery region).
[0047] In addition to indicating the authenticity of at least some data frames and / or at least some data in the target multimedia data, the identification result may also include, but is not limited to, explanations of the authenticity information, target responses to target questions, and / or relevant source information of the target multimedia data.
[0048] The explanation of the authenticity information may include, but is not limited to, the reasons for judging the authenticity of the target multimedia data, the reasons for judging the authenticity of local spatial regions (such as local video frames) or local time periods in the target multimedia data, and / or the forgery technology or generation model used for forged regions or forged audio segments, etc., so that users can more easily understand the logic behind the authenticity judgment.
[0049] The target question information may include, but is not limited to, questions from users regarding the authenticity of the target multimedia data, such as "Is there a fake segment in the video from the 10th to the 30th second? If so, please locate the fake time period, determine how the fake image was generated, and mark the fake area in the video."
[0050] The relevant traceability information for the target multimedia data may include, but is not limited to, the source information of the target multimedia data (such as which device it came from, or which person or device it originally came from), or, in the case of the target multimedia data being forged, which model was used to forge the target multimedia data, etc.
[0051] The source information related to the target multimedia data can also be understood as part of the explanatory information regarding the authenticity of the information.
[0052] For example, the output of the recognition result for the target multimedia data can be implemented as any one of the following a:
[0053] a. Output authenticity information for at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, and an explanation of the authenticity information.
[0054] The at least partial data frames may include, but are not limited to, at least partial video frames, audio frames, or audio-visual frames in the target audio-visual data. The at least partial region data may include at least partial screen regions in at least partial video frames of the target audio-visual data.
[0055] In this embodiment a, the specific information may include, but is not limited to, the authenticity type of at least some video frames, audio frames, or audio-visual frames in the target audio-visual data, the corresponding forgery time period and / or the spatial positioning information of the forged area in the forged video frame, as well as the source information of the forged video segment / audio-visual segment and related reasons / basis (such as the forgery technology or generation model used), etc. The explanation of the source information is the reason for drawing the conclusion of authenticity.
[0056] b. Output authenticity information and a first control for at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, wherein the first control can be triggered to display an explanation of the authenticity information.
[0057] In this implementation method b, instead of directly displaying the explanation, the first control (such as "View More") is displayed as a first control or a thumbnail display. The full explanation is then displayed after the control is triggered (such as by clicking).
[0058] c. Output authenticity information for at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, and, when the authenticity information is triggered, output an explanation of the authenticity information.
[0059] In this embodiment c, the authenticity information is output in the form of a control, for example, a control is output and the authenticity information is displayed within the area of the control. After the user further operates the control of the authenticity information (such as clicking the control), the action of outputting an explanation can be triggered, thereby calling and outputting an explanation for the authenticity information.
[0060] d. Output the authenticity information of at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, and the response information to the target question information.
[0061] In this implementation method d, in addition to displaying authenticity information, response information for target questions can also be displayed. Target questions can be questions about the authenticity of target audio and video data, or questions about other aspects besides authenticity identification. For example, if the user's question is not about the authenticity of the target audio and video data, but about other aspects such as the generation time, acquisition time, or the photographer, the corresponding response information can be timestamp information or other information, such as photographer information, generation location information, etc.
[0062] e. Output authenticity information and a second control for at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, wherein the second control can be triggered to display the associated information of the authenticity information.
[0063] The second control may be, but is not limited to, a recommendation control for related information concerning the authenticity of the information. After the second control is triggered, related information concerning the authenticity of the information is displayed to facilitate the user's viewing or use of the related information.
[0064] The information associated with the authenticity of the data may include, but is not limited to, relevant introductory or recommendation information for the target audio and video data, such as purchase links for items involved in the target audio and video data, configuration interfaces for setting image content in the target audio and video data as travel destinations, sharing interfaces for sharing image content, etc.
[0065] In summary, the data processing method provided in this embodiment, in response to a target trigger operation, allows the target application or the target application to call a target program file to identify and process the target multimedia data, and generates and outputs identification results that can at least indicate the authenticity of at least some data frames and / or at least some data in the data frames of the target multimedia data. This enables the identification of the authenticity of at least some data frames and / or at least some data in the data frames of the target multimedia data, thereby effectively addressing the challenges posed by deepfake audio and video, protecting user privacy, and ensuring user information security.
[0066] In one optional embodiment, the target multimedia data is target audio and video data, which includes video data and audio data. Optionally, the video data and audio data included in the target audio and video data are synchronized in time. For example, the mouth movements of the speaker in the video frame are aligned with the speaking sounds in the audio data in time.
[0067] In step 101 of the method provided in this application, the identification and processing of the target multimedia data can be implemented as follows: steps “1-1”-“1-3”:
[0068] 1-1: Obtain the audio feature data and image feature data of the target audio and video data.
[0069] Video data in the target audio and video data can be processed using a video encoder, but is not limited to, to obtain a video embedding representation (such as a video embedding vector) of the video data, and this video embedding representation can be used as image feature data of the target audio and video data.
[0070] In practice, optionally, the video data in the target audio and video data can first be converted into a superimage, and features can be extracted from the superimage based on the video encoder. The extracted features of the superimage can be used as the video embedding representation of the video data, and correspondingly as the image feature data of the target audio and video data. The superimage includes the pixel information of each pixel in the target audio and video data in all video frames. More specifically, the information of each pixel in the superimage can include the pixel information of each pixel at the same position as the pixel in each video frame of the target audio and video data.
[0071] Similarly, an audio encoder can be used, but is not limited to, to process the audio data in the target audio and video data to obtain the audio embedding representation (such as an audio embedding vector) of the audio data, and the audio embedding representation can be used as the audio feature data of the target audio and video data.
[0072] The image feature data includes overall and local information of each video frame in the target audio-visual data. Optionally, the overall information may include detection features of each video frame in the target audio-visual data, and the local information may include segmentation features of each video frame in the target audio-visual data.
[0073] Video frame detection features primarily focus on identifying and locating specific objects or events within a video frame. These features typically include spatial location, category labels, and motion information. Detection features often emphasize global context, such as the bounding box coordinates and category probabilities of objects, to support subsequent tracking, behavior analysis, and authenticity determination.
[0074] Video frame segmentation features focus on pixel-level partitioning, classifying each pixel in a video frame into a corresponding semantic category (such as foreground, background, or a specific object). Segmentation tasks rely on more refined local features, such as edges, textures, and color distributions. Common segmentation features include, but are not limited to, threshold-based segmentation, region growing, or edge detection, which utilize pixel intensity differences or spatial continuity to divide regions.
[0075] 1-2: When the target question information is obtained, the text feature data of the target question information is acquired, and the first fused feature data obtained by fusing the text feature data and the image feature data is used for reasoning to obtain the authenticity information of at least a portion of the video frame data in the target audio and video data, and / or to obtain the authenticity information of at least a portion of the video frame in the target audio and video data.
[0076] The target questions may include, but are not limited to, questions about the authenticity and type of the target audio and video data, the temporal and spatial location of the forged video segment, the temporal location of the forged audio segment, and the forgery methods / evidence.
[0077] In practical applications, target question information can be obtained based on any one or more of the following methods, including but not limited to text input, speech-to-text processing, image-to-text processing, or semantic recognition.
[0078] After obtaining the target query information, it can be input into a text encoder for processing to obtain text feature data such as word embeddings. Then, the text feature data and image feature data can be further fused to obtain first fused feature data, and inference can be performed based on this first fused feature data. Alternatively, inference can be performed by combining audio feature data to obtain authenticity information for at least a portion of the video frames in the target audio / video data, and / or, to obtain authenticity information for at least a portion of the video frames in the target audio / video data.
[0079] For example, by reasoning based on the first fused feature data, spatial positioning information of the forged video segment in the target audio and video data (such as the pixel position of the forged area in the forged video segment), the authenticity type of the video, and the response information to the target's questions can be obtained. Alternatively, by combining audio feature data for reasoning, temporal positioning information of the forged audio / video in the target audio and video data (such as the forged time period) can be obtained. This part will be described in detail in the embodiments below.
[0080] 1-3: Without obtaining the target question information, reasoning is performed on the second fused feature data obtained by fusing the audio feature data and the image feature data to obtain the authenticity information of at least some audio and video frames in the target audio and video data.
[0081] In the absence of target query information, optionally, the second fused feature data can be obtained by processing the image feature data through a fully connected layer and then fusing it with the audio feature data. For example, the video embedding representation can be processed through a fully connected layer and then fused with the audio embedding representation. Based on this, the authenticity information of at least some audio and video frames, such as the audio forgery time period and / or video forgery time period in the target audio and video data, can be inferred from the second fused feature data.
[0082] This embodiment can perform different inferences based on the audio feature data and image feature data of the target audio and video data, depending on whether the target question information is obtained or not, and obtain identification information such as authenticity identification of the target audio and video data under the condition that the target question information is obtained or not. In this way, it can realize the authenticity identification of at least some data frames and / or at least some data in the data frames of the target audio and video data, effectively cope with the challenges brought by deepfake audio and video, and thus protect user privacy and ensure user information security.
[0083] In an optional embodiment, a further provided process is provided for performing identification processing such as authenticity verification on the target audio and video data when the target question information is obtained.
[0084] See Figure 2An example of the implementation architecture is provided. Given the target question information, the overall implementation architecture for identifying and processing target audio and video data, such as distinguishing between genuine and fake data, can include relevant models for visual text processing, video codecs, audio encoders, audio classifiers, and related functional modules such as fully connected layers, attention modules, and time period filtering modules. Based on the collaboration between relevant processing models, codecs, and / or functional modules, the identification and processing of target audio and video data, such as distinguishing between genuine and fake data, can be achieved.
[0085] In practical applications, a target application or target program file capable of performing the recognition processing can be implemented based on this architecture, so as to provide the ability to recognize and process target multimedia data such as target audio and video data in the target application or target program file.
[0086] Among them, the relevant models for visual text processing may include, but are not limited to, large visual text models that can be used for multimodal data processing of visual and text information.
[0087] See Figure 2 In this embodiment, the video data and audio data in the target audio and video data, as well as the target question information (such as...) in response to the target audio and video data can be included. Figure 2 (Text question in the text) as Figure 2 The architecture shown takes overall input information and, after passing through its recognition logic, outputs at least one of the following as the recognition result for the target audio and video data:
[0088] a. Output the authenticity type of the video and audio, as the result of judging the authenticity of the audio and video.
[0089] b. Output spatial positioning video as the spatial positioning result of the fake video.
[0090] Spatial positioning video can refer to a forged video segment in the target audio and video data where the forged area of the forged video image has been marked.
[0091] c. Output fake time periods, which may include fake audio time periods and / or fake video time periods, to show users the fake audio / video time intervals identified in the audio / video time series of the target audio and video data.
[0092] d. Output text responses as answers to the target question, and also as explanations or reasons / basis for determining the authenticity of the video.
[0093] Among them, such as Figure 2 As shown, the overall working principle of this architecture is as follows:
[0094] The target question information and video data from the target audio / video data are input into the visual-text big data model. The model extracts features from the video data to obtain video embedding representations as image feature data. It then extracts text features from the target question information to obtain text feature data. The image and text feature data are fused to obtain the first fused feature data. Based on this first fused feature data, inference is performed to generate textual responses to the target question information. The final hidden layer of the visual-text big data model generates target segmentation features and target detection features (such as...). Figure 2 The segmentation feature representation and detection feature representation shown in the diagram are processed by a fully connected layer to generate video authenticity type, then by another fully connected layer, and finally fused with the target segmentation feature through an attention module and residual connections. The fused feature representation is then input into the video decoder. The video data in the target audio-visual data is processed by a video encoder to generate image feature data such as video embedding representation. The video embedding representation and the fused feature representation are then processed by the video decoder to generate spatially localized video. Simultaneously, the audio data in the target audio-visual data is processed by an audio encoder to generate audio embedding representation and other audio features. The video embedding representation undergoes feature transformation through a fully connected layer. The transformed video embedding representation and the audio embedding representation have the same dimension. Both are input into the audio-video fusion module for fusion. After the audio embedding representation and the transformed video embedding representation pass through the audio-video fusion module, a candidate forgery time period set is generated. Redundant forgery time periods in the candidate forgery time period set can be removed, and accurate and effective forgery time periods can be filtered and output.
[0095] In addition, the audio embedding representation can be processed by an audio classifier to generate audio authenticity information.
[0096] In this embodiment, the first fused feature data obtained by fusing the text feature data and the image feature data in steps 1-2 above is used for inference to obtain authenticity information for at least a portion of the data in the video frames of the target audio and video data, and / or to obtain authenticity information for at least a portion of the video frames of the target audio and video data. This can be implemented as at least some of the steps in steps "2-1"-"2-3" below:
[0097] 2-1. Input the first fused feature into the question answering model, and extract the target detection feature and target segmentation feature from the hidden layer when the question answering model processes the first fused feature.
[0098] The question-answering model can be a component of a large visual text model, used to determine the response information corresponding to the target question information based on the input information.
[0099] The target detection feature may include a fusion feature of the overall information in the image feature data and the text feature data; the target segmentation feature may include a fusion feature of the local information in the image feature data and the text feature data.
[0100] 2-2. Based on the target detection features, determine the authenticity information of at least a portion of the video frames in the target audio and video data.
[0101] like Figure 2 As shown, optionally, the target detection features can be integrated and classified through a fully connected layer (the first fully connected layer) to obtain the authenticity type of the video data in the target audio and video data.
[0102] 2-3. Based on the target detection features, the target segmentation features, and the image feature data, determine the spatial positioning information corresponding to the forged image content in the video frame of the target audio and video data.
[0103] like Figure 2 As shown, the target detection features can be processed through another fully connected layer (the second fully connected layer) for feature integration, and then the processed target detection features are fused with the target segmentation features after attention processing by the attention module and residual connection by the residual module to obtain the third fused feature. The fused feature representation, i.e. the third fused feature, and the image feature data of the target audio and video data are input into the video decoder to obtain the spatial positioning information of the forged screen content in the video frame of the target audio and video data determined by the video decoder based on the target interrogation information, such as the pixel position of the region corresponding to the forged screen area.
[0104] Based on this embodiment, target audio and video data can be identified and processed when target question information is obtained. Based on this identification and processing, information such as the authenticity of the target audio and video data, the time period of the forged audio, the time period of the forged video, the spatial location of the forged video, and / or the response to the target question information can be obtained. This can effectively address the challenges brought by deepfake audio and video, protect user privacy, and ensure user information security.
[0105] In an alternative embodiment, the following are provided: Figure 2 The diagram illustrates the composition and processing of the large visual text model within the implementation architecture.
[0106] The visual-text big model can not only understand and align textual semantics with visual information and generate responses to target questions to provide textual explanations for judging fake videos, but also provide feature representations for judging fake videos and locating the spatial regions of fake videos (such as the pixel positions of fake video regions).
[0107] See Figure 3 The network structure diagram of the visual-text large model is shown below. Optionally, the visual-text large model may include a visual encoder, a text encoder, and a question-answering model. The question-answering model may be, but is not limited to, a large language model, used to determine the response information corresponding to the target question information based on the input information.
[0108] The input to the visual-text large model includes video data from the target audio-visual data and target question information, such as... Figure 3 The "video" and "text questions" in the text are used to define the video data dimension. Where C represents the number of color channels of a pixel in an image frame in the video, T represents the number of image frames in the video data, and H and W represent the height and width of the image frame, respectively. Optionally, this embodiment uses a numerical stacking method to stack the color values of all frames of the video data at the same pixel position to form the color information of that pixel position, thereby constructing a superimage. The color information of each pixel in the superimage includes the pixel information of each pixel at the same position in each video frame of the video data. The data dimension of the superimage is... ,in .
[0109] The superimage is input to the visual encoder, which extracts the visual features of the superimage as a video embedding representation, that is, as the image feature data of the video data in the target audio data. The target question information is input to the text encoder. Optionally, the text encoder tokenizes the target question information to generate a word embedding token, which can be used as the text feature data of the target question information.
[0110] After obtaining the video embedding representation, a linear projection matrix can be used to transform the video embedding representation into the word embedding token space, generating video embedding tokens. These video embedding tokens and word embedding tokens have the same dimension. Then, the word embedding tokens and video embedding tokens can be concatenated to form a joint embedding token, which is the concrete implementation of the first fused feature data.
[0111] Based on this, the first fused feature data, such as the joint embedded token, is input into a question-answering model such as a large language model. The question-answering model receives the first fused feature data and performs inference based on the first fused features to generate and output response information to the target question, such as... Figure 3 The text response in the video is used to answer the user's targeted question and can also explain why the video is judged to be real or fake.
[0112] Furthermore, in order to use the visual text large model for detecting the authenticity of videos and locating the spatial position of fake videos, this embodiment uses two feature locators, namely a first feature locator and a second feature locator, to expand the vocabulary of the visual text large model. The first and second feature locators can be text characters, such as text characters. <det>and <seg>The first feature locator is as follows: <det>The second feature locator is used to locate the detection features in the intermediate feature representation of question-answering models such as large language models. <seg>These two locators are used to locate segmentation features in the intermediate state feature representation of question-answering models such as large language models. They can be applied to tasks such as detecting the authenticity of videos and locating the spatial location of fake videos, respectively.
[0113] For example, given a video dataset and a target question, a question-answering model such as a large language model generates a response to the target question after the aforementioned multimodal data processing, feature fusion, and inference. This response is contained in the intermediate feature representation corresponding to the question-answering model. <det>and <seg>The text character, based on which the corresponding text can be extracted from the last hidden layer of the question-answering model, is respectively related to... <det>and <seg>Corresponding feature representations, and based on <det>The extracted feature representations will be used as object detection features, based on <seg>The extracted features are represented as target segmentation features. Subsequent target detection features will be applied to the task of detecting the authenticity of videos. The target detection features and target segmentation features will be used together to locate the spatial position of fake videos.
[0114] In practice, the pre-trained CLIP (Contrastive Language-Image Pre-training) visual encoder ViT-L / 14 can be used as the visual encoder in the large visual text model, and the pre-trained Vicuna model can be used as the large language model in the large visual text model, etc., as question answering models.
[0115] Based on the composition and functional design of the visual text large model in this embodiment, it is possible to generate corresponding responses to target questions in the target audio and video data, facilitating the explanation of relevant reasons or judgment criteria to users. Simultaneously, it can provide feature representations for subsequent judgments of the authenticity of video data. Furthermore, by processing the video frame sequence in the video data into a superimage, this embodiment enables the use of an image encoder (such as the ViT-L / 14 visual encoder) to extract features from the video data. This improves the flexibility of video feature extraction, balances feature extraction performance with resource consumption, and ensures feature extraction efficiency, performance, and practicality.
[0116] In an alternative embodiment, the following are provided: Figure 2 The diagram illustrates the composition and processing of the video encoder in the implementation architecture.
[0117] The video encoder can be implemented using, but is not limited to, a pre-trained ViT-H image encoder.
[0118] See Figure 4 The network structure diagram of the video encoder is shown below. Optionally, the video encoder includes an image segmentation module, a linear projection module, and a Transformer encoder. The image segmentation module is used to segment the hyperimage into blocks, and the linear projection module is used to perform linear projection on the tiling result of the image blocks. The Transformer encoder may include, but is not limited to, components such as a multilayer perceptron, a layer normalization module, and a self-attention module. It is mainly used to transform the input information into a high-dimensional vector representation rich in contextual information based on the collaboration of these components, thereby capturing the semantic relationships and structural patterns in the input.
[0119] The video encoder can receive a sequence of video frames from the target audio and video data as input, and process the input data to generate a video embedding representation, which serves as the image feature data of the video data.
[0120] It should be noted that, Figure 2 Both the video encoder in the big data model and the visual encoder in the visual text model can extract image feature data from the target audio and video data. In practical applications, the image feature data extracted by the video encoder or the image feature data extracted by the visual encoder in the big data model can be selected based on the processing stage or processing requirements.
[0121] Let the video data dimension be . Where C represents the number of color channels of pixels in an image frame in the video, T represents the number of image frames in the video data, and H and W represent the height and width of the image frame, respectively.
[0122] Optionally, before inputting the video frame sequence of the video data into the video encoder, a numerical stacking method is first used to stack the color values of all frames of the video data at the same pixel position to form the color information of the pixel at that pixel position, thereby constructing a superimage. Then, the superimage is input into the video encoder, where the image segmentation module in the video encoder performs block processing on the superimage to obtain each block / image block of the superimage. After that, each image block is tiled and linearly projected based on the linear projection module. Finally, after the linear projection result is connected and the position is encoded, it is input into the Transformer encoder to be converted into the corresponding high-dimensional vector representation. This high-dimensional vector representation is the video embedding representation of the video data, which is also the image feature data of the video data.
[0123] See also Figure 2 Subsequently, a fully connected layer can be used to transform the video embedding representation output by the video encoder to align the dimensions of the video embedding representation with the audio embedding representation generated by the audio encoder. The transformed video embedding representation and audio embedding representation can then be input into the audio-video fusion module to generate fake time period information for audio and video.
[0124] Based on the composition and functional design of the video encoder in this embodiment, feature extraction of video data from target audio data can be achieved, resulting in image feature data of the video data. This provides support for subsequent identification processing such as authenticity verification, addressing the challenges posed by deepfake audio and video, protecting user privacy, and ensuring user information security. Furthermore, by processing the video frame sequence in the video data into a superimage, this embodiment enables feature extraction of the video data using an image encoder (such as the ViT-H image encoder). This improves the flexibility of video feature extraction, balancing feature extraction performance and resource consumption, and ensuring feature extraction efficiency, performance, and practicality.
[0125] In an alternative embodiment, the following are provided: Figure 2 The diagram illustrates the composition and processing of the video decoder in the implementation architecture.
[0126] See Figure 5 The network structure diagram of the video decoder is shown below. The video decoder may include, but is not limited to, various components such as decoder layers, transposed convolutions, token and video cross-attention, multiple perceptrons, and video frame rearrangement. The following describes the process of video decoding based on the collaborative implementation of these components.
[0127] The "segmentation and detection fusion feature representation" formed by fusing the target segmentation features and target detection features generated by the visual text large model is used as a token input to the video decoder. The video decoder maps the video embedding representation output by the video encoder and the "segmentation and detection fusion feature representation" into a masked video and an IoU (Intersection over Union) score. The masked video contains spatial localization information of the fake video segments in the video data (such as the pixel positions of fake regions in the fake video segments) to be used to output spatially localized videos. During the model training phase, the IoU score is used to calculate the loss error of the network model and optimize the network model parameters through the backpropagation mechanism. During the model inference phase, the IoU score is used to represent the accuracy and confidence of the model in predicting the spatial localization of fake videos.
[0128] The core of the video decoder is located in Figure 5 Optionally, in this embodiment, two cascaded decoder layers are used in the left-hand decoder layer. The first decoder layer outputs an updated token and video embedding representation as input to the second decoder layer.
[0129] In practical applications, it is not limited to using two decoder layers. The number of decoder layers can also be one or more. There is no restriction on this, and it can be determined according to the actual application requirements.
[0130] The processing of each decoder layer includes four steps: First, self-attention is performed on the token; second, the token is used as the query input for cross-attention, and cross-attention is performed between the token and the video embedding representation output by the video encoder; third, a multilayer perceptron is used to update the token; fourth, the video embedding representation is used as the query input for cross-attention, and cross-attention is performed between the video embedding representation and the token, which updates the video embedding representation. When the video embedding representation participates in the cross-attention operation, this embodiment adds positional encoding to the video embedding representation so that the decoder layer can fully utilize the spatial positional information of the video.
[0131] After two decoder layers, this embodiment uses two transposed convolutions to upsample the updated video embedding representation so that the size of the masked video is aligned with the size of the input video. The activation function of each transposed convolution can be, but is not limited to, GELU (Gaussian Error Linear Unit). A normalization layer is used between the two transposed convolution layers to normalize the feature map.
[0132] Subsequently, this embodiment uses the updated token as the query input for cross-attention, and performs cross-attention operations using the updated token and the updated video embedding representation to generate a mask token and an IoU token. The IoU token passes through a multilayer perceptron to generate an IoU score. The mask token passes through a multilayer perceptron to generate a weight matrix (optionally, the elements of the weight matrix are values between 0 and 1, representing the probability that the corresponding pixel is real / fake). The channel dimension of this weight matrix matches the channel dimension of the upsampled video embedding. The weight matrix can be spatially multiplied with the upsampled video embedding to generate a feature representation representing the masked video. This feature representation passes through a video frame rearrangement module to generate a masked video. The spatial size of the masked video is the same as the spatial size of the input video. The masked region in the masked video represents the spatial region corresponding to the forged image content in the corresponding forged video segment of the video data. Optionally, the video frame reordering module includes a convolutional layer (such as a Conv1×1 convolutional layer) and two operation modules: numerical reordering. The convolutional layer adjusts the number of channels in the masked video feature representation, while numerical reordering unfolds the values in the channel dimension according to the number of video frames. The numerical reordering process is the inverse of the numerical stacking process in the video encoder (i.e., the process of generating a hyperimage).
[0133] Based on the composition and functional design of the video decoder in this embodiment, it is possible to decode the video embedding representation (image feature data of video data) output by the video encoder and the "segmentation and detection fusion feature representation" obtained based on the visual text large model, so as to obtain the spatial positioning information of the fake video segment in the video data (such as the spatial positioning video containing the fake video segment and marked with the fake screen area in it), thereby effectively identifying deepfake videos and locating and marking the fake screen area in the fake video segment, providing support for protecting user privacy and ensuring user information security.
[0134] In an alternative embodiment, the following are provided: Figure 2 The diagram illustrates the composition and processing of the audio encoder in the implementation architecture.
[0135] See Figure 6 The network structure diagram of the audio encoder is shown. The audio encoder may include, but is not limited to, various functional modules such as spectralization (e.g., Log-Mel spectralization), convolutional layers (e.g., two-dimensional convolutional layers), and max pooling layers.
[0136] The following describes the process by which an audio encoder extracts audio features based on the collaboration of its components.
[0137] The role of the audio encoder is to learn audio features from the audio frame sequence of audio data to generate audio embedding representations. In the model training phase of the audio encoder, optionally, this embodiment uses a contrastive loss function to align the semantics of the audio embedding representation with the video embedding representation.
[0138] In the audio encoder, firstly, the spectral information of the audio frame sequence is generated based on the spectralization module. For example, the Log-Mel spectrum of the audio frame sequence is generated in the Log space based on the Log-Mel spectralization module. Then, the Log-Mel spectrum is processed sequentially through two two-dimensional convolutional layers and a max pooling layer to shorten the time dimension and generate an audio embedding representation, thereby obtaining the audio feature data of the audio data.
[0139] Subsequently, the audio embedding representation, i.e., the audio feature data, can be input into an audio classifier for inference to generate the audio authenticity type, or the audio embedding representation and video embedding representation can be input into an audio-video fusion module to generate a fused audio-video boundary, so as to locate the temporal information of the forged audio / video based on the fused audio-video boundary.
[0140] The fused audio and video boundary is represented by the set of forged time periods corresponding to the forged audio segments and forged video segments in the target audio and video data. Optionally, each forged time period is represented by the start and end times of the forged time period.
[0141] Based on the composition and functional design of the audio encoder in this embodiment, feature extraction of audio data can be achieved to obtain audio feature data, thereby providing data support for subsequent processing such as audio data authenticity identification and forgery time location. This facilitates the effective identification of deepfake audio and the location of forgery time, so as to protect user privacy and ensure user information security.
[0142] In an alternative embodiment, the following are provided: Figure 2 The diagram illustrates the composition and processing of the audio-video fusion module in the implementation architecture.
[0143] See Figure 7 The network structure diagram of the audio-video fusion module shows that the audio-video fusion module can include a boundary matching layer and a boundary fusion module.
[0144] The audio-video fusion module achieves audio-video fusion through the collaboration of its various components as follows:
[0145] The video embedding representation (image feature data of video data) passes through a fully connected layer and then through a boundary matching layer to generate a video boundary map. Similarly, the audio embedding representation (audio feature data of audio data) passes through a boundary matching layer to generate an audio boundary map. The video boundary map includes the mapping relationship between different time periods in the video data and the authenticity of the corresponding video segment content, while the audio boundary map includes the mapping relationship between different time periods in the audio data and the authenticity of the corresponding audio segment content.
[0146] Subsequently, the video embedding representation, audio embedding representation, video boundary mapping, and audio boundary mapping are processed by the boundary fusion module to generate fused audio and video boundaries. The fused audio and video boundaries are represented as a set of fake time segments. Each fake time segment is represented by its start and end times. The set of fake time segments is then processed by a time segment filtering module to remove redundant time segment intervals, resulting in the output fake time segments. The output fake time segments can contain fake audio time segments and fake video time segments.
[0147] The time period filtering module may, but is not limited to, use a soft non-maximum suppression algorithm (Soft-NMS) to remove redundant fake time periods from the candidate fake time period set, so as to filter out accurate and effective fake time periods for output.
[0148] See Figure 8 It provides a network structure diagram of the boundary matching layer in the audio and video fusion module. The role of the boundary matching layer is to uniformly sample in the feature embedding to generate features with temporal context information as boundary mappings. The sampling weights are predefined weight matrices, and the feature embeddings and sampling weights are multiplied by a dot product in the time dimension.
[0149] See Figure 9 This document provides a network structure diagram of the boundary fusion module within the audio-video fusion module, which includes two sets of one-dimensional convolutional modules. Specifically, the video boundary mapping, video embedding representation, and audio embedding representation are each processed by the first set of one-dimensional convolutions to generate corresponding feature matrices (i.e., three feature matrices). The elements corresponding to these three feature matrices (elements at the same position in the three matrices) are then averaged to generate the video weight matrix. Similarly, the video embedding representation, audio embedding representation, and audio boundary mapping are processed by the second set of one-dimensional convolutions to generate corresponding feature matrices (again, three feature matrices). The elements corresponding to these three feature matrices (elements at the same position in the three matrices) are also averaged to generate the audio weight matrix.
[0150] The video weight matrix is used as the weight for the video boundary mapping, and the audio weight matrix is used as the weight for the audio boundary mapping. Based on this, the elements at corresponding positions in the video boundary mapping and audio boundary mapping can be weighted, summed, and normalized one by one according to the weights corresponding to the video boundary mapping and audio boundary mapping, so as to generate the fused audio and video boundary.
[0151] Based on the composition and functional design of the audio-video fusion module in this embodiment, the image feature data of video data and the audio feature data of audio data can be fused to obtain the fused audio-video boundary, thereby providing support for the subsequent location of the forged time of video and audio data. This facilitates the effective identification of deepfake audio and video and the location of the forgery time, so as to protect user privacy and ensure user information security.
[0152] In an alternative embodiment, the following are provided: Figure 2 The diagram illustrates the composition and processing of the audio classifier in the implementation architecture.
[0153] See Figure 10 The network structure diagram of the audio classifier is shown. Optionally, the audio classifier includes fully connected layers and activation function layers. The activation function layers can be, but are not limited to, Sigmoid activation function layers.
[0154] Among them, the fully connected layer plays a core role in the classification task. Its main function is to integrate and map the extracted distributed feature representations to the sample label space (such as class labels) to achieve the final classification decision. The activation function layer plays a core role in the classifier by introducing nonlinear factors, thereby improving the model's expressive power and enabling it to solve linearly inseparable problems.
[0155] The audio classifier receives the audio embedding representation (audio feature data) output by the audio encoder as input, and processes the input information through a fully connected layer and an activation function layer to output the audio authenticity type. Thus, the solution based on this embodiment can effectively identify deepfake audio, providing support for protecting user privacy and ensuring user information security.
[0156] In an optional embodiment, a process for identifying and processing target audio and video data, such as verifying authenticity, is further provided when the target question information is not obtained.
[0157] Even without obtaining the target query information, the target audio and video data can be processed for authenticity verification based on the overall architecture shown in Figure 11. This implementation architecture can include a video encoder, an audio encoder, an audio classifier, a fully connected layer, an audio-video fusion module, a time segment filtering module, and other components. Based on the collaboration between these components, the target audio and video data can be processed for authenticity verification.
[0158] in, Figure 11 The structure and processing of the various components, including the mid-video encoder, audio encoder, audio classifier, fully connected layer, audio-video fusion module, and time segment filtering module, are discussed in relation to their respective functions. Figure 2 The structure and processing of the corresponding objects (such as encoders, classifiers, or related modules) are the same; please refer to the section above for details. Figure 2 The descriptions of the corresponding components will not be repeated here.
[0159] In this embodiment, in the case where the target question information is not obtained, before reasoning on the second fused feature data obtained by fusing audio feature data and image feature data, the audio feature data and image feature data can be fused to obtain the second fused feature data.
[0160] Optionally, the fusion processing of audio feature data and image feature data can be further implemented as follows: determining video boundary mapping based on image feature data, determining audio boundary mapping based on audio feature data, and fusing video data features, audio data features, video boundary mapping, and audio boundary mapping to obtain the second fused feature data. The meanings of video boundary mapping and audio boundary mapping are detailed in the descriptions of the corresponding embodiments above and will not be repeated here.
[0161] Specifically, the image feature data of the video data is passed through a fully connected layer, and then further passed through the boundary matching layer in the audio-video fusion module to generate a video boundary map. Similarly, the audio feature data of the audio data is passed through the boundary matching layer in the audio-video fusion module to generate an audio boundary map. Then, the image feature data (such as video embedding representation), audio feature data (such as audio embedding representation), video boundary map, and audio boundary map are passed through the boundary fusion module in the audio-video fusion module to generate a fused audio-video boundary. This fused audio-video boundary can then be used as the second fused feature data.
[0162] Based on this, inference can be performed based on the second fused feature data to obtain authenticity information for at least a portion of the audio and video frames in the target audio and video data. This process can be further implemented as follows: determining a candidate forgery time period set based on the second fused feature data, wherein the candidate forgery time period in the candidate forgery time period set represents that the corresponding video segment picture content and / or audio segment contains forged data content; removing redundant candidate forgery time periods from the candidate forgery time period set to obtain a non-redundant candidate forgery time period set; and determining the time position information corresponding to the forged picture content in the target audio and video data and the time position information corresponding to the forged audio content in the target audio and video data based on the non-redundant candidate forgery time period set.
[0163] The fused audio-visual boundary, also known as the second fused feature, is represented by a set of forged time segments. Each forged time segment is represented by its start and end times, based on... Figure 11 The implementation architecture, after obtaining the fused audio and video boundaries, can input the set of fake time periods contained in the fused audio and video boundaries into the time period filtering module, so as to use the time period filtering module to remove redundant time period intervals in the set of fake time periods, thereby outputting the final fake time period after removing redundancy. The final fake time period can include fake audio time periods (time position information corresponding to fake audio content) and fake video time periods (time position information corresponding to fake video content).
[0164] In addition, such as Figure 11 As shown, the audio feature data output by the audio encoder can also be input into the audio classifier, so that the audio classifier can determine the authenticity of the audio data based on the audio feature data, thereby obtaining the authenticity type of the audio data.
[0165] Based on this embodiment, in cases where target question information is not obtained, target audio and video data can be identified and processed. Based on this identification and processing, the authenticity of audio data in the target audio and video data and information such as the time period of the forged audio and the time period of the forged video can be obtained, which can effectively address the challenges brought by deepfake audio and video, protect user privacy, and ensure user information security.
[0166] It is worth noting that the facial data and other biometric data used in this application are from legal sources and comply with privacy regulations. In practical applications, for example, but not limited to, pop-up prompts, users can be reminded that facial data and other biometric data need to be collected, and this data should only be collected after obtaining the user's authorization, to ensure the legality and compliance of the facial data and other biometric data used.
[0167] In an optional embodiment, an application example of using the method of this application for deepfake audio and video detection, spatiotemporal localization, and text interpretation is provided.
[0168] See this example. Figure 12 Given an audio / video file, extract the video and audio data from it. At the same time, the user inputs a text question: "Is there a fake segment in the video from the 10th to the 30th second? If so, please locate the fake time period, determine how the fake image was generated, and mark the fake area in the video frame."
[0169] The implementing entity of this application's scheme, such as based on Figure 2 The target application or program file implemented by the architecture shown can receive video data, audio data, and user text queries as input. Based on this, the target application and other executing entities can perform identification processing on the video and audio data, including authenticity verification, spatiotemporal localization, and text interpretation. For example, the identification processing results include: identifying the video type as fake and the audio type as fake; locating the fake video and audio time periods within the entire audio-visual sequence from the start to the end time, with each fake time period represented by a corresponding start and end time; and generating a text response: "The footage from second 12 to second 16 of the video is fake; the male face has been altered using face-swapping technology, and the altered facial area is marked in the output video."
[0170] Based on this, the corresponding recognition and processing results can be output. See the output results in this example. Figure 12 As shown, the text response "the footage from the 12th to the 16th second of the video is fake" corresponds to the start and end times of the fake video period. The text response "the male face was altered using face-swapping technology, and the altered facial area is marked in the output video" corresponds to the output spatial positioning video. The text response not only explains that the male facial area in the video data was altered, but also explains that the alteration technology is face-swapping technology, and the altered facial area is located and marked with a dashed box in the output video.
[0171] See Figure 13 The diagram shown illustrates the input / output of deepfake audio / video detection, spatiotemporal localization, and text interpretation. Based on the input target audio / video data (including video and audio data) and the question information, it can ultimately output various recognition results such as text answers, video authenticity type, spatially located video, fake time period, and audio authenticity type. This can meet the comprehensive needs of users in audio / video application scenarios and ensure user information security.
[0172] This application also discloses an electronic device, including at least one processor and at least one processing model capable of running on the processor, wherein the processing model can be invoked by a target application to perform the data processing method provided in any of the embodiments above.
[0173] The processor can be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a neural network processor (NPU), a deep learning processor (DPU), or other programmable logic devices.
[0174] Optionally, the electronic device may also include a display device that can be used to output and display information.
[0175] Optionally, electronic devices may also include storage resources such as memory, RAM, and cache.
[0176] Optionally, the electronic device may also include an image acquisition device.
[0177] In addition to these components, electronic devices may also include communication interfaces, communication buses, and other parts. Memory, processor, and communication interface communicate with each other through the communication bus.
[0178] Communication interfaces are used for communication between electronic devices and other devices. Communication buses can be Peripheral Component Interconnect (PCI) buses or Extended Industry Standard Architecture (EISA) buses, and can be categorized into address buses, data buses, control buses, etc.
[0179] This application also discloses a storage medium carrying one or more computer instruction sets, which, when executed by an electronic device, enable the electronic device to implement the data processing method provided in any of the above method embodiments.
[0180] In summary, the solution of this application embodiment has at least the following technical advantages compared with the prior art:
[0181] 1. This application provides an end-to-end solution for identifying forged audio and video and spatiotemporal positioning. It can not only determine the authenticity of video and audio, but also locate the forged time period of audio and video and the forged spatial area of video frame, providing fine-grained and transparent judgment results. This allows users to quickly find the forged spatiotemporal area, which can meet the comprehensive needs of users in audio and video communication scenarios, and also plays an important role in judicial evidence collection.
[0182] 2. This application allows users to ask questions about the authenticity of videos by inputting text statements. The scope of text questions not only supports the video's spatial frame but also custom time intervals, meeting users' personalized needs. Furthermore, this application's solution provides a technical foundation for further developing more intelligent multimodal (e.g., user-generated questions via voice) deepfake audio and video detection and localization methods.
[0183] 3. This application explains the reasons for judging the authenticity of a video by providing a textual answer. This not only allows for an overall judgment on the authenticity of the video, but also enables analysis and judgment on the authenticity of specific spatial areas and time periods of the video based on the questions raised by the user. This makes it easier for users to understand the logic behind the authenticity judgment, thereby improving the user experience.
[0184] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0185] For ease of description, the above systems or devices are described separately as various modules or units based on their functions. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware components.
[0186] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence or the part that makes a creative contribution, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0187] Finally, it should be noted that in this document, relational terms such as first, second, third, and fourth are used to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one…" does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0188] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.< / seg> < / det> < / seg> < / det> < / seg> < / det> < / seg> < / det> < / seg> < / det>
Claims
1. A data processing method, comprising: In response to a target trigger operation, target multimedia data is identified and processed. The target multimedia data includes at least one data frame input to the target application or currently output by the target application. The target application is an application capable of performing the identification and processing or an application capable of calling a target program file to perform the identification and processing. Output the identification result for the target multimedia data, the identification result being able to indicate at least some data frames and / or at least some data in the target multimedia data the authenticity information.
2. The data processing method according to claim 1, wherein the target multimedia data is target audio and video data, wherein, The target multimedia data is identified and processed, including: Acquire audio feature data and image feature data of the target audio and video data, wherein the image feature data includes overall information and local information of each video frame of the target audio and video data; Given the target question information, the text feature data of the target question information is obtained, and the first fused feature data obtained by fusing the text feature data and the image feature data is used for reasoning to obtain the authenticity information of at least a portion of the video frame data in the target audio and video data, and / or to obtain the authenticity information of at least a portion of the video frame data in the target audio and video data. Alternatively, in the absence of target question information, reasoning can be performed on the second fused feature data obtained by fusing the audio feature data and the image feature data to obtain authenticity information for at least a portion of the audio and video frames in the target audio and video data.
3. The data processing method according to claim 2, acquiring the audio feature data and image feature data of the target audio and video data, includes: The video data in the target audio and video data is converted into a superimage, and the superimage is used to extract features based on the video encoder to obtain the image feature data of the target audio and video data; the superimage includes the pixel information of each pixel in the target audio and video data in all video frames. The audio feature data is obtained by extracting the audio features of the audio data in the target audio and video data based on the audio encoder.
4. The data processing method according to claim 2, wherein reasoning is performed on the first fused feature data obtained by fusing the text feature data and the image feature data to obtain authenticity information for at least a portion of the video frames in the target audio-visual data, and / or, to obtain authenticity information for at least a portion of the video frames in the target audio-visual data, comprising: The first fused feature is input into the question answering model, and target detection features and target segmentation features are extracted from the hidden layer when the question answering model processes the first fused feature; the question answering model is used to determine the response information corresponding to the target question information, the target detection feature includes the fused feature of the overall information and the text feature data, and the target segmentation feature includes the fused feature of the local information and the text feature data; Based on the target detection features, determine the authenticity information of at least a portion of the video frames in the target audio and video data; And / or, Based on the target detection features, the target segmentation features, and the image feature data, the spatial positioning information corresponding to the forged image content in the video frame of the target audio and video data is determined.
5. The data processing method according to claim 4, based on the target detection features, determining the authenticity information of at least a portion of video frames in the target audio and video data, including: The target detection features are processed based on the first fully connected layer to obtain the authenticity information of the video data in the target audio and video data.
6. The data processing method according to claim 4, based on the target detection features, the target segmentation features, and the image feature data, determining the spatial positioning information corresponding to the forged scene content in the video frame of the target audio and video data, including: The target detection features are processed based on the second fully connected layer, and the processed target detection features are fused with the target segmentation features after passing through an attention module and residual connections to obtain the third fused feature; The image feature data and the third fusion feature are input into the video decoder to obtain the spatial positioning information of the forged image content in the video frame of the target audio and video data determined by the video decoder based on the target interrogation information.
7. The data processing method according to claim 2, further comprising, before performing inference on the second fused feature data obtained by fusing the audio feature data and the image feature data, the method includes: The audio feature data and the image feature data are fused together. The process of fusing the audio feature data and the image feature data includes: Based on the image feature data, a video boundary mapping is determined. The video boundary mapping includes the mapping relationship between the authenticity of different time periods in the video data of the target audio and video data and the corresponding video segment content. Based on the audio feature data, an audio boundary mapping is determined, which includes the mapping relationship between the authenticity of different time periods in the audio data of the target audio and video data and the corresponding audio segment content. The video data features, the audio data features, the video boundary mapping, and the audio boundary mapping are fused to obtain the second fused feature data.
8. The data processing method according to claim 7, wherein reasoning is performed on the second fused feature data obtained by fusing the audio feature data and the image feature data to obtain authenticity information for at least a portion of the audio and video frames in the target audio and video data, including: A candidate forgery time period set is determined based on the second fused feature data; the candidate forgery time period in the candidate forgery time period set indicates that the corresponding video segment and / or audio segment contains forged data content; Remove redundant candidate forgery time periods from the candidate forgery time period set to obtain a non-redundant candidate forgery time period set; Based on the non-redundant candidate forgery time period set, the time location information corresponding to the forged visual content in the target audio and video data, and the time location information corresponding to the forged audio content in the target audio and video data are determined.
9. The method according to claim 1, wherein, Output the recognition result for the target multimedia data, including at least one of the following: Output authenticity information for at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, and an explanation of the authenticity information; Output authenticity information and a first control for at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, wherein the first control can be triggered to display an explanation of the authenticity information; Output authenticity information for at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, and, when the authenticity information is triggered, output an explanation of the authenticity information. Output authenticity information for at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, and response information for the target query information; Output authenticity information and a second control for at least a portion of the data frames and / or at least a portion of the region data in the target audio and video data, wherein the second control can be triggered to display the associated information of the authenticity information.
10. An electronic device comprising at least one processor and at least one processing model capable of running on said processor, said processing model being invoked by a target application to perform the following operations: In response to a target trigger operation, target multimedia data is identified and processed. The target multimedia data includes at least one data frame input to the target application or currently output by the target application. The target application is an application capable of performing the identification and processing or an application capable of calling a target program file to perform the identification and processing. Output the identification result for the target multimedia data, the identification result being able to indicate at least some data frames and / or at least some data in the target multimedia data the authenticity information.