Video detection method and device based on multi-modal data, equipment and storage medium
By using a multimodal data video detection method, image, voice, and text elements are automatically extracted and analyzed, solving the problem of low efficiency in manual review and achieving efficient and accurate review of AI-generated videos, thereby improving the content quality of video platforms.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2026-03-10
AI Technical Summary
Existing video detection methods rely on manual review, resulting in low efficiency and accuracy, especially for low-quality AI-generated videos, which negatively impacts user experience.
By extracting multimodal data (image, voice, and text) elements, using preset algorithms to detect video elements, and combining the detection results of multiple data modalities to determine the video generation method, automated review is achieved.
It improves the accuracy and efficiency of video detection, reduces the time and manpower costs of manual review, and ensures the quality of content displayed on video platforms.
Smart Images

Figure CN121640332A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of video processing in the field of artificial intelligence, and particularly relates to a video detection method and device based on multi-modal data, equipment and a storage medium. BACKGROUND
[0002] With the development of AIGC (AI Generated Content) technology, more and more videos are generated by AI (Artificial Intelligence). The quality of AI-generated videos is generally low, and the user's viewing experience is poor.
[0003] At present, the detection and review of videos on video platforms are mostly performed by humans, so as to filter out AI-synthesized videos. This method requires manual viewing of complete videos, and the detection process is time-consuming and laborious, and the accuracy and efficiency of video detection are low. SUMMARY
[0004] The present disclosure provides a video detection method and device based on multi-modal data, equipment and a storage medium.
[0005] According to a first aspect of the present disclosure, a video detection method based on multi-modal data is provided, comprising:
[0006] obtaining a to-be-detected video;
[0007] extracting video elements of multiple data modalities from the to-be-detected video; wherein the video elements represent the content expressed by the to-be-detected video;
[0008] performing detection processing on the to-be-detected video according to the video elements to obtain a detection result of the to-be-detected video; wherein the detection result represents the generation mode of the to-be-detected video.
[0009] According to a second aspect of the present disclosure, a video detection device based on multi-modal data is provided, comprising:
[0010] an obtaining unit configured to obtain a to-be-detected video;
[0011] an extracting unit configured to extract video elements of multiple data modalities from the to-be-detected video; wherein the video elements represent the content expressed by the to-be-detected video;
[0012] a detection unit configured to perform detection processing on the to-be-detected video according to the video elements to obtain a detection result of the to-be-detected video; wherein the detection result represents the generation mode of the to-be-detected video.
[0013] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0014] at least one processor; and
[0015] a memory communicatively connected to the at least one processor;
[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect of the present disclosure.
[0017] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, the computer instructions being used to cause the computer to perform the method according to the first aspect of the present disclosure.
[0018] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the steps of the method according to the first aspect of the present disclosure.
[0019] According to the technology of the present disclosure, the accuracy and efficiency of video detection are improved.
[0020] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0022] Figure 1 is a flowchart of a video detection method based on multi-modal data according to an embodiment of the present disclosure;
[0023] Figure 2 is a framework diagram of video detection according to an embodiment of the present disclosure;
[0024] Figure 3 is a flowchart of a video detection method based on multi-modal data according to an embodiment of the present disclosure;
[0025] Figure 4 is a determination diagram of a detection score corresponding to image data according to an embodiment of the present disclosure;
[0026] Figure 5 is a structural block diagram of a video detection device based on multi-modal data according to an embodiment of the present disclosure;
[0027] Figure 6 is a structural block diagram of a video detection device based on multi-modal data according to an embodiment of the present disclosure;
[0028] Figure 7 is a block diagram of an electronic device for implementing the video detection method based on multi-modal data according to an embodiment of the disclosure;
[0029] Figure 8 is a block diagram of an electronic device for implementing the video detection method based on multi-modal data according to an embodiment of the disclosure. DETAILED DESCRIPTION
[0030] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are presented for the purpose of illustration and description. They are intended to be considered as examples only, and thus are not intended to limit the scope of the disclosure. Those skilled in the art will recognize that various modifications and changes can be made to the embodiments described herein without departing from the scope and spirit of the disclosure. Also, the description below includes various details of the embodiments of the present disclosure to assist in understanding thereof. Thus, it will be apparent to those skilled in the art that various changes and modifications can be made therein without departing from the scope and spirit of the disclosure. For the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.
[0031] As a kind of information medium, short video is more and more popular among young people. Due to its rich media form, videos in the form of short videos and the like can provide an immersive experience, and therefore various institutions are focusing on videos, and use AIGC technology to obtain AI-generated videos. Unlike artificially fine-tuned videos, AI-generated videos have problems such as fixed audio, lack of dynamic elements in pictures, and illusion of content in scripts, and therefore it is necessary to detect and distinguish such videos to avoid excessive AI-generated videos affecting the video viewing experience of users.
[0032] Current video detection methods are mostly manually audited, but this approach requires staff to consider multiple dimensions such as scripts, pictures, and audio, and watch complete videos. In particular, for long videos, the auditing process is time-consuming and labor-intensive, and the precision and efficiency of video detection are low.
[0033] The present disclosure provides a video detection method and device based on multi-modal data, and a device and a storage medium, which are applied to the field of video processing in the field of artificial intelligence to improve the precision of video detection.
[0034] It should be noted that the model in this embodiment is not a model for a specific user, and cannot reflect the personal information of a specific user. It should be noted that the video in this embodiment comes from a public data set.
[0035] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.
[0036] To enable the reader to have a more profound understanding of the implementation principle of the present disclosure, the following Figures 1-8 The embodiments are further refined.
[0037] Figure 1 A flowchart of a video detection method based on multi-modal data according to an embodiment of the present disclosure can be executed by a video detection device based on multi-modal data. As shown in the figure, the method comprises the following steps: Figure 1
[0038] S101, obtaining a video to be detected.
[0039] Exemplarily, there are various video platforms on the market, and users can upload videos to the video platforms for everyone to watch. Each video platform is provided with an audit personnel, who can audit the videos uploaded by each user to determine whether the video is qualified, and display the qualified video in the video platform, and the unqualified video fails to be uploaded.
[0040] With the development of AI technology, more and more videos are synthesized by AI, i.e. AIGC generated patchwork videos, which have low quality. If the videos on the video platform are all AI generated videos, it will affect the viewing experience of platform consumers. Therefore, it is necessary to determine whether the video is AI generated or carefully made by human beings during the audit, and to determine the video that needs to be audited as the video to be detected.
[0041] The videos uploaded by each user to the video platform can be obtained in real time or at a fixed time as the video to be detected.
[0042] S102, extracting video elements of multiple data modalities from the video to be detected; wherein the video elements represent the content expressed by the video to be detected.
[0043] Exemplarily, data extraction is performed on the video to be detected, and the extracted data is determined as the video element. The video element can be in the form of multiple data modalities, for example, it can be an image, a video, an audio, a document, a text, a point cloud, etc. In this embodiment, the video element can be image data, voice data, and text data, etc. The video element can represent the content expressed in the video to be detected, for example, the video frames in the video can be extracted as image modal video elements, the subtitles in the video can be extracted as text modal video elements, and the audio in the video can be extracted as voice modal video elements.
[0044] The algorithm for data extraction can be preset, and different algorithms can be used for extraction of video elements of different data modalities. For example, one or more video frames can be extracted from the video as image data; a preset voice extraction algorithm can be used to extract voice data; and a preset text extraction algorithm can be used to extract text data. In this embodiment, the algorithm for data extraction is not specifically limited.
[0045] In this embodiment, the data modality of the video element is image data; the video element of multiple data modalities is extracted from the to-be-detected video, comprising: performing a segmentation processing on the to-be-detected video to obtain a video segment in the to-be-detected video; wherein the video segment represents part of the video content in the to-be-detected video; the video frame at a preset position in the video segment is determined as the video element with the data modality of image data.
[0046] Specifically, the data modality of the video element can be image data. When extracting image data from the to-be-detected video, the to-be-detected video can be first segmented to divide the to-be-detected video into multiple video segments, each video segment representing part of the video content in the to-be-detected video, and all the video segments being spliced to form the complete to-be-detected video. The segmentation can be performed according to the shots in the to-be-detected video, and a shot segmentation algorithm is pre-set, for example, the shot segmentation algorithm can be TransNet V2 algorithm. The shot segmentation algorithm is used to perform shot segmentation to divide the to-be-detected video into N video segments, where N is an integer.
[0047] One or more video frames can be extracted from the video segment as image data of the video element. The video frame at a preset position in the video segment can be determined as the video element with the data modality of image data. For example, the preset position can be the first, middle and last three positions, that is, for each video segment, the first, middle and last three video frames can be extracted as the video element with the data modality of image data.
[0048] The beneficial effect of such setting is that by performing segmentation and video frame extraction on the to-be-detected video, image data that can accurately express the content of the to-be-detected video can be obtained, comprehensive extraction of the video frames of the to-be-detected video is realized, the extraction accuracy of the image data is improved, and the detection accuracy of the to-be-detected video is further improved.
[0049] In this embodiment, the data modality of the video element is speech data; the video element of multiple data modalities is extracted from the to-be-detected video, comprising: extracting audio in the to-be-detected video according to a pre-set audio extraction algorithm, as the video element with the data modality of speech data.
[0050] Specifically, the data modality of the video element can be speech data. When extracting speech data from the to-be-detected video, an audio extraction algorithm can be pre-set, and the audio extraction algorithm is used to extract audio in the to-be-detected video, and the extracted audio is taken as the video element with the data modality of speech data. The extracted audio can be the background sound of the to-be-detected video. In this embodiment, the pre-set audio extraction algorithm is not specifically limited.
[0051] The beneficial effect of such an arrangement is that voice data in the to-be-detected video is extracted as a video element that needs to be detected, comprehensive detection of the to-be-detected video is achieved, and the detection accuracy of the video is improved.
[0052] In this embodiment, the data modality of the video element is text data; the video element of multiple data modalities is extracted from the to-be-detected video, including: performing voice-to-text recognition processing on the voice data in the video element to obtain the video element with text data as the data modality.
[0053] Specifically, the data modality of the video element can be text data, and the text data can be the script of the to-be-detected video. The script can be embodied as the content of the subtitles of the to-be-detected video, or can be embodied as the content of the audio of the to-be-detected video.
[0054] When extracting the text data from the to-be-detected video, it can be determined whether there are subtitles in the to-be-detected video, and the subtitles of the to-be-detected video are extracted as text data in the video element. The audio of the to-be-detected video can also be extracted first, and the audio is converted into text according to a preset voice recognition algorithm, and the converted text is taken as the video element with text data as the data modality. For example, the preset voice recognition algorithm can be an ASR (Automatic Speech Recognition) algorithm, and the voice-to-text recognition processing is performed on the voice data in the video element by using the ASR algorithm to obtain the text data.
[0055] The beneficial effect of such an arrangement is that by extracting the text data, the script content of the to-be-detected video can be detected, the comprehensiveness of video detection is improved, and the accuracy of video detection is further improved.
[0056] S103, according to the video element, performing detection processing on the to-be-detected video to obtain a detection result of the to-be-detected video; wherein the detection result represents a generation mode of the to-be-detected video.
[0057] Exemplarily, after obtaining the video elements of multiple modalities, detection processing is performed on each kind of video element, and the detection results of various video elements are comprehensively obtained to obtain the detection result of the to-be-detected video. The detection result can represent the generation mode of the to-be-detected video. That is, it can be determined whether the to-be-detected video is artificially generated or AI generated. For example, for each video element, it can be determined whether the video element is artificially generated or AI generated. If all video elements are artificially generated, it is determined that the detection result of the to-be-detected video is artificially generated. If at least one video element is AI generated, it is determined that the detection result of the to-be-detected video is AI generated.
[0058] The detection rule of the video element can be preset, and different video elements can correspond to different detection rules. For example, for image data, the detection rule can be whether the image data contains a dynamic object; if yes, the image data is artificially made; if no, the image data is AI generated. For voice data, the detection rule can be whether the audio of the voice data is relatively fixed, if yes, the voice data is AI generated; if no, the voice data is artificially made. For text data, the detection rule can be whether the video script corresponding to the text data has a script content illusion, if yes, the text data is AI generated; if no, the voice data is artificially made. The script content illusion can refer to problems such as incoherence or logical contradiction between the front and back of the script. In this embodiment, the preset detection rule is not specifically limited.
[0059] If it is determined that the to-be-detected video is not AI generated, it is determined that the to-be-detected video is uploaded successfully, and the to-be-detected video is displayed on the video platform for users to watch; if it is determined that the to-be-detected video is AI generated, it is determined that the to-be-detected video fails to be uploaded, and the uploader of the to-be-detected video is prompted to re-upload a new video, thereby improving the quality of the videos in the video platform.
[0060] Figure 2 A framework diagram of video detection provided in this embodiment. Figure 2 In this embodiment, the to-be-detected video is first input, and the to-be-detected video is detected from three aspects, which are image aspect, text aspect, and voice aspect. For the image aspect, the to-be-detected video is lens cut, and image data is obtained after cutting, that is, the picture of the video is obtained, and the picture is detected; for the text aspect, the script of the video is extracted through an ASR algorithm, and the script is detected; for the voice aspect, the audio of the to-be-detected video is extracted, and the audio is detected. The overall score of the to-be-detected video is obtained by comprehensively detecting the three aspects, that is, the detection result is obtained.
[0061] In this embodiment of the disclosure, the generation mode of the to-be-detected video is determined automatically by acquiring the to-be-detected video. A plurality of data modal video elements can be extracted from the to-be-detected video, for example, image data, voice data, text data, etc. can be extracted. According to these data modal video elements, the to-be-detected video is detected and processed to obtain the detection result of the to-be-detected video, thereby determining the generation mode of the to-be-detected video. The automatic review of the to-be-detected video is realized, manpower and time are saved, and the video review efficiency is improved. By comprehensively detecting a plurality of data modal video elements, comprehensive detection of the to-be-detected video is realized, and the accuracy of video detection is improved.
[0062] Figure 3 A flowchart of a video detection method based on multi-modal data provided in this embodiment of the disclosure.
[0063] In this embodiment, the detection process of the video to be detected is performed based on video elements to obtain the detection result of the video to be detected, including: performing detection processing on the video to be detected based on video elements to obtain the detection score corresponding to the video elements; wherein, the detection score represents the rationality of the video elements in the video to be detected; and determining the detection result of the video to be detected based on the detection score corresponding to each video element.
[0064] like Figure 3 As shown, the method includes the following steps:
[0065] S301. Obtain the video to be tested.
[0066] For example, this step can refer to step S101 above, and will not be repeated here.
[0067] S302. Extract video elements of multiple data modalities from the video to be detected; wherein, the video elements represent the content expressed by the video to be detected.
[0068] For example, this step can refer to step S102 above, and will not be repeated here.
[0069] S303. Based on the video elements, perform detection processing on the video to be detected to obtain the detection score corresponding to the video elements; wherein, the detection score represents the reasonableness of the video elements in the video to be detected.
[0070] For example, for each video element, and for each video to be detected, each video element can correspond to a detection score, which can characterize the reasonableness of the video element. For instance, if the video element is image data, then the image data can be detected to obtain the detection score corresponding to the image data.
[0071] Each video element can have its own corresponding detection method, and targeted detection can be performed based on the detection method corresponding to the video element. For example, for text data, it can detect whether there are too many duplicate texts in the text data; the more duplicates, the lower the detection score. In this embodiment, a neural network model can also be pre-built for different video elements. The video elements are input into the preset model, and the corresponding detection score is output. In this embodiment, the structure of the preset model is not specifically limited.
[0072] In this embodiment, the data modality of video elements is image data, and the image data is a video frame in the video to be detected. Based on the video elements, detection processing is performed on the video to be detected to obtain the detection score corresponding to the video elements. This includes: performing semantic recognition processing on the image data of the video to be detected to obtain semantic information of the image data, which is the first semantic information; obtaining the text data corresponding to the video frame to which the image data belongs; wherein the text data represents the text of the audio played by the video frame; performing semantic recognition processing on the text data corresponding to the video frame to which the image data belongs to obtain the second semantic information; and determining the detection score corresponding to the image data of the video to be detected based on the first semantic information and the second semantic information.
[0073] Specifically, for video elements in image data, the image data can be video frames from the video to be detected. One or more video frames can be identified as image data within the video elements. For example, the video to be detected can be first divided into multiple video segments, and then video frames can be extracted from each video segment as video elements of the image data.
[0074] An image semantic detection algorithm can be pre-set. Based on the pre-set image semantic detection algorithm, the extracted image data is subjected to semantic recognition processing to obtain the semantic information of the image data, which serves as the first semantic information. If the image data consists of multiple video frames, each video frame corresponds to its own first semantic information.
[0075] For each video frame in the image data, determine the text data corresponding to the video frame to which the image data belongs. The text data can represent the text of the audio played in the video frame; for example, the audio corresponding to the video frame can be obtained and converted into text data. If the video frame displays subtitles, these subtitles can be directly extracted as the text data corresponding to that video frame. Semantic recognition processing is then performed on this text data; for example, a pre-defined text semantic recognition algorithm can be used to obtain the semantic information corresponding to the text data, which serves as secondary semantic information.
[0076] The detection score of the image data of the video to be detected is determined by comparing the first semantic information and the second semantic information based on their consistency. The higher the consistency between the first semantic information and the second semantic information, the higher the detection score of the image data of the video to be detected. If the image data consists of multiple video frames, the detection score of the entire image data can be determined based on the detection scores of those multiple video frames. For example, the detection scores of multiple video frames can be added together or the average value can be calculated as the detection score of the image data.
[0077] The advantage of this setup is that AI-synthesized videos may have mismatched images and text. By performing image-text consistency checks on the image data, the reasonableness of the images and text in the video to be detected can be determined, thereby improving the accuracy of video detection.
[0078] In this embodiment, semantic recognition processing is performed on the image data of the video to be detected to obtain semantic information of the image data, which is the first semantic information. This includes: if image content of a preset format is detected from the image data, then semantic recognition processing is performed on the image content of the preset format to obtain semantic information of the image content of the preset format, which is the first semantic information.
[0079] Specifically, it can be determined whether image content in a preset format exists in the image data. For example, the preset format could be an emoji format or a fancy font format; that is, it can be determined whether emojis or fancy fonts exist in the video frames of the image data. A preset object detection algorithm can be used to detect image content in the preset format. For example, the preset object detection algorithm could be the YOLO (You Only Look Once) algorithm. If image content in the preset format does not exist, it is determined that the content in the image data is insufficient, and a lower detection score can be assigned to the image data. Alternatively, semantic recognition processing can be performed directly on the image data to obtain initial semantic information before calculating subsequent detection scores.
[0080] If image content in a preset format is identified, semantic recognition processing can be performed on the image content in the preset format, and the semantic information of the image content in the preset format can be determined as the first semantic information. For example, if the preset format is an emoji, semantic recognition can be performed on the emoji to determine the meaning or emotion expressed by the emoji, which is used as the first semantic information; if the preset format is cursive text, the content of the cursive text can be extracted, and semantic recognition can be performed on the text in the cursive text to obtain the first semantic information.
[0081] The advantage of this setup is that it allows for personalized detection of specific elements in image data, thereby obtaining important information from the image data, improving the semantic accuracy of image information, and ultimately improving the detection accuracy of video.
[0082] In this embodiment, there are multiple preset formats; based on the video elements, the video to be detected is processed to obtain the detection score corresponding to the video elements, including: if image content of a preset format is detected from the image data, then the detection score corresponding to the image data of the video to be detected is determined according to the number of types of image content of the preset format in the image data.
[0083] Specifically, multiple preset formats can be set. For example, preset formats may include cursive text, emoticons, and special prompt boxes. Special prompt boxes can be target boxes of preset shapes. Based on a preset target detection algorithm, it is determined whether image content of preset formats exists in each video frame corresponding to the image data, and which preset formats exist. For example, for each video segment, the first, middle, and last three video frames are extracted as image data, and it is determined which preset format image content exists in these video frames.
[0084] For a single video frame, if it contains image content in a preset format, the detection score is determined based on the number of preset formats present in that frame. This detection score characterizes the richness of the image data. For example, if the video frame contains three preset formats—cursive text, emoticons, and special prompts—then the detection score can be relatively high.
[0085] For each video frame in the image data, these three preset formats can be scored separately. Combining these three scores yields the detection score for the video frame. Then, combining this score with the detection scores of all video frames in the image data yields the detection score for the image data itself. For example, to determine if cursive text exists in a video frame, if not, the first score for the cursive text is 0; if it exists, the content of the cursive text and the corresponding text data for that video frame are extracted, and a consistency comparison is performed between the cursive text content and the corresponding text data to obtain the first score. Similarly, to determine if an emoji exists in a video frame, if not, the second score for the emoji is 0; if it exists, semantic recognition is performed on both the emoji and the corresponding text data, resulting in a consistency comparison and the second score. Finally, to determine if a special prompt box exists in a video frame, if not, the third score for the special prompt box is 0; if it exists, the third score for the special prompt box is 1. The first, second, and third scores are then weighted and summed to obtain the detection score for the video frame. Then, the detection scores of each video frame are summed or averaged to obtain the detection score corresponding to the image data.
[0086] The advantage of this setup is that AI-synthesized videos often have low image richness. By recognizing image content in a preset format within the image data, the richness of the video to be detected is improved, thus enhancing the accuracy of video detection.
[0087] In this embodiment, the data modality of the video element is speech data; based on the video element, the video to be detected is processed to obtain the detection score corresponding to the video element, including: performing feature extraction processing on the speech data of the video to be detected to obtain the frequency domain features corresponding to the speech data of the video to be detected; wherein, the frequency domain features characterize the spectral fluctuation of the speech data; based on the frequency domain features corresponding to the speech data of the video to be detected, the detection score corresponding to the speech data of the video to be detected is determined.
[0088] Specifically, for video elements of speech data, an audio feature extraction algorithm can be preset to perform feature extraction processing on the speech data of the video to be detected, obtaining the frequency domain features corresponding to the speech data of the video to be detected. The frequency domain features can characterize the spectral fluctuations of the speech data. For example, they can represent the timbre, pitch fluctuations, etc. of the speech data. In this embodiment, the audio feature extraction algorithm is not specifically limited.
[0089] Based on the frequency domain features of the speech data in the video to be tested, the detection score corresponding to the speech data is determined. Videos created by real people consist of multiple audio sources, resulting in rich frequency domain features. AI-generated videos tend to use the same type of TTS (Text-to-Speech) audio, and the speech generated by TTS is often flat and lacks richness. Therefore, the detection score corresponding to the speech data can be determined through frequency domain features. The richer the frequency domain features, the higher the detection score.
[0090] You can preset a PANNS (Pretrained Audio Neural Networks) model, input speech data into the PANANS model, and output the detection score corresponding to the speech data.
[0091] The advantage of this setup is that by extracting frequency domain features, it can determine whether the audio of the video to be detected is linear and straightforward, thereby determining the richness of the speech data and improving the accuracy of video detection.
[0092] In this embodiment, feature extraction processing is performed on the audio data of the video to be detected to obtain the frequency domain features corresponding to the audio data of the video to be detected. This includes: segmenting the audio data of the video to be detected according to a preset time interval to obtain multiple audio segments; wherein, each audio segment represents part of the audio data; and performing feature extraction processing on each audio segment to obtain the frequency domain features corresponding to each audio segment.
[0093] Specifically, the audio data of the video to be tested is segmented into multiple audio segments. Each audio segment represents a portion of the audio data, and the combination of all audio segments constitutes the complete audio data. A time interval can be preset, and the audio data of the video to be tested is segmented into equal time intervals according to the preset time interval, thereby obtaining multiple audio segments.
[0094] For each speech segment, frequency domain features are extracted. By combining the frequency domain features of all speech segments, a detection score is obtained for the corresponding speech data. For example, the frequency domain features of all speech segments can be concatenated into a sequence, input into the PAANS model, and the output will be the detection score for the corresponding speech data.
[0095] The advantage of this setup is that by segmenting the speech data, the accuracy of frequency domain feature extraction can be improved, thereby gaining a more comprehensive understanding of the richness of the speech data and improving the accuracy of video detection.
[0096] In this embodiment, the data modality of the video element is text data. Based on the video element, the video to be detected is processed to obtain the detection score corresponding to the video element, including: performing feature extraction processing on the text data of the video to be detected to obtain the semantic features corresponding to the text data of the video to be detected; wherein, the semantic features characterize the logic and content quality of the text data; and determining the detection score corresponding to the text data of the video to be detected based on the semantic features corresponding to the text data of the video to be detected.
[0097] Specifically, for video elements with text data, it can be determined whether the text data was automatically generated by a large model. Text created by real people is high-quality, logically clear, and rich in precise information, while text generated by a large model is rejected. The criteria for rejection include large-model-generated text containing excessive repetitive information, unclear subjects, incoherent sentences, too many viewpoints, and logical contradictions.
[0098] Feature extraction can be performed on the text data of the video to be tested to extract semantic features. Semantic features characterize the logic and content quality of the text data, i.e., whether the text data is logically clear and grammatically correct. Based on the semantic features corresponding to the text data of the video to be tested, a detection score is determined. If the quality of the text data is low, the detection score will be low.
[0099] The advantage of this setup is that by extracting features from the text data, the quality of the text in the video to be detected can be determined, thereby determining whether the text was generated by a large model and improving the accuracy of video detection.
[0100] S304. Determine the detection result of the video to be detected based on the detection score corresponding to each video element.
[0101] For example, after obtaining the detection scores corresponding to each video element, these detection scores are combined to obtain the final detection result of the video to be detected, that is, to determine whether the video to be detected is a human-made video or an AI-synthesized video. For example, the lowest value can be determined from these detection scores. If the lowest value is greater than a preset score threshold, the video to be detected is determined to be a human-made video; otherwise, it is an AI-synthesized video.
[0102] In this embodiment, comprehensive identification is performed from the aspects of images, text, and audio to determine the generation method of the video, thereby improving the detection accuracy of the video and thus improving the efficiency and accuracy of video review on the platform.
[0103] In this embodiment, the detection result of the video to be detected is determined based on the detection score corresponding to each video element, including: determining the target score of the video to be detected based on the detection score corresponding to the image data, the detection score corresponding to the audio data, and the detection score corresponding to the text data; wherein, the target score represents the probability that the video to be detected is generated by artificial intelligence; and the detection result of the video to be detected is determined based on the target score of the video to be detected.
[0104] Specifically, the detection scores for image data, speech data, and text data are combined. For example, these three scores can be added together or averaged, and the result is used as the target score for the video to be detected. The target score characterizes the probability that the video to be detected was generated by AI, that is, the probability that the video to be detected was generated by a human. The higher the target score, the more likely the video to be generated by a human; the lower the target score, the more likely the video to be generated by AI.
[0105] Alternatively, weights can be preset for image data (first weight), speech data (second weight), and text data (third weight). A weighted sum is calculated based on the detection scores for image data, speech data, and text data, along with the preset weights, and the result is used as the target score. A pre-set score threshold is then compared to the target score. If the target score is greater than the preset threshold, the detection result is considered human-generated; if the target score is less than or equal to the preset threshold, the detection result is considered AI-generated.
[0106] The advantage of this setup is that by integrating the detection scores of multiple modalities, the final detection result is obtained, thus improving the accuracy of video detection.
[0107] Figure 4 This is a schematic diagram illustrating the determination of the detection score corresponding to the image data. Figure 4 The system contains three video clips. The first frame of each clip can be extracted as image data, designated as Video Frame 1, Video Frame 2, and Video Frame 3. The text of the first frame of each clip is then obtained: Video Frame 1 corresponds to Text 1, Video Frame 2 to Text 2, and Video Frame 3 to Text 3. A pre-defined Vision Transformer (VIT) network is used to extract image features from the image data, yielding image feature vectors (the first semantic information). Similarly, text features are extracted from the video frames, yielding text feature vectors (the second semantic information). Based on these image feature vectors, an image feature sequence is formed. The text feature vectors are also serialized to form a text feature sequence. Both the image and text feature sequences are then input into a pre-defined large language model, which performs image-text consistency and relevance detection, outputting a detection score for each image data. Alternatively, after inputting the image data into the VIT, the VIT output can be input into a Multi-Layer Perceptron (MLP) to obtain another image feature sequence.
[0108] In this embodiment, the generation method of the video to be detected is automatically determined by acquiring the video to be detected. Multiple data modalities of video elements can be extracted from the video to be detected, such as image data, audio data, and text data. Based on these data modalities of video elements, the video to be detected is processed to obtain the detection result, thereby determining the generation method of the video to be detected. This achieves automatic review of the video to be detected, saving manpower and time, and improving video review efficiency. By integrating video elements of multiple data modalities, comprehensive detection of the video to be detected is achieved, improving the accuracy of video detection.
[0109] Figure 5 This is a structural block diagram of a video detection device based on multimodal data, provided in an embodiment of this disclosure. For ease of explanation, only the parts relevant to the embodiments of this disclosure are shown. (Refer to...) Figure 5 The video detection device 500 based on multimodal data includes: an acquisition unit 501, an extraction unit 502, and a detection unit 503.
[0110] Acquisition unit 501 is used to acquire the video to be detected;
[0111] Extraction unit 502 is used to extract video elements of multiple data modalities from the video to be detected; wherein the video elements represent the content expressed by the video to be detected;
[0112] The detection unit 503 is used to perform detection processing on the video to be detected based on the video elements to obtain the detection result of the video to be detected; wherein the detection result represents the generation method of the video to be detected.
[0113] Figure 6 A structural block diagram of a video detection device based on multimodal data provided in this disclosure embodiment is shown below. Figure 6 As shown, the video detection device 600 based on multimodal data includes an acquisition unit 601, an extraction unit 602, and a detection unit 603. The detection unit 603 includes an element detection module 6031 and a result determination module 6032.
[0114] The element detection module 6031 is used to perform detection processing on the video to be detected based on the video elements, and obtain the detection score corresponding to the video elements; wherein, the detection score represents the rationality of the video elements in the video to be detected;
[0115] The result determination module 6032 is used to determine the detection result of the video to be detected based on the detection score corresponding to each video element.
[0116] In one example, the data modality of the video element is image data, and the image data is a video frame in the video to be detected; the element detection module 6031 includes:
[0117] The first determining submodule is used to perform semantic recognition processing on the image data of the video to be detected to obtain the semantic information of the image data, which is the first semantic information.
[0118] The text acquisition submodule is used to acquire the text data corresponding to the video frame to which the image data belongs; wherein, the text data represents the text of the audio played by the video frame;
[0119] The second determining submodule is used to perform semantic recognition processing on the text data corresponding to the video frame to which the image data belongs, and obtain the second semantic information;
[0120] The image detection submodule is used to determine the detection score corresponding to the image data of the video to be detected based on the first semantic information and the second semantic information.
[0121] In one example, the first determined submodule is specifically used for:
[0122] If image content of a preset format is detected from the image data, then semantic recognition processing is performed on the image content of the preset format to obtain the semantic information of the image content of the preset format, which is the first semantic information.
[0123] In one example, there are multiple preset formats; the element detection module 6031 includes:
[0124] The format determination submodule is used to determine the detection score corresponding to the image data of the video to be detected based on the number of types of image content in the preset format detected from the image data.
[0125] In one example, the data modality of the video element is audio data; the element detection module 6031 includes:
[0126] The frequency domain determination submodule is used to perform feature extraction processing on the speech data of the video to be detected to obtain the frequency domain features corresponding to the speech data of the video to be detected; wherein, the frequency domain features characterize the spectral fluctuation of the speech data;
[0127] The speech detection submodule is used to determine the detection score corresponding to the speech data of the video to be detected based on the frequency domain features corresponding to the speech data of the video to be detected.
[0128] In one example, the frequency domain determination submodule is specifically used for:
[0129] According to a preset time interval, the audio data of the video to be detected is segmented to obtain multiple audio segments; wherein, the audio segment represents part of the audio data;
[0130] Feature extraction is performed on each speech segment to obtain the frequency domain features corresponding to each speech segment.
[0131] In one example, the data modality of the video element is text data; the element detection module 6031 includes:
[0132] The semantic extraction submodule is used to perform feature extraction processing on the text data of the video to be detected to obtain the semantic features corresponding to the text data of the video to be detected; wherein, the semantic features characterize the logic and content quality of the text data;
[0133] The text detection submodule is used to determine the detection score corresponding to the text data of the video to be detected based on the semantic features corresponding to the text data of the video to be detected.
[0134] In one example, the result determination module 6032 includes:
[0135] The target determination submodule is used to determine the target score of the video to be detected based on the detection score corresponding to the image data, the detection score corresponding to the audio data, and the detection score corresponding to the text data; wherein, the target score represents the probability that the video to be detected is generated by artificial intelligence.
[0136] The result determination submodule is used to determine the detection result of the video to be detected based on the target score of the video to be detected.
[0137] In one example, the data modality of the video element is image data; the extraction unit 602 includes:
[0138] The video segmentation module is used to segment the video to be detected to obtain video segments from the video to be detected; wherein, the video segments represent a portion of the video content in the video to be detected.
[0139] The image determination module is used to determine the video frame at a preset position in the video segment as a video element with image data as its data modality.
[0140] In one example, the data modality of the video element is audio data; the extraction unit 602 includes:
[0141] The voice determination module is used to extract the audio from the video to be detected according to a preset audio extraction algorithm, and to identify video elements whose data modality is voice data.
[0142] In one example, the data modality of the video element is text data; extraction unit 602 includes:
[0143] The text determination module is used to perform speech-to-text recognition processing on the audio data in the video element to obtain the video element with text data as its data modality.
[0144] According to embodiments of this disclosure, this disclosure also provides an electronic device.
[0145] Figure 7 A structural block diagram of an electronic device provided in this disclosure embodiment, such as... Figure 7 As shown, the electronic device 700 includes: at least one processor 702; and a memory 701 communicatively connected to the at least one processor 702; wherein the memory stores instructions executable by the at least one processor 702, the instructions being executed by the at least one processor 702 to enable the at least one processor 702 to execute the video detection method based on multimodal data of this disclosure.
[0146] The electronic device 700 also includes a receiver 703 and a transmitter 704. The receiver 703 is used to receive instructions and data sent by other devices, and the transmitter 704 is used to send instructions and data to external devices.
[0147] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0148] According to embodiments of this disclosure, this disclosure also provides a computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program causing the electronic device to perform the scheme provided in any of the above embodiments.
[0149] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0150] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0151] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0152] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a video detection method based on multimodal data. For example, in some embodiments, the video detection method based on multimodal data can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the video detection method based on multimodal data described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform a video detection method based on multimodal data by any other suitable means (e.g., by means of firmware).
[0153] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0154] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0155] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0156] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0157] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0158] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.
[0159] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0160] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A video detection method based on multi-modal data, comprising: obtaining a video to be detected; extracting video elements of multiple data modalities from the video to be detected; wherein the video elements represent the content expressed by the video to be detected; performing detection processing on the video to be detected according to the video elements to obtain a detection result of the video to be detected; wherein the detection result represents the generation mode of the video to be detected.
2. The method of claim 1, wherein, The detection processing on the video to be detected according to the video elements to obtain a detection result of the video to be detected comprises: performing detection processing on the video to be detected according to the video elements to obtain a detection score corresponding to the video elements; wherein the detection score represents the rationality of the video elements in the video to be detected; determining the detection result of the video to be detected according to the detection scores corresponding to the video elements.
3. The method of claim 2, wherein, The data modality of the video elements is image data, and the image data is a video frame in the video to be detected; the detection processing on the video to be detected according to the video elements to obtain a detection score corresponding to the video elements comprises: performing semantic recognition processing on the image data of the video to be detected to obtain semantic information of the image data, which is first semantic information; obtaining text data corresponding to the video frame to which the image data belongs; wherein the text data represents the script of the audio played by the video frame; performing semantic recognition processing on the text data corresponding to the video frame to which the image data belongs to obtain second semantic information; determining the detection score corresponding to the image data of the video to be detected according to the first semantic information and the second semantic information.
4. The method of claim 3, wherein, The semantic recognition processing on the image data of the video to be detected to obtain the semantic information of the image data, which is first semantic information, comprises: if a preset format of image content is detected from the image data, performing semantic recognition processing on the preset format of image content to obtain semantic information of the preset format of image content, which is the first semantic information.
5. The method of claim 4, wherein, The preset format has multiple types. The detection processing on the video to be detected according to the video elements to obtain a detection score corresponding to the video elements comprises: if a preset format of image content is detected from the image data, determining the detection score corresponding to the image data of the video to be detected according to the number of types of the preset format of image content in the image data.
6. The method of any one of claims 2-5, wherein, The data modality of the video elements is voice data; the detection processing on the video to be detected according to the video elements to obtain a detection score corresponding to the video elements comprises: performing feature extraction processing on the voice data of the video to be detected to obtain a frequency domain feature corresponding to the voice data of the video to be detected; wherein the frequency domain feature represents the spectral fluctuation of the voice data; determining the detection score corresponding to the voice data of the video to be detected according to the frequency domain feature corresponding to the voice data of the video to be detected.
7. The method of claim 6, wherein, The feature extraction processing is performed on the voice data of the to-be-detected video to obtain frequency domain features corresponding to the voice data of the to-be-detected video, and the feature extraction processing comprises: According to a preset time interval, the voice data of the to-be-detected video is segmented to obtain a plurality of voice segments; wherein the voice segments represent part of the voice data; The feature extraction processing is performed on each voice segment to obtain frequency domain features corresponding to each voice segment.
8. The method of any one of claims 2-7, wherein, The data mode of the video element is text data; the detection processing is performed on the to-be-detected video according to the video element to obtain a detection score corresponding to the video element, and the detection processing comprises: The feature extraction processing is performed on the text data of the to-be-detected video to obtain semantic features corresponding to the text data of the to-be-detected video; wherein the semantic features represent the logical and content quality of the text data; According to the semantic features corresponding to the text data of the to-be-detected video, the detection score corresponding to the text data of the to-be-detected video is determined.
9. The method of claim 8, wherein, According to the detection score corresponding to each video element, the detection result of the to-be-detected video is determined, and the detection result comprises: According to the detection score corresponding to the image data of the to-be-detected video, the detection score corresponding to the voice data of the to-be-detected video, and the detection score corresponding to the text data of the to-be-detected video, a target score of the to-be-detected video is determined; wherein the target score represents the possibility that the to-be-detected video is generated by artificial intelligence; According to the target score of the to-be-detected video, the detection result of the to-be-detected video is determined.
10. The method of any one of claims 1-9, wherein, The data mode of the video element is image data; the video element of multiple data modes is extracted from the to-be-detected video, and the video element comprises: The to-be-detected video is segmented to obtain video segments in the to-be-detected video; wherein the video segments represent part of the video content in the to-be-detected video; The video frame at a preset position in the video segment is determined as a video element with image data as the data mode.
11. The method of any one of claims 1-10, wherein, The data mode of the video element is voice data; the video element of multiple data modes is extracted from the to-be-detected video, and the video element comprises: According to a preset audio extraction algorithm, audio in the to-be-detected video is extracted to be a video element with voice data as the data mode.
12. The method of claim 11, wherein, The data mode of the video element is text data; the video element of multiple data modes is extracted from the to-be-detected video, and the video element comprises: The voice data in the video element is recognized by voice-to-text processing to obtain a video element with text data as the data mode.
13. A video detection device based on multi-modal data, comprising: An acquisition unit configured to acquire a to-be-detected video; An extraction unit configured to extract video elements of multiple data modes from the to-be-detected video; wherein the video elements represent the content expressed by the to-be-detected video; A detection unit configured to perform detection processing on the to-be-detected video according to the video elements to obtain a detection result of the to-be-detected video; wherein the detection result represents the generation mode of the to-be-detected video.
14. The apparatus of claim 13, wherein, The detection unit comprises: An element detection module is configured to perform detection processing on the video to be detected according to the video element, so as to obtain a detection score corresponding to the video element; wherein the detection score represents rationality of the video element in the video to be detected. A result determination module is configured to determine a detection result of the video to be detected according to the detection score corresponding to each video element.
15. The apparatus of claim 14, wherein, The data mode of the video element is image data, and the image data is a video frame in the video to be detected. The element detection module comprises: A first determination sub-module is configured to perform semantic recognition processing on the image data of the video to be detected, so as to obtain semantic information of the image data, which is first semantic information. A text acquisition sub-module is configured to acquire text data corresponding to a video frame to which the image data belongs; wherein the text data represents a script of audio played by the video frame. A second determination sub-module is configured to perform semantic recognition processing on the text data corresponding to the video frame to which the image data belongs, so as to obtain second semantic information. An image detection sub-module is configured to determine a detection score corresponding to the image data of the video to be detected according to the first semantic information and the second semantic information.
16. The apparatus of claim 15, wherein, The first determination sub-module is specifically configured to: If a preset format of image content is detected from the image data, perform semantic recognition processing on the preset format of image content, so as to obtain semantic information of the preset format of image content, which is the first semantic information.
17. The apparatus of claim 16, wherein, The preset format has multiple types. The element detection module comprises: A format determination sub-module is configured to, if a preset format of image content is detected from the image data, determine a detection score corresponding to the image data of the video to be detected according to a type number of the preset format of image content in the image data.
18. The apparatus of any of claims 14-17, wherein, The data mode of the video element is voice data; and the element detection module comprises: A frequency domain determination sub-module is configured to perform feature extraction processing on the voice data of the video to be detected, so as to obtain a frequency domain feature corresponding to the voice data of the video to be detected; wherein the frequency domain feature represents a frequency spectrum fluctuation of the voice data. A voice detection sub-module is configured to determine a detection score corresponding to the voice data of the video to be detected according to the frequency domain feature corresponding to the voice data of the video to be detected.
19. The apparatus of claim 18, wherein, The frequency domain determination sub-module is specifically configured to: Perform segmentation processing on the voice data of the video to be detected according to a preset time interval, so as to obtain a plurality of voice segments; wherein the voice segment represents part of the voice data. Perform feature extraction processing on each voice segment, so as to obtain a frequency domain feature corresponding to each voice segment.
20. The apparatus of any of claims 14-19, wherein, The data mode of the video element is text data; and the element detection module comprises: A semantic extraction sub-module is configured to perform feature extraction processing on the text data of the video to be detected, so as to obtain a semantic feature corresponding to the text data of the video to be detected; wherein the semantic feature represents a logic and a quality of content of the text data. A text detection sub-module is configured to determine a detection score corresponding to the text data of the video to be detected according to the semantic feature corresponding to the text data of the video to be detected.
21. The apparatus of claim 20, wherein, The result determination module comprises: A target determination sub-module is configured to determine a target score of the to-be-detected video according to a detection score corresponding to image data of the to-be-detected video, a detection score corresponding to voice data of the to-be-detected video, and a detection score corresponding to text data of the to-be-detected video, wherein the target score represents a possibility that the to-be-detected video is artificially generated. A result determination sub-module is configured to determine a detection result of the to-be-detected video according to the target score of the to-be-detected video.
22. The apparatus of any of claims 13-21, wherein, The data modality of the video element is image data; the extraction unit comprises: A video segmentation module is configured to perform segmentation processing on the to-be-detected video to obtain a video segment in the to-be-detected video, wherein the video segment represents part of video content in the to-be-detected video. An image determination module is configured to determine a video frame at a preset position in the video segment as a video element with image data as the data modality.
23. The apparatus of any of claims 13-22, wherein, The data modality of the video element is voice data; the extraction unit comprises: A voice determination module is configured to extract audio in the to-be-detected video according to a preset audio extraction algorithm to obtain a video element with voice data as the data modality.
24. The apparatus of claim 23, wherein, The data modality of the video element is text data; the extraction unit comprises: A text determination module is configured to perform voice-to-text recognition processing on voice data in the video element to obtain a video element with text data as the data modality.
25. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-12.
26. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-12.
27. A computer program product, wherein, The computer program is executed by the processor to implement the steps of the method of any one of claims 1-12.
Citation Information
Patent Citations
Pretraining multi-modal model-based forged video detection method and system
CN114782858A
Video detection method, device and equipment and computer readable storage medium
CN114817636A
Video detection method and training method and device of video detection model
CN114842399A
Video detection method and device, equipment and storage medium
CN116823726A
Counterfeit content detection method, related device and storage medium
CN117409344A