Method, apparatus, device, and storage medium for identifying the integrity of video content

By extracting and splicing the audio features and text features of video files, identifying the completeness of short videos, the problem of low-efficiency review of video content in the existing technology is solved, and a more efficient and accurate review process is achieved.

CN112418011BActive Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202011237365.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-09
Publication Date
2025-05-27
Estimated Expiration
2040-11-09

AI Technical Summary

Technical Problem

When performing video content completeness review in the short video platform, the prior art is inefficient and relies on manual review, making it difficult to effectively identify and filter incomplete video content.

Method used

By extracting the audio and text features of the video file, splicing them and identifying them, the completeness of the video content is determined in combination with the characteristics of multiple dimensions, thereby improving the accuracy and efficiency of the audit.

Benefits of technology

It has achieved the accuracy and efficiency of video content completeness recognition, and can screen out incomplete video content more quickly and improve the video quality received by users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112418011B_ABST
    Figure CN112418011B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, device, and storage medium for identifying the integrity of video content, relating to the field of deep learning. An integrity identification model for video is constructed through artificial intelligence technology, and the function of identifying the integrity of a video is implemented by a computer device. The method includes: obtaining a video file and video release information of the video file, where the video release information represents the information provided when releasing the video content corresponding to the video file; separating audio data from the video file; extracting audio features from the audio data and text features from the video release information; splicing the audio features and the text features to obtain spliced features; and identifying the spliced features to obtain the integrity of the video content corresponding to the video file. By identifying the vector obtained by splicing the audio features and text features corresponding to the video file, the integrity of the video content is determined by integrating features from multiple dimensions, improving the accuracy of video integrity review.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A method for identifying the integrity of video content, characterized in that, the method is executed by a server, and the method includes: obtaining a video file and video release information of the video file, where the video release information represents information required to be provided by the uploader of the video file when publishing the video content corresponding to the video file in a video application; separating audio data from the video file; extracting audio features from the audio data and extracting text features from the video release information; concatenating the audio features and the text features to obtain concatenated features; identifying the concatenated features to obtain the integrity of the video content corresponding to the video file, where the integrity is used to indicate any one of a normal video, a truncated ending video, a non-truncated ending video, and a non-human voice ending video; the truncated ending video is used to indicate that the pronunciation of a single word is incomplete, the non-truncated ending video is used to indicate that the time interval between the ending moment of the last word and the video ending moment is less than a preset time interval, and the non-human voice ending video is used to indicate that the non-human voice suddenly ends and the non-human voice is incomplete; when the integrity of the video content corresponding to the video file is used to indicate that the video file is a normal video, recommending the video file with complete video content to a terminal.

2. The method according to claim 1, characterized in that, the identifying the concatenated features to obtain the integrity of the video content corresponding to the video file includes: invoking a video integrity identification model to identify the concatenated features to obtain the prediction probability that the video content corresponding to the video file belongs to complete video content; obtaining the integrity of the video content corresponding to the video file according to the prediction probability.

3. The method according to claim 2, characterized in that, the video integrity identification model is obtained in the following manner: obtaining a sample video file and sample video release information of the sample video file, where the sample video corresponding to the sample video file is labeled with video content integrity, and the sample video release information represents the information provided when publishing the video content corresponding to the sample video file; extracting sample audio features from the audio data corresponding to the sample video and extracting sample text features from the sample video release information; concatenating the sample audio features and the sample text features to obtain concatenated sample features; identifying the concatenated sample features to obtain the predicted integrity of the content of the sample video corresponding to the sample video file; training the video integrity identification model according to the predicted integrity of the content and the video content integrity labeled in the sample video to obtain a trained video integrity identification model.

4. The method according to claim 3, characterized in that, the training the video integrity identification model according to the predicted integrity of the content and the video content integrity labeled in the sample video to obtain a trained video integrity identification model includes: Calculate the error loss between the predicted content completeness and the video content completeness; Train the video completeness recognition model according to the error loss to obtain the trained video completeness recognition model.

5. The method according to claim 4, wherein, the calculating the error loss between the predicted content completeness and the video content completeness includes: Obtain the activation function corresponding to the video completeness recognition model; According to the activation function, the predicted content completeness and the video content completeness, obtain a cross-entropy loss function for binary classification; Calculate the error loss between the predicted content completeness and the video content completeness according to the cross-entropy loss function for binary classification.

6. The method according to claim 4, wherein, the training the video completeness recognition model according to the error loss to obtain the trained video completeness recognition model includes: Calculate the error loss through the cross-entropy loss function for binary classification, and the cross-entropy loss function for binary classification is obtained through the activation function corresponding to the video completeness recognition model, the predicted content completeness and the video content completeness; In response to the convergence of the error loss, obtain the weight matrix and the offset vector corresponding to the video completeness recognition model, where the weight matrix is used to characterize the influence degree of the sample video file on the output of the predicted content completeness of the video completeness recognition model, and the offset vector is used to characterize the deviation between the predicted content completeness and the video completeness; Obtain the trained video completeness recognition model according to the weight matrix and the offset vector.

7. The method according to any one of claims 1 to 6, wherein, the extracting audio features from the audio data and extracting text features from the video release information includes: Call an audio feature extraction model to extract the audio features from the audio data; Call a text feature extraction model to extract the text features from the video release information.

8. The method according to claim 7, wherein, the calling an audio feature extraction model to extract the audio features from the audio data includes: Call the Visual Geometry Group network model VGGish to extract the first audio features from the audio data; or, Extract the second audio features from the audio data through the Mel Frequency Cepstral Coefficient algorithm MFCC; or, Call the VGGish model to extract the first audio features from the audio data; extract the second audio features from the audio data through the MFCC algorithm.

9. The method according to claim 7, wherein, the video release information includes at least one of a video title, a video tag, and a user account; the calling a text feature extraction model to extract the text features from the video release information includes: In response to the video release information including the video title, a Bidirectional Encoder Representations from Transformers (BERT) model based on a transformer model is invoked to process the video title, obtaining first text features corresponding to the video title, where the video title is the video title corresponding to the video content in the video file; In response to the video release information including the video tags, the BERT model is invoked to process the video tags, obtaining second text features corresponding to the video tags, where the video tags are the categories to which the video content in the video file belongs; In response to the video release information including the user account, the BERT model is invoked to process the user account, obtaining third text features corresponding to the user account, where the user account is the user account that publishes the video content in the video file.

10. An apparatus for identifying the integrity of video content, characterized in that, the apparatus is provided by a server, and the apparatus includes: an acquisition module, configured to acquire a video file and video release information of the video file, where the video release information represents information required to be provided by the uploader of the video file when publishing the video content corresponding to the video file in a video application; a processing module, configured to separate audio data from the video file; a feature extraction module, configured to extract audio features from the audio data and extract text features from the video release information; the processing module, configured to splice the audio features and the text features to obtain spliced features; an identification module, configured to identify the spliced features to obtain the integrity of the video content corresponding to the video file, where the integrity is used to indicate any one of a normal video, a truncated ending video, a non-truncated ending video, and a non-human voice ending video for the video file; the truncated ending video is used to indicate that the pronunciation of a single word is incomplete, the non-truncated ending video is used to indicate that the time interval between the ending moment of the last word and the video ending moment is less than a preset time interval, and the non-human voice ending video is used to indicate that the non-human voice suddenly ends, making the non-human voice incomplete; When the integrity of the video content corresponding to the video file is used to indicate that the video file is a normal video, recommend the video file with complete video content to a terminal.

11. A computer device, characterized in that, the computer device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method for identifying the integrity of video content according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, at least one instruction, at least one program, a code set or an instruction set is stored in the readable storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the method for identifying the integrity of video content according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Video classification method and device, storage medium and server

    CN111209970A