AI-driven convergence media digital audio and video content security auditing method

Through AI-driven multimodal feature extraction and fusion technology, combined with user behavior analysis and scene recognition, the problem of low efficiency and insufficient accuracy of digital audio-visual content review on integrated media platforms is solved, and intelligent and dynamic review of video, audio, and graphic content is realized to adapt to diverse content review needs.

CN120455732APending Publication Date: 2025-08-08广西日报社
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510590507.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The digital audio-visual content review methods of existing integrated media platforms mainly rely on single-modal data analysis, resulting in misjudgment and low efficiency, making it difficult to meet the needs of large-scale real-time audits, and lacks the ability to comprehensively analyze multimodal content, and cannot adapt to dynamic scenarios and user behavior differences.

Method used

The AI-driven integrated media digital audio-visual content security audit method is adopted, through multimodal feature extraction and fusion, combined with user historical behavior and scene recognition, the audit strategy is dynamically adjusted, and deep learning and natural language processing technology is used to conduct intelligent review of video, audio, and graphic content.

Benefits of technology

It improves the accuracy and efficiency of audits, can adapt to differences in different scenarios and user behaviors, ensure content security and compliance, and meet the diversified needs of integrated media platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455732A_ABST
    Figure CN120455732A_ABST
Patent Text Reader

Abstract

The invention discloses an AI-driven convergence media digital audio and video content security auditing method, which comprises the following steps: acquiring audio and video contents published by a platform and user information of a publisher in real time, the audio and video contents comprising live streaming, video files, audio files and image-text contents; feature extraction is carried out on the video file, the text data, the video data and the audio data, and then fusion and vectorization processing are carried out to obtain a comprehensive feature vector; and inputting the comprehensive feature vector into an AI model for detection, determining whether illegal content exists or not, and processing according to a detection result. The AI technology is introduced, intelligent auditing of various modal contents such as videos, audios and images and texts is achieved, the auditing efficiency and accuracy are improved, meanwhile, the auditing strategy is dynamically adjusted according to different scenes and user behaviors, the diversified content auditing requirements of the convergence media platform are met, and the safety and compliance of digital audio and video contents are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of converged media technology, and specifically relates to an AI-driven converged media digital audio and video content security review method. Background Art

[0002] With the rapid development of converged media technology, the dissemination channels and formats of digital audiovisual content are becoming increasingly diverse. Video, audio, graphics, text, live broadcasts, and other content are now widely disseminated across various platforms. While the rapid dissemination of this content has enabled efficient information flow, it has also brought about content security issues. For example, the dissemination of sensitive or illegal content can pose a potential threat to social order and the public interest. Currently, the review of digital audiovisual content on converged media platforms primarily relies on solutions such as deep learning-based video content review and natural language processing (NLP)-based text content review. While these two AI-based content review technologies offer certain advantages in their respective fields, they both have significant drawbacks. They primarily rely on single-modal data (video or text) and lack the ability to comprehensively analyze multimodal content, which can easily lead to misjudgments. Furthermore, deep learning models typically require large amounts of labeled data for training, resulting in high data acquisition and annotation costs and complex model computations, making them difficult to meet the needs of large-scale, real-time review. Therefore, a comprehensive review method combining multimodal analysis, dynamic threshold adjustment, and user behavior assessment is needed to improve review accuracy and efficiency and adapt to the diverse content review needs of converged media platforms. Summary of the Invention

[0003] In response to the above-mentioned shortcomings, the present invention discloses an AI-driven integrated media digital audio-visual content security review method, which introduces AI technology to realize intelligent review of multi-modal content such as video, audio, graphics, etc., improves review efficiency and accuracy, and dynamically adjusts review strategies according to different scenarios and user behaviors to meet the diverse content review needs of integrated media platforms, ensure the security and compliance of digital audio-visual content, and solves the problems of low efficiency, insufficient accuracy, difficulty in multi-modal content review, and difficulty in adapting to dynamic scenarios and user behavior differences in the security review of digital audio-visual content on integrated media platforms.

[0004] The present invention is achieved by adopting the following technical solutions: An AI-driven integrated media digital audio and video content security review method includes the following steps: (1) Real-time acquisition of audio and video content published by the platform and user information of the publisher, the audio and video content including live streams, video files, audio files, and graphic content; frame rate adjustment, resolution normalization, and watermark removal of video files to obtain video data; noise reduction and volume normalization of audio files to obtain audio data; and text and image data extraction of graphic content to obtain text data; After word segmentation, part-of-speech tagging, and semantic analysis of text data, text features are extracted; Extract key frame features, motion features, and color features from the video data to obtain video features, and at the same time treat the image data as single-frame video data for feature extraction processing; Extracting spectrum features, rhythm features, and speech features from audio data to obtain audio features; The text features, video features and audio features are integrated and vectorized to obtain a comprehensive feature vector; (2) Inputting the comprehensive feature vector obtained in step (1) into an AI model for detection, wherein the AI model includes a user history behavior analysis module, a scene recognition module, an audit threshold allocation module, a video content detection module, an audio content detection module, a text content detection module, and a comprehensive output module; The user history behavior analysis module searches for user information to determine whether the user has posted any illegal content, and determines the user's credibility based on the number of times the illegal content has been posted. The scene recognition module is built based on a deep learning network and identifies the scene of the audio and video content based on comprehensive feature vector analysis, and the scenes include news, entertainment, and education; The audit threshold allocation module allocates audit thresholds to the video content detection module, the audio content detection module, and the text content detection module according to the user's credibility and the scene of the audio and video content; The video content detection module is constructed based on a neural network and analyzes and processes the comprehensive feature vector according to the assigned review threshold to identify the content of the picture and confirm whether it is a video type of non-violation picture, suspected violation picture, or violation picture; The audio content detection module is constructed based on the extension network and analyzes and processes the comprehensive feature vector according to the assigned review threshold to identify the audio content and confirm it as one of the audio types of non-violation audio, suspected violation audio, and violation audio; The text content detection module uses a natural language processing method and analyzes and processes the comprehensive feature vector according to the assigned review threshold to identify the text content and confirm whether it is a text type selected from the group consisting of non-violation text, suspected violation text, and violation text; The comprehensive output module integrates the confirmation results and confidence levels obtained by the video content detection module, the audio content detection module, and the text content detection module, and outputs the detection results. When the confirmation result is one of the illegal images, illegal audio, and illegal text, and the confidence level is greater than 90%, it is determined to be obvious illegal content; when the confirmation result is one of the illegal images, illegal audio, and illegal text, and the confidence level is less than 90%, it is determined to be suspected illegal content; when the confirmation result is one of the suspected illegal images, suspected illegal audio, and suspected illegal text, it is determined to be suspected illegal content; when the confirmation result is one of the non-illegal images, non-illegal audio, and non-illegal text, it is determined to be normal content; (3) Processing will be carried out based on the detection results. Normal content will be directly approved for review and allowed to be published or disseminated. Suspected illegal content will be marked as pending review and forwarded to manual reviewers for manual review. Obviously illegal content will be marked as illegal and blocked. The detection results will be recorded and sent to the user for notification. After receiving the notification, the user can improve the audio and video content to be published and re-review it, or send a complaint to the administrator for review.

[0005] Furthermore, in step (1), for a video file containing subtitles or text, the subtitle text is extracted to obtain text data or the text data is obtained by recognizing the text in the image through OCR.

[0006] Furthermore, the video content detection module is constructed based on a convolutional neural network (CNN), and the audio content detection module is constructed based on a recurrent neural network (RNN).

[0007] Furthermore, the video content detection module compares and analyzes the comprehensive feature vector with the feature vectors of pornographic, violent, and terrifying images to obtain similarities, and determines the video type based on whether the similarities are within a threshold range.

[0008] Furthermore, the audio content detection module compares and analyzes the comprehensive feature vector with the audio feature vector of the sensitive word to obtain similarity, and determines the audio-visual type based on whether the similarity is within a threshold range.

[0009] Furthermore, in step (3), relevant information about the detection results and processing methods is recorded in the system log for use in training and optimization of the AI model.

[0010] Compared with the existing technology, this technical solution has the following beneficial effects: 1. This invention is primarily used in the converged media industry, aiming to address issues such as low efficiency, insufficient accuracy, difficulty in multimodal content review, and difficulty adapting to dynamic scenarios and user behavior differences in the security review of digital audio and video content on converged media platforms. By introducing AI technology, it enables intelligent review of multimodal content such as video, audio, and graphics, improving review efficiency and accuracy. At the same time, it dynamically adjusts review strategies based on different scenarios and user behavior to meet the diverse content review needs of converged media platforms and ensure the security and compliance of digital audio and video content.

[0011] 2. This invention uses AI technology to achieve automated review, improve review efficiency, and meet the real-time review needs of large-scale content. This addresses the problem that manual review is difficult to meet the real-time review needs of large-scale content and is prone to missed or delayed reviews. Furthermore, the deep learning capabilities of AI models are used to conduct comprehensive analysis of content in multiple modalities, such as video, audio, and text, to improve the accuracy of identifying illegal content. This addresses the problem that rule-based review methods are prone to misjudgment and cannot effectively identify complex scenarios and semantics. For example, simple keyword filtering may mistakenly identify normal content as illegal content.

[0012] 3. The present invention uses multimodal feature extraction and fusion technology to achieve comprehensive review of content in multiple modalities, such as video, audio, and text, and comprehensively evaluate the security of the content. This solves the problem that most existing technologies focus on content review in a single modality, lack a comprehensive review mechanism for content in multiple modalities, such as video, audio, and text, and have difficulty in comprehensively evaluating the security of the content. At the same time, the present invention dynamically adjusts the review threshold based on the specific scenario of the content and user behavior, flexibly adapting to the review needs of different scenarios and user groups, and solving the problem that existing review systems cannot flexibly adjust review strategies based on different scenarios and user behavior. For example, news content has higher review requirements for politically sensitive content, while entertainment content has a higher tolerance for vulgar and violent content.

[0013] 4. The present invention analyzes the user's historical behavior and credibility, sets personalized review thresholds for different user groups, improves the fairness and effectiveness of the review, and solves the problem that the existing system lacks an analysis mechanism for user historical behavior and credibility, and is difficult to dynamically adjust the review threshold according to the user's credibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a flowchart of the AI-driven integrated media digital audio-visual content security review method described in Example 1.

[0015] Figure 2 This is a flowchart of step (3) in the AI-driven integrated media digital audio-visual content security review method described in Example 1. DETAILED DESCRIPTION

[0016] The present invention is further illustrated by the following examples, which are not intended to limit the present invention. Specific experimental conditions and methods not specified in the following examples are conventional methods well known to those skilled in the art.

[0017] The following examples utilize a three-tiered, interconnected knowledge base established by the applicants of this invention. Relevant data resources can be retrieved from this base, which covers the entire Guangxi region's media history and newly added manuscript resources. Furthermore, the acquisition of user data and integrated media content is performed with the authorization of both the user and publisher, and a security protection mechanism is implemented.

[0018] Example 1: An AI-driven method for security review of converged media digital audio and video content, comprising the following steps: (1) Real-time acquisition of audio and video content published by the platform and user information of the publisher, the audio and video content including live streams, video files, audio files, and graphic content; frame rate adjustment, resolution normalization, and watermark removal of video files to obtain video data; noise reduction and volume normalization of audio files to obtain audio data; and text and image data extraction of graphic content to obtain text data; Text features are extracted from text data after word segmentation, part-of-speech tagging, and semantic analysis. Video features are extracted from keyframe features, motion features, and color features, while image data is treated as single-frame video data for feature extraction. For video files containing subtitles or text, subtitle text is extracted to obtain text data, or text data is obtained by OCR recognition of text in images. Extracting spectrum features, rhythm features, and speech features from audio data to obtain audio features; The text features, video features and audio features are integrated and vectorized to obtain a comprehensive feature vector; (2) Inputting the comprehensive feature vector obtained in step (1) into an AI model for detection, wherein the AI model includes a user history behavior analysis module, a scene recognition module, an audit threshold allocation module, a video content detection module, an audio content detection module, a text content detection module, and a comprehensive output module; The user history behavior analysis module searches for user information to determine whether the user has posted any illegal content, and determines the user's credibility based on the number of times the illegal content has been posted. The scene recognition module is built based on a deep learning network and identifies the scene of the audio and video content based on comprehensive feature vector analysis, and the scenes include news, entertainment, and education; The audit threshold assignment module assigns audit thresholds to the video content detection module, the audio content detection module, and the text content detection module based on the user's credibility and the context of the audio and video content. For example, the audit threshold for politically sensitive content in news content can be set to be stricter; while the audit threshold for vulgar and violent content in entertainment content is higher due to its entertainment nature; and for children's educational videos, the threshold for pornographic and violent content can be set to be lower to ensure absolute content safety. The video content detection module is constructed based on a product neural network and analyzes and processes the comprehensive feature vector according to the assigned review threshold to identify the content of the picture and confirm whether it is a video type of non-violation picture, suspected violation picture, or violation picture. The video content detection module compares and analyzes the comprehensive feature vector with the feature vectors of pornographic, violent, and terrifying images to obtain similarity, and determines the video type based on whether the similarity is within the threshold range. The audio content detection module is constructed based on a recurrent neural network and analyzes and processes the comprehensive feature vector according to an assigned review threshold to identify the audio content and confirm whether it is an audio type of non-violation audio, suspected violation audio, or violation audio. The audio content detection module compares and analyzes the comprehensive feature vector with the audio feature vector of the sensitive word to obtain similarity, and determines the audio / video type based on the similarity within the threshold range. The text content detection module uses a natural language processing method and analyzes and processes the comprehensive feature vector according to the assigned review threshold to identify the text content and confirm whether it is a text type selected from the group consisting of non-violation text, suspected violation text, and violation text; The comprehensive output module integrates the confirmation results and confidence levels obtained by the video content detection module, the audio content detection module, and the text content detection module, and outputs the detection results. When the confirmation result is one of the illegal images, illegal audio, and illegal text, and the confidence level is greater than 90%, it is determined to be obvious illegal content; when the confirmation result is one of the illegal images, illegal audio, and illegal text, and the confidence level is less than 90%, it is determined to be suspected illegal content; when the confirmation result is one of the suspected illegal images, suspected illegal audio, and suspected illegal text, it is determined to be suspected illegal content; when the confirmation result is one of the non-illegal images, non-illegal audio, and non-illegal text, it is determined to be normal content; (3) Processing will be carried out based on the detection results. Normal content will be directly approved for review and allowed to be published or disseminated. Suspected illegal content will be marked as pending review and forwarded to manual reviewers for manual review. Obviously illegal content will be marked as illegal and blocked. The detection results will be recorded and sent to the user for notification. After receiving the notification, the user can improve the audio and video content to be published and re-review it, or send a complaint to the administrator for review.

[0019] Example 2: The AI-driven integrated media digital audio-visual content security review method described in this example differs from the method described in Example 1 only in that, in step (3), relevant information about the detection results and processing methods is recorded in the system log for use in AI model training and optimization.

[0020] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. An AI-driven method for security review of converged media digital audio and video content, characterized by: The following steps are involved: (1) Real-time acquisition of audio and video content published by the platform and user information of the publisher, the audio and video content including live streaming, video files, audio files, and graphic content; frame rate adjustment, resolution normalization, and watermark removal processing are performed on the video files to obtain video data, and noise reduction and volume normalization processing are performed on the audio files to obtain audio data; text data and image data are extracted from the graphic content; text features are extracted after word segmentation, part-of-speech tagging, and semantic analysis of the text data; key frame features, motion features, and color features are extracted from the video data to obtain video features, and the image data is treated as a single frame of video data for feature extraction processing; spectrum features, rhythm features, and speech features are extracted from the audio data to obtain audio features; The text features, video features and audio features are integrated and vectorized to obtain a comprehensive feature vector; (2) Inputting the comprehensive feature vector obtained in step (1) into an AI model for detection, the AI model includes a user history behavior analysis module, a scene recognition module, an audit threshold allocation module, a video content detection module, an audio content detection module, a text content detection module, and a comprehensive output module; the user history behavior analysis module searches for whether the user has posted illegal content based on user information, and confirms the user's credibility based on the number of times the illegal content has been posted; The scene recognition module is constructed based on a deep learning network, and confirms the scene of the audio-visual content according to the comprehensive feature vector analysis, and the scene includes news, entertainment, and education; the audit threshold allocation module allocates audit thresholds to the video content detection module, the audio content detection module, and the text content detection module according to the user credibility and the scene of the audio-visual content; the video content detection module is constructed based on a neural network, and analyzes and processes the comprehensive feature vector according to the allocated audit threshold, identifies the picture content and confirms it as a video type among non-violation pictures, suspected violation pictures, and violation pictures; the audio content detection module is constructed based on an extension network, and analyzes and processes the comprehensive feature vector according to the allocated audit threshold, identifies the audio content and confirms it as an audio type among non-violation audio, suspected violation audio, and violation audio; the text content detection module adopts natural language processing to analyze and process the comprehensive feature vector according to the allocated audit threshold, identifies the audio content and confirms it as an audio type among non-violation audio, suspected violation audio, and violation audio; A language processing method is provided, and the comprehensive feature vector is analyzed and processed according to the assigned review threshold, the text content is identified and confirmed as a text type among non-violation text, suspected violation text and violation text; the comprehensive output module integrates the confirmation results and confidence levels obtained by the video content detection module, the audio content detection module and the text content detection module, and outputs the detection results. When the confirmation result is one of violation images, violation audio and violation text and the confidence level is greater than 90%, it is determined to be obvious violation content; when the confirmation result is one of violation images, violation audio and violation text and the confidence level is less than 90%, it is determined to be suspected violation content; when the confirmation result is one of suspected violation images, suspected violation audio and suspected violation text, it is determined to be suspected violation content; when the confirmation result is one of non-violation images, non-violation audio and non-violation text, it is determined to be normal content; (3) Processing will be carried out based on the detection results. Normal content will be directly approved for review and allowed to be published or disseminated. Suspected illegal content will be marked as pending review and forwarded to manual reviewers for manual review. Obviously illegal content will be marked as illegal and blocked. The detection results will be recorded and sent to the user for notification. After receiving the notification, the user can improve the audio and video content to be published and re-review it, or send a complaint to the administrator for review.

2. The AI-driven integrated media digital audio and video content security audit method according to claim 1 is characterized by: In step (1), for a video file containing subtitles or text, the subtitle text is extracted to obtain text data or the text data is obtained by recognizing the text in the image through OCR.

3. The AI-driven integrated media digital audio and video content security review method according to claim 1 is characterized by: The video content detection module is constructed based on a product neural network, and the audio content detection module is constructed based on a recurrent neural network.

4. The AI-driven integrated media digital audio and video content security audit method according to claim 1 is characterized by: The video content detection module compares and analyzes the comprehensive feature vector with the feature vectors of pornographic, violent, and terrifying images to obtain similarities, and determines the video type based on whether the similarities are within a threshold range.

5. The AI-driven integrated media digital audio and video content security audit method according to claim 1 is characterized by: The audio content detection module compares and analyzes the comprehensive feature vector with the audio feature vector of the sensitive word to obtain similarity, and determines the audio and video type according to the threshold range of the similarity.

6. The AI-driven integrated media digital audio and video content security audit method according to claim 1 is characterized by: In step (3), relevant information about the detection results and processing methods is recorded in the system log for training and optimization of the AI model.

Citation Information

Cited By

  • Data security verification method and system based on AI

    CN120785658A

  • Media stream security detection method and device

    CN121397295A

  • Video release management method and system based on artificial intelligence

    CN121691751A