Video content automatic auditing method and system based on AI drive

Through multi-level video content segmentation and structured node modeling, combined with a large language model and temporal Transformer network, multimodal feature fusion and intelligent collaborative review are achieved, which solves the complexity problem of multimodal content review and improves the efficiency and accuracy of video content review.

CN120726547AActive Publication Date: 2025-09-30JIANGSU BROADCASTING CORPORATION

Patent Information

Application Number
CN202511233429.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-09-30
Estimated Expiration
2045-09-01

AI Technical Summary

Technical Problem

Existing video review technology is unable to cope with the collaborative analysis of multimodal content, node segmentation and structured modeling are insufficient, the identification of complex violations is incomplete, and the efficiency and accuracy of content review needs to be improved.

Method used

It adopts multi-level video content segmentation and structured node modeling, combines large language models and temporal Transformer networks for multimodal feature fusion, generates video content review reports through collaborative review by multiple virtual agents and dynamic allocation of preference weights, and utilizes knowledge graph analysis and principal component analysis.

Benefits of technology

It significantly improves the multimodal collaborative recognition capability of video content review, improves the accuracy of identifying complex violations and the robustness of review results, reduces the misjudgment rate and missed judgment rate, and realizes efficient and accurate automated review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726547A_ABST
    Figure CN120726547A_ABST
Patent Text Reader

Abstract

The invention discloses a video content automatic auditing method and system based on AI drive, and relates to the technical field of video content analysis, and the method comprises the steps: obtaining video information to be audited, and carrying out the multi-level segmentation of the video information; in the slices, node simulation is carried out according to the content independence of the video information; analyzing associated branches and historical features of each node to generate attribute features of the nodes; content auditing is carried out through an auditing strategy of multiple virtual agents, and illegal space analysis is carried out according to an unstable part in an auditing result; and generating a final video content auditing report according to the auditing result and the analysis result of the illegal space. According to the method, the structured recognition and intelligent auditing efficiency of the multi-modal video content is improved, accurate positioning, label attribution and automatic early warning of complex violation behaviors are realized, and the content security management capability is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of video content analysis technology, and in particular to an AI-driven automatic video content review method and system. Background Art

[0002] With the rapid adoption of mobile internet and smart devices, the volume and influence of video content on various platforms, including social, entertainment, and education, are increasing. While the richness and diversity of online video content has boosted the efficiency of information dissemination, it has also created a series of social governance challenges, including content compliance, information security, and youth protection. This is particularly true in scenarios like short videos, live broadcasts, and film and television sharing, where video content often includes multimodal information such as images, audio, subtitles, and dubbing. Traditional rule-based automated review or manual review methods are unable to rapidly screen and judge content at this scale.

[0003] Most existing video review technologies focus on image recognition or text analysis, lacking the ability to deeply understand multimodal content and determine cross-modal consistency. With the increasing sophistication of video content editing, and the widespread use of editing, splicing, and special effects, the concealment, mutation, and disguise of illegal information have become more subtle and diverse. Single-modal or single-level review methods struggle to capture the complex semantic connections between clips and fail to balance global structure with local details.

[0004] Furthermore, improving the real-time, accuracy, and scalability of content review, reducing the cost of manual intervention, and enhancing the model's adaptability to new types of violations have become key technical areas of focus for the industry. Faced with increasingly complex application scenarios and compliance requirements, building a more intelligent, structured, and multi-layered collaborative video content review system has become a major technical challenge for the industry. Summary of the Invention

[0005] In view of the above existing problems, the present invention is proposed.

[0006] Therefore, the present invention provides an AI-driven automatic video content review method to solve problems such as the difficulty in collaborative analysis of multimodal content, insufficient node segmentation and structured modeling, incomplete identification of complex violations, and the need to improve content review efficiency and accuracy.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0008] In a first aspect, the present invention provides an AI-driven automatic video content review method, which includes obtaining video information to be reviewed and performing multi-level segmentation on the video information;

[0009] In the slice, node simulation is performed according to the content independence of the video information; and the associated branches and historical features of each node are analyzed to generate the attribute features of the node;

[0010] Conduct content review through a multi-virtual agent review strategy, and analyze illegal content based on the unstable parts of the review results;

[0011] Generate a final video content review report based on the review results and the analysis results of the illegal space;

[0012] The audit strategy of the multi-virtual agent includes matching different preferences for all agents, voting based on the audit results of each agent, and thus generating an audit result; wherein the agent preference matching process is: randomly matching the attention weight of each agent in the preference library.

[0013] As a preferred solution of the AI-driven automatic video content review method of the present invention, wherein: the video information includes video frame sequence data and audio stream data in the video to be reviewed;

[0014] After aligning the video frame sequence data and the audio stream data according to the timestamps, inserting segmentation points on the time axis of the video to be reviewed, and implementing multi-level segmentation of the video information based on the inserted segmentation points:

[0015] Step 1: Insert the first-level segmentation point according to the exit point of each program in the video content;

[0016] Step 2: In the video frame sequence data, identifying the shot switching moment and inserting the secondary segmentation point;

[0017] Step 3: Using an AI-driven large language model, identify each sentence in the audio stream data and, based on the recognition results of the sentence semantics, identify consecutive sentences with semantic connections. Insert a secondary segmentation point after each consecutive sentence. Insert a tertiary segmentation point after each sentence.

[0018] Step 4: Segment the timeline according to the first-level segmentation points to obtain first-level slices of the video;

[0019] The second-level segmentation points obtained in steps 2 and 3 are segmented on the time axis of the first-level slice to obtain second-level slices under each first-level slice; according to the third-level segmentation points, the second-level slice is segmented on the time axis to obtain third-level slices under each second-level slice.

[0020] As a preferred solution of the AI-driven automatic video content review method of the present invention, the node simulation includes treating each level of the slice of the video information on the timeline as a node; and synchronizing the hierarchical relationship of each node based on the hierarchical relationship between the slices to obtain three levels of nodes: ;in, Indicates the first-level node corresponding to the i-th first-level slice; Indicates the secondary node corresponding to the jth secondary slice under the i-th primary slice; It shows the j-th second-level slice under the i-th first-level slice, and the third-level nodes corresponding to the third-level slices included.

[0021] As a preferred solution of the AI-driven automatic video content review method of the present invention, the attribute features include: using the knowledge graph to construct the associated branches and analyzing the node features in the video information environment; analyzing the habitual features of the current node through the features of similar nodes in the historical records;

[0022] The AI-driven large language model is used to encode the audio portion of each node to obtain the feature vector of the audio portion. The temporal Transformer network is used to encode the video portion of each node to obtain the feature vector of the video portion. The feature representation of each node is: ;in, A feature vector representing the audio portion of the node, A feature vector representing the video portion of the node; generating edges between nodes by analyzing the correlation between the feature vectors between nodes, and constructing the knowledge graph;

[0023] The correlation analysis includes generating edges between each two nodes through principal component analysis and temporal correlation analysis, wherein the edge relationship includes: temporal correlation between two similarity correlation sequences and four feature vector pairing results;

[0024] The principal component analysis strategy includes measuring the similarity of the two feature vectors of the audio part and the video part in the node, respectively, to obtain the similarity correlation sequence of the audio part and the similarity correlation sequence of the video part; the measurement process for each feature vector includes: step 1, calculating the similarity of the feature vectors in each two nodes, and obtaining the initial similarity of each two nodes, which is recorded as Represents the initial similarity between node w and node e on the rth feature vector, 0 represents the initial index of the similarity sequence, and r contains two feature vectors of the audio part and the video part of the node;

[0025] Step 2: Based on the same eigenvectors on different nodes, construct an eigenvector matrix; use the preset cumulative explained variance threshold to reduce the dimension of the eigenvector matrix through principal component analysis to obtain the eigenvector of each node after dimensionality reduction; use the eigenvector of each node after dimensionality reduction to calculate the similarity between every two nodes to obtain the similarity between every two nodes after the first dimensionality reduction: ;

[0026] Step 3: Repeat step 2 until the change in the similarity values ​​after two consecutive dimensionality reductions is less than the preset minimum fluctuation threshold; the sequence is obtained: ;

[0027] Among them, n represents the number of dimensionality reduction operations, and n-2 is used to remove the repeated similarity parts of the sequence;

[0028] The temporal correlation analysis includes inputting the feature representation F of each node into a pre-trained temporal modeling network and outputting the correlation probability of the four feature vector pairing results between two nodes: ;in, Represents the correlation probability between the feature vectors of the audio part of node w and node e after pairing; Represents the probability of correlation between the feature vector of the audio part of node w and the feature vector of the video part of node e after pairing; represents the correlation probability between the feature vector of the audio part of node e and the feature vector of the video part of node w after pairing; Represents the probability of correlation between the feature vectors of node w and node e after pairing the video part.

[0029] As a preferred solution of the AI-driven automatic video content review method of the present invention, the feature representation F of each node is and , calculate the prior probability respectively, and generate and The probability distribution of each violation situation is used as the characteristic feature of each node.

[0030] As a preferred solution of the AI-driven automatic video content review method of the present invention, the review process of each intelligent agent includes two rounds of review stages: in the first stage, the intelligent agent traverses each node and obtains and For each violation probability distribution, if the probability of a violation is greater than a preset value, the video content at the node is marked and an alert is issued. Otherwise, the second stage begins: using a pre-trained neural network, the video content of each node in the knowledge graph is reviewed to generate a probability of violation for each node.

[0031] The voting according to the audit results of each agent includes: if any agent identifies that there is a node violation, it votes +1 for the identified illegal node; if the number of votes at the node is greater than 1, it is judged that there is an unstable part in the video content; if the number of votes at the node is 0, it is judged that the video content has passed the review;

[0032] The preference library is for each edge relationship The preset weight value range is such that the sum of the weights of the four elements matched each time is 1;

[0033] The agent audits every two nodes in each sequence. In the example, find the maximum and minimum values, and generate the guide vector of the node feature according to the maximum and minimum values; splice the two guide vectors obtained from the audio and video parts; obtain two splicing vectors by splicing in different orders, and use and as weight;

[0034] The pre-trained neural network is used to determine the probability distribution of each guide vector and each violation, as well as the probability distribution of each splicing vector and each violation. The weighted summation is performed using the weight preference of the intelligent agent to obtain the final probability distribution result.

[0035] As a preferred solution of the AI-driven automatic video content review method of the present invention, the analysis of the violation space includes: when determining that the node content is in violation, according to the voting results, calculating the proportion of different violations; generating the distribution of the illegal content in different violation situations, inserting corresponding violation labels for the illegal nodes, and issuing an early warning;

[0036] When the video content review report is generated, the node range of the violation label is narrowed down, and only the lowest node referred to by each violation label is retained.

[0037] In a second aspect, the present invention provides an AI-driven automatic video content review system, comprising: an acquisition unit for acquiring video information to be reviewed and performing multi-level segmentation on the video information;

[0038] The simulation unit simulates nodes in the slices according to the content independence of the video information; and analyzes the associated branches and historical features of each node to generate attribute features of the node;

[0039] The analysis unit conducts content review through the review strategy of multiple virtual agents, and analyzes the violation space based on the unstable parts in the review results;

[0040] The output unit generates a final video content review report based on the review results and the analysis results of the illegal space.

[0041] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the AI-driven automatic video content review method as described in the first aspect of the present invention is implemented.

[0042] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the AI-driven automatic video content review method as described in the first aspect of the present invention.

[0043] The beneficial effects of the present invention are as follows: through multi-level segmentation and structured node modeling, it is possible to fully capture the complex temporal and semantic relationships of video content, and achieve fine-grained management of video clips. The use of multimodal feature fusion and knowledge graph analysis significantly improves the collaborative recognition capabilities of multi-source information such as audio and video. Combining principal component analysis with time series networks effectively enhances the accuracy of identifying content change trends, abnormal behaviors, and cross-segment associations. The multi-virtual agent collaborative review and voting mechanism not only improves the robustness of the review results, but also can adaptively respond to diverse violations. Through fault tolerance and violation space analysis, the present invention significantly reduces the misjudgment rate and missed judgment rate, achieves efficient, accurate, and intelligent automatic review of complex video content, and significantly improves content security management and platform compliance operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0045] Figure 1 Flowchart of the AI-driven automatic video content review method. DETAILED DESCRIPTION

[0046] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0047] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0048] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.

[0049] Reference Figure 1 , is an embodiment of the present invention, which provides an AI-driven automatic video content review method, including the following steps:

[0050] S1: Obtain video information to be reviewed, and perform multi-level segmentation on the video information.

[0051] Furthermore, the video information includes the video frame sequence data and audio stream data in the video to be reviewed. After aligning the video frame sequence data and the audio stream data according to the timestamps, segmentation points are inserted on the timeline of the video to be reviewed, and the video information is segmented into multiple levels according to the inserted segmentation points:

[0052] Step 1: Insert a first-level segmentation point based on the exit point of each program in the video content. "Each program's exit point" generally refers to the starting and ending points of each independent program unit in the video content stream. This can be the program's start point, end point, or content switching boundary. In this invention, this is measured as the program's start point.

[0053] Step 2: In the video frame sequence data, identify the shot switching moment (the moment between two frames) and insert a secondary segmentation point.

[0054] Step 3: Use the AI-driven large language model to identify each sentence in the audio stream data, and based on the recognition results of the sentence semantics, identify semantically connected consecutive sentences; insert a secondary segmentation point at the moment after the consecutive sentences; and insert a tertiary segmentation point at the moment after each sentence.

[0055] In practice, an AI-driven large language model generates a feature vector for each encoded sentence. Each feature vector represents a sentence. The relationships between these feature vectors are further utilized to identify semantic concatenation relationships. In this embodiment, the AI-driven large language model is preferably a large, pre-trained natural language processing model based on the Transformer architecture, used to perform sentence-level understanding and feature encoding on the text transcribed from the audio stream. Optional models include, but are not limited to, general-purpose language models such as BERT, RoBERTa, ALBERT, ERNIE, XLNet, and T5. Models optimized for conversation and speech-to-text, such as ChatGPT, Llama, and Baidu-ERNIE-Bot, can also be used. Each sentence is input into the aforementioned model to obtain a sentence feature vector representing its semantic information. In a preferred embodiment, a long short-term memory (LSTM) layer is introduced into the neural network architecture to perform temporal modeling on the sequence of sentence-level feature vectors. The LSTM layer effectively captures contextual dependencies and semantic cohesion between sentences, thereby automatically learning and discerning semantic concatenation relationships between consecutive sentences. Specifically, the sentence feature vectors arranged in chronological order are input into the LSTM layer in sequence, and its hidden state output is used to analyze the correlation between the current sentence and the context sentences.

[0056] Step 4: Split the timeline according to the first-level segmentation points to obtain first-level slices of the video. Split the timeline of the first-level slices using the second-level segmentation points obtained in steps 2 and 3 to obtain second-level slices under each first-level slice; and split the timeline of the second-level slices according to the third-level segmentation points to obtain third-level slices under each second-level slice.

[0057] First, by setting the first-level segmentation point through the "exit point" of the program unit, the video stream can be effectively divided into logically independent program segments, improving the accuracy of content attribution and structure recognition. Secondly, by identifying the second-level segmentation point through the lens switching point, the visual scene changes within the program can be further refined and segmented to facilitate the capture of important picture changes, abnormal scenes and potential risks. Furthermore, a large language model is used to perform sentence-level and semantic-level recognition of audio content, and the third-level and auxiliary segmentation points are set in combination with the semantic series relationship to achieve precise positioning of semantic information and context tracing under audio and video synchronization. The multi-level segmentation mechanism organically combines the temporal structure, semantic structure and visual structure of the video content, providing a solid foundation for subsequent steps such as node modeling, feature extraction, intelligent review and violation detection, and effectively improving the granularity and intelligence level of content analysis.

[0058] S2: In the slice, node simulation is performed according to the content independence of the video information; and associated branches and historical features of each node are analyzed to generate attribute features of the node.

[0059] It is important to note that the attribute features include using the knowledge graph to construct the associated branches and analyze the node characteristics in the video information environment; and analyzing the habitual characteristics of the current node through the characteristics of similar nodes in historical records. By modeling the structured correlation of multimodal features between nodes, deep semantic connections are achieved in the audio and visual dimensions of node-level video content, providing highly robust and high-resolution knowledge support for intelligent review and anomaly detection.

[0060] The AI-driven large language model is used to encode the audio portion of each node to obtain the feature vector of the audio portion (using the vector generated by the large language model in S1 to save additional calculations). The temporal Transformer network is used to encode the video portion of each node to obtain the feature vector of the video portion. The feature representation of each node is: ;in, A feature vector representing the audio portion of the node, A feature vector representing the video portion of the node; generating edges between nodes by analyzing the correlation between the feature representations of the nodes, and constructing the knowledge graph.

[0061] Specifically, the correlation analysis includes generating edges between every two nodes through principal component analysis strategy and temporal correlation analysis respectively; wherein the edge relationship includes: temporal correlation of two similarity correlation sequences and four feature vector pairing results.

[0062] Furthermore, the principal component analysis strategy includes measuring the similarity of the two feature vectors of the audio part and the video part in the node, respectively, to obtain the similarity correlation sequence of the audio part and the similarity correlation sequence of the video part; the measurement process for each feature vector includes (that is, the analysis of any feature vector of audio or video is performed in the following text. By analyzing the two feature vectors separately, two sequences are generated): Step 1, calculate the similarity of the feature vectors in each two nodes, and obtain the initial similarity of each two nodes, which is recorded as Represents the initial similarity between node w and node e on the rth feature vector, 0 represents the initial index of the similarity sequence, and r contains two feature vectors of the audio part and the video part in the node.

[0063] Step 2: Based on the same eigenvectors on different nodes, construct an eigenvector matrix; use the preset cumulative explained variance threshold to reduce the dimension of the eigenvector matrix through principal component analysis to obtain the eigenvector of each node after dimensionality reduction; use the eigenvector of each node after dimensionality reduction to calculate the similarity between every two nodes to obtain the similarity between every two nodes after the first dimensionality reduction: .

[0064] Step 3: Repeat step 2 until the change in the similarity values ​​after two consecutive dimensionality reductions is less than the preset minimum fluctuation threshold; the sequence is obtained: .

[0065] Here, n represents the number of dimensionality reduction operations. By using n-2, the sequence will remove the repeated similarities to reduce the amount of calculation.

[0066] It should be noted that the analysis of temporal correlation includes inputting the feature representation F of each node into a pre-trained temporal modeling network (in this embodiment, a long short-term memory neural network (LSTM) is preferably used. In other optional embodiments, a gated recurrent unit (GRU), a temporal transformer, or other models can also be used to globally model the node sequence and learn the temporal dependencies and contextual dynamic associations between nodes), and outputting the correlation probabilities of the four feature vector pairing results between two nodes: ;in, Represents the correlation probability between the feature vectors of the audio part of node w and node e after pairing; Represents the probability of correlation between the feature vector of the audio part of node w and the feature vector of the video part of node e after pairing; represents the correlation probability between the feature vector of the audio part of node e and the feature vector of the video part of node w after pairing; Represents the probability of correlation between the feature vectors of node w and node e after pairing the video part.

[0067] Specifically, AI-driven large language models and temporal transformers are used to encode the audio and video parts of the nodes respectively, obtain a unified feature representation, and achieve efficient fusion of multimodal information. The principal component analysis strategy can eliminate redundancy and noise in high-dimensional features, retain only representative main information, improve the effectiveness and explanatory power of feature similarity measurements between different nodes, and dynamically optimize the feature space through an adaptive dimensionality reduction process. Further combined with temporal modeling networks (such as LSTM / GRU / temporal transformers), the complex dependencies and dynamic coupling features between nodes are captured in the global temporal context, and the correlation probabilities under various pairing conditions are output, significantly enhancing the expressive power of the knowledge graph edge weights. Through the above-mentioned multi-dimensional and fine-grained structured correlation analysis, the deep logical connections and abnormal links between video content segments can be accurately revealed, providing a solid decision-making basis for subsequent intelligent collaborative review, illegal space identification, and content security management.

[0068] The feature representation of each node in F and , calculate the prior probability respectively, and generate and The probability distribution of each violation situation is used as the characteristic feature of each node.

[0069] By modeling prior probabilities for each node's audio and video feature vectors, corresponding probability distributions are generated for various violation scenarios, providing an efficient, interpretable, and traceable basis for risk assessment in the subsequent content review process. Leveraging large-scale historical data or expert knowledge, prior probability distribution models are established for different violation types (such as pornography, violence, sensitive topics, and copyright risks). Node features are then matched with known violation labels to enable rapid quantitative assessment of the inherent risk attributes of node content. This mechanism not only improves the review process's ability to provide early warning of potential violations but also reduces reliance on real-time large-scale model inference, reducing system computational pressure. By supplementing the inherent feature probability distribution, the robustness and fault tolerance of node discrimination are further enhanced. This allows the entire review system to combine prior knowledge and real-time features for joint decision-making when faced with new or complex violations, enhancing the intelligence and sophistication of content security management.

[0070] S3: Conduct content review through the review strategy of multiple virtual agents, and analyze the violation space based on the unstable parts in the review results.

[0071] It is important to know that the audit strategy of multiple virtual agents includes matching different preferences for all agents, voting based on the audit results of each agent, and thus generating audit results; among them, the agent preference matching process is: randomly matching the attention weight of each agent in the preference library.

[0072] Furthermore, the audit process of each agent includes two rounds of audit phases: in the first phase, the agent traverses each node and obtains and For the probability distribution of each violation, if the probability of a violation is greater than the preset value, the video content at the node will be marked and an early warning will be issued (this is an audit result); otherwise, the second stage will be entered: using the pre-trained neural network, the video content of each node in the knowledge graph will be audited to generate the probability of violation for each node.

[0073] The voting based on the review results of each intelligent agent includes: if any intelligent agent recognizes that there is a node violation (generally speaking, the probability of violation is greater than a certain threshold), then the vote for the identified illegal node is +1; if the number of votes at the node is greater than 1, it is judged to be an unstable part in the video content; if the number of votes at the node is 0, then the video content is judged to have passed the review.

[0074] The preference library is for each edge relationship The preset weight value range satisfies that the sum of the weights of the four elements matched each time is 1.

[0075] The agent audits every two nodes in each sequence. In the dataset, the maximum and minimum values ​​are found and guiding vectors for node features are generated based on these values. Specifically, the vector difference corresponding to the maximum value is used as a translation, and the starting and ending positions of the vector difference corresponding to the minimum value are shifted (effectively, the difference between the vector difference corresponding to the minimum value minus the difference between the vector difference corresponding to the maximum value) to obtain a guiding vector representing feature change. By carefully analyzing the extreme similarity values ​​between each pair of nodes in the feature sequence, directional modeling of risk evolution trends between nodes is achieved. Generally, the maximum similarity in the sequence often corresponds to low-risk segments with highly similar features and close semantic connections between nodes. In these areas, the vector difference is small and cannot effectively reflect abnormal changes. In contrast, the minimum similarity often reveals significant divergence or sudden changes in the multimodal attributes or semantics of the content between nodes. This is the "demarcation zone" where violations, abnormal evolution, and potential risks are most likely to occur. Specifically, the two node feature vectors corresponding to the minimum similarity are selected and the vector difference is used to determine the direction of change at the "most semantic level." This direction of change is the dominant path for dramatic divergence or abnormal jumps in content attributes between nodes, making it highly likely to carry characteristics such as violations, sensitivity, and derailment. Compared to single extreme value or mean statistics, this guidance vector uses the "global farthest point" as the starting point for risk analysis, which can effectively amplify the performance of potential violations in the feature space and significantly improve the system's ability to identify complex and highly concealed content. The vector difference corresponding to the maximum similarity is used as a translation benchmark to ensure that the extracted guidance vector focuses on the actual content changes between nodes, rather than local noise or irrelevant disturbances. Ultimately, combined with the neural network's mapping of guidance vectors to the probability distribution of various types of violations, the system can intelligently determine whether such changes point to violation risks, thereby achieving directional guidance and high-confidence early warning of the content evolution process.

[0076] The vector difference between nodes with the minimum similarity corresponds to the direction of the furthest semantic divergence between nodes, which serves as a guiding vector for identifying the evolutionary trend of content violations. This vector accurately reflects abnormal attribute jumps and risk mutations between nodes, effectively improving the content review system's ability to intelligently identify and warn of complex violations.

[0077] The two guide vectors (set as H and T) obtained from the audio and video parts are spliced; two splicing vectors [T,H] and [H,T] are obtained by splicing in different orders, and are used respectively. and As a weight.

[0078] The pre-trained neural network is used to determine the probability distribution of each guide vector and each violation, as well as the probability distribution of each splicing vector and each violation. The weighted summation is performed using the weight preference of the intelligent agent to obtain the final probability distribution result.

[0079] By matching diverse preferences with attention weights, the system simulates the differentiated focus of multiple agents in assessing video content risk in real-world scenarios, effectively improving the system's adaptability and robustness to complex and volatile violations. The two-stage review mechanism ensures hierarchical identification of explicit and implicit violation risks, effectively capturing high-risk nodes while also enabling deeper, context-enhanced analysis of challenging samples. By dynamically assigning different attention weights to each agent in the preference library, it ensures that each agent forms complementary judgments within a multidimensional feature space. The design of guidance vectors (based on extreme values ​​and their variation in feature sequences) effectively captures representative directions and anomalous characteristics of attribute changes between nodes, enhancing sensitivity to changing trends and anomalous behavior in video content. The splicing and weighting mechanism further integrates multimodal features and their variation relationships, making probabilistic judgments more holistic and targeted. Ultimately, the integrated voting and probabilistic weighted decision-making effectively improves the accuracy, robustness, and traceability of review results, significantly reducing the risk of missed and misjudgment, and meeting the high standards required for intelligent and refined video content security review in real-world scenarios.

[0080] Furthermore, if node content is determined to be in violation, the voting results are used to calculate the percentage of different violations. The distribution of infringing content across different violation scenarios is generated, and corresponding violation labels are assigned to the offending nodes, triggering an alert. When violations are detected in node content, the system can perform detailed statistics and labeling of the violation type, enabling hierarchical management and intelligent tracing of multiple violation risks. By calculating the voting distribution across all agent review results, the system not only determines whether a node is in violation but also quantifies the percentage of different violation types within the same node, thereby characterizing the diversity and complexity of content risks. This mechanism provides a structured data foundation for subsequent content governance, tiered resolution, and automated alerting. Furthermore, by automatically inserting specific violation labels into infringing nodes, the system enables precise tracing and location, facilitating the interpretation and review of review decisions, and providing a basis for platform compliance operations and user protection. Furthermore, this real-time alert mechanism helps the platform respond promptly to potential risks, reduce the spread and impact of infringing content, and enhance overall content security and intelligent management.

[0081] S4: Generate a final video content review report based on the review results and the analysis results of the illegal space.

[0082] When the video content review report is generated, the node range of the violation label is narrowed down, and only the lowest node referred to by each violation label is retained.

[0083] When generating a video content review report, the system collects the inserted illegal labels and their distribution across all nodes (including nodes at all levels in a multi-layered structure). For each illegal label type, the system checks the distribution of similar illegal labels across all subnodes (such as second- and third-level slices) starting from the top node (such as the first-level slice).

[0084] If the illegal label on a certain level node does not have the same label in all its direct child nodes (next level) (that is, the distribution of illegal content can be considered consistent within the preset fluctuation range), the label will remain on the current node.

[0085] If there is an illegal label on the next level that is consistent with the current label, the label on the current node is "hidden" and only the label on the next level child node is retained.

[0086] Recurse in sequence until the bottom layer (the finest-grained node).

[0087] Ultimately, the offending label is retained only on the lowest-level node where no finer-grained child nodes share the same offending label, achieving minimum coverage and unique labeling for violation localization. However, it should be noted that during the node scope reduction process, the offending label "blanking" operation adopts a locally recursive, on-demand phasing mechanism. Specifically, for each node with an offending label, the system compares only that node with all its directly subordinate nodes. If a node A in a higher-level layer has an offending label that is identical to A's (e.g., determined to be "distributed consistently" or "consistent within a preset fluctuation range"), only the label of node A is phasing out (i.e., the label is no longer displayed at the A level, and only the labels of the subordinate nodes are retained). If no subordinate nodes of A have identical labels, the A-level label is retained. This phasing determination and operation applies only to node A and its subordinate nodes and does not affect label processing on other nodes in the same level (such as node B). In other words, if no subordinate nodes of node B in the same level have identical offending labels, the label on node B is retained, and only the label on node A is phasing out. This "node-by-node, branch-by-branch" recursive processing ensures that label positioning is both precise and avoids unnecessary loss, and does not cause global blanking that leads to omission of important information.

[0088] By recursively eliminating multi-level node labels and preserving the uniqueness of the lowest level, the system can effectively prevent the repeated labeling of the same violation in the upper-level node and all child nodes, simplifying the violation tracing process and improving the clarity and operability of content review results. This mechanism helps platforms or regulators quickly identify the most representative and detailed violation fragments when faced with complex multi-level content, improving processing efficiency and the scientific nature of compliance decisions, and preventing misjudgments and multiple judgments. It also provides high-quality structured review data support for downstream applications such as automated content distribution, risk visualization, and big data governance.

[0089] This embodiment also provides an AI-driven automatic video content review system, including:

[0090] The acquisition unit obtains the video information to be reviewed and performs multi-level segmentation on the video information.

[0091] The simulation unit performs node simulation in the slice according to the content independence of the video information; and analyzes the associated branches and historical features of each node to generate attribute features of the node.

[0092] The analysis unit conducts content review through the review strategy of multiple virtual agents, and analyzes the violation space based on the unstable parts in the review results.

[0093] The output unit generates a final video content review report based on the review results and the analysis results of the illegal space.

[0094] This embodiment also provides a computer device, which is suitable for the case of an AI-driven automatic video content review method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the AI-driven automatic video content review method proposed in the above embodiment.

[0095] The computer device may be a terminal, comprising a processor, memory, a communication interface, a display, and an input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and computer programs. The internal memory provides an environment for the operating system and computer programs stored in the non-volatile storage media. The communication interface of the computer device is used to communicate with external terminals via wired or wireless communication. Wireless communication may be achieved via Wi-Fi, a carrier network, NFC (near-field communication), or other technologies. The display of the computer device may be a liquid crystal display or an electronic ink display. The input device may be a touchscreen overlay on the display, buttons, a trackball, or a touchpad on the computer device housing, or an external keyboard, touchpad, or mouse.

[0096] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the AI-driven automatic video content review method proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0097] In summary, the present invention achieves this through: multi-level video content segmentation and structured node modeling, combined with a large language model and a temporal Transformer network to deeply encode the multimodal features of audio and video, and uses principal component analysis and temporal correlation modeling methods to achieve multi-dimensional, fine-grained correlation analysis and knowledge graph construction between nodes; further, through multi-virtual agent collaborative review, dynamic allocation of preference weights, guidance vectors and a two-stage review and judgment mechanism, the recognition capability and judgment robustness of content violations are improved. Finally, based on violation space statistics and recursive label elimination mechanism, the system can accurately locate high-risk nodes in video content, realize automated, intelligent and traceable content security review, and effectively improve the intelligence level and governance efficiency of video content compliance management.

[0098] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. An AI-driven automatic video content review method, characterized by: Including, obtaining video information to be reviewed and performing multi-level segmentation on the video information; In the slice, node simulation is performed according to the content independence of the video information; and the associated branches and historical features of each node are analyzed to generate the attribute features of the node; Conduct content review through a multi-virtual agent review strategy, and analyze illegal content based on the unstable parts of the review results; Generate a final video content review report based on the review results and the analysis results of the illegal space; The audit strategy of the multi-virtual agent includes matching different preferences for all agents, voting based on the audit results of each agent, and thus generating an audit result; wherein the agent preference matching process is: randomly matching the attention weight of each agent in the preference library.

2. The AI-driven automatic video content review method according to claim 1, characterized in that: The video information includes video frame sequence data and audio stream data in the video to be reviewed; After aligning the video frame sequence data and the audio stream data according to the timestamps, inserting segmentation points on the time axis of the video to be reviewed, and implementing multi-level segmentation of the video information based on the inserted segmentation points: Step 1: Insert the first-level segmentation point according to the exit point of each program in the video content; Step 2: In the video frame sequence data, identifying the shot switching moment and inserting the secondary segmentation point; Step 3: Using an AI-driven large language model, identify each sentence in the audio stream data and, based on the recognition results of the sentence semantics, identify semantically connected consecutive sentences; insert a secondary segmentation point after the consecutive sentences; and insert a tertiary segmentation point after each sentence; Step 4: Segment the timeline according to the first-level segmentation points to obtain first-level slices of the video; The second-level segmentation points obtained in steps 2 and 3 are segmented on the time axis of the first-level slice to obtain second-level slices under each first-level slice; according to the third-level segmentation points, the second-level slice is segmented on the time axis to obtain third-level slices under each second-level slice.

3. The AI-driven automatic video content review method according to claim 2, characterized in that: The node simulation includes treating each level of the slice of the video information on the time axis as a node; and synchronizing the hierarchical relationship of each node according to the hierarchical relationship between the slices to obtain three levels of nodes: ;in, Indicates the first-level node corresponding to the i-th first-level slice; Indicates the secondary node corresponding to the jth secondary slice under the i-th primary slice; It shows the j-th second-level slice under the i-th first-level slice, and the third-level nodes corresponding to the third-level slices included.

4. The AI-driven automatic video content review method according to claim 3, characterized in that: The attribute features include constructing the associated branches using the knowledge graph and analyzing the node features in the video information environment; Analyze the characteristics of the current node through the characteristics of similar nodes in the historical records; Using the AI-driven large language model, the audio part in each node is encoded to obtain the feature vector of the audio part; using The temporal Transformer network encodes the video part in each node and obtains the feature vector of the video part; the feature representation of each node is: ;in, a feature vector representing the audio portion of the node, A feature vector representing the video portion of the node; generating edges between nodes by analyzing the correlation between feature representations between nodes, and constructing the knowledge graph; The correlation analysis includes generating edges between each two nodes through principal component analysis and temporal correlation analysis, wherein the edge relationship includes: temporal correlation between two similarity correlation sequences and four feature vector pairing results; The principal component analysis strategy includes measuring the similarity of the two feature vectors of the audio part and the video part in the node, respectively, to obtain the similarity correlation sequence of the audio part and the similarity correlation sequence of the video part; the measurement process for each feature vector includes: step 1, calculating the similarity of the feature vectors in each two nodes, and obtaining the initial similarity of each two nodes, which is recorded as Represents the initial similarity between node w and node e on the rth feature vector, 0 represents the initial index of the similarity sequence, and r contains two feature vectors of the audio part and the video part of the node; Step 2: Based on the same eigenvectors on different nodes, construct an eigenvector matrix; use the preset cumulative explained variance threshold to reduce the dimension of the eigenvector matrix through principal component analysis to obtain the eigenvector of each node after dimensionality reduction; use the eigenvector of each node after dimensionality reduction to calculate the similarity between every two nodes to obtain the similarity between every two nodes after the first dimensionality reduction: ; Step 3: Repeat step 2 until the change in the similarity values ​​after two consecutive dimensionality reductions is less than the preset minimum fluctuation threshold; the sequence is obtained: ; Among them, n represents the number of dimensionality reduction operations, and n-2 is used to remove the repeated similarity parts of the sequence; The temporal correlation analysis includes inputting the feature representation F of each node into a pre-trained temporal modeling network and outputting the correlation probability of the four feature vector pairing results between two nodes: ;in, Represents the correlation probability between the feature vectors of the audio part of node w and node e after pairing; Represents the probability of correlation between the feature vector of the audio part of node w and the feature vector of the video part of node e after pairing; represents the correlation probability between the feature vector of the audio part of node e and the feature vector of the video part of node w after pairing; Represents the probability of correlation between the feature vectors of node w and node e after pairing the video part.

5. The AI-driven automatic video content review method according to claim 4, characterized in that: The feature representation of each node in F and , calculate the prior probability respectively, and generate and The probability distribution of each violation situation is used as the characteristic feature of each node.

6. The AI-driven automatic video content review method according to claim 5, characterized in that: The audit process of each agent includes two rounds of audit phases: in the first phase, the agent traverses each node and obtains and For the probability distribution of each violation, if the probability of a violation is greater than a preset value, the video content at the node is marked and an early warning is issued; Otherwise, the second stage begins: using a pre-trained neural network to review the video content of each node in the knowledge graph and generate a probability of violation for each node; The voting according to the audit results of each agent includes: if any agent identifies that there is a node violation, it votes +1 for the identified illegal node; if the number of votes at the node is greater than 1, it is judged that there is an unstable part in the video content; if the number of votes at the node is 0, it is judged that the video content has passed the review; The preference library is for each edge relationship The preset weight value range is such that the sum of the weights of the four elements matched each time is 1; The agent audits every two nodes in each sequence. In the example, find the maximum and minimum values, and generate the guide vector of the node feature according to the maximum and minimum values; splice the two guide vectors obtained from the audio and video parts; obtain two splicing vectors by splicing in different orders, and use and as weight; The pre-trained neural network is used to determine the probability distribution of each guide vector and each violation, as well as the probability distribution of each splicing vector and each violation. The weighted summation is performed using the weight preference of the intelligent agent to obtain the final probability distribution result.

7. The AI-driven automatic video content review method according to claim 6, characterized in that: The analysis of the violation space includes, when determining that the node content is in violation, counting the proportion of different violations based on the voting results; Generate the distribution of illegal content in different violation situations, insert corresponding violation labels for illegal nodes, and issue warnings; When the video content review report is generated, the node range of the violation label is narrowed down, and only the lowest node referred to by each violation label is retained.

8. An AI-driven automatic video content review system, based on the AI-driven automatic video content review method according to any one of claims 1 to 7, characterized in that: The system comprises: a collection unit for acquiring video information to be reviewed and performing multi-level segmentation on the video information; The simulation unit simulates nodes in the slice according to the content independence of the video information; and analyzes the associated branches and historical features of each node to generate attribute features of the node; The analysis unit conducts content review through the review strategy of multiple virtual agents, and analyzes the violation space based on the unstable parts in the review results; The output unit generates a final video content review report based on the review results and the analysis results of the illegal space.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the AI-driven automatic video content review method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the AI-driven automatic video content review method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Video auditing method based on multiple levels and multiple models, medium and computer equipment

    CN111385602A

  • Short video auditing method based on multiple modes

    CN115512259A

  • Video content auditing method and system

    CN117173608A

  • Multi-mode network content security intelligent auditing system and method thereof

    CN118312922A

  • Video auditing method based on multi-modal large model

    CN118968380A

Cited By

  • Security test method, system and equipment for text video model and medium

    CN122019395A