Method and device for identifying video violation risk based on multi-agent cooperation, computer device, storage medium and computer program product

CN122761256APending Publication Date: 2026-09-15广东南方新媒体股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610996342.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-15

Smart Images

  • Figure CN122761256A_ABST
    Figure CN122761256A_ABST
Patent Text Reader

Abstract

The application provides a generative video violation risk identification method and device based on multi-agent cooperation, a computer device, a storage medium and a computer program product. The method comprises the following steps: obtaining a generative video to be audited, and performing preliminary auditing on the generative video based on a preset auditing rule; in the case that the generative video passes the preliminary auditing, auditing the generative video by multiple agents to obtain an auditing result output by each agent; the auditing result comprises a confidence score; the confidence score represents the certainty degree of the agent determining that the generative video contains violation content; determining a comprehensive confidence score according to the confidence score output by each agent; and determining the violation risk of the generative video according to the comprehensive confidence score. The method can accurately detect the implicit violation risk of the generative video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to a generative video violation risk identification method, apparatus, computer equipment, storage medium and computer program product based on multi-agent collaboration. Background Technology

[0002] With the rapid development of generative artificial intelligence technology, a large number of generative videos created using video generation models have emerged, and the content of these generative videos is diverse.

[0003] Currently, the identification of violations in generated videos largely relies on single-modal analysis methods, such as image feature matching for visuals, keyword detection for audio, and sensitive word filtering for subtitles. However, the violation risk characteristics of generated videos often differ from those of naturally shot videos. Some generated videos may not exhibit obvious violation characteristics in their individual modalities, but when the content from all three modalities is combined, it can convey satirical, distorted, or allegorical meanings. When traditional violation risk identification methods are used to identify violations in generated videos, it is often difficult to detect these implicit violations, resulting in a large number of potentially harmful generated videos passing review and detection.

[0004] Therefore, traditional technologies have the problem of being unable to accurately detect the hidden risks of violations in generated videos. Summary of the Invention

[0005] Based on this, the purpose of this application is to at least solve one of the above-mentioned technical defects, especially the technical defect that the prior art cannot accurately detect the implicit violation risk of generative videos. This application provides a generative video violation risk identification method, device, computer equipment, storage medium and computer program product based on multi-agent collaboration.

[0006] Firstly, this application provides a generative video violation risk identification method based on multi-agent collaboration, including: Obtain generated videos to be reviewed, and conduct a preliminary review of the generated videos based on preset review rules; If the generative video passes the initial review, multiple agents review the generative video and obtain the review results output by each agent. The review results include a confidence score, which represents the degree of certainty with which the agent determines that the generative video contains illegal content. The overall confidence score is determined based on the confidence scores output by each agent. The risk of violations in generated videos is determined based on the overall confidence score.

[0007] In an exemplary embodiment, a comprehensive confidence score is determined based on the confidence scores output by each agent, including: Determine the maximum difference between the confidence scores output by each agent; If the maximum difference is less than or equal to the preset threshold, the confidence score output by each agent will be used as the target confidence score output by each agent. If the maximum difference is greater than the preset threshold, the confidence score output by each agent is corrected based on the collaborative debate results among the agents, and the corrected confidence score output by each agent is obtained. The corrected confidence score output by each agent is then used as the target confidence score output by each agent. Obtain the confidence weight coefficients of each agent, and determine the comprehensive confidence score based on the confidence weight coefficients of each agent and the target confidence score output by each agent.

[0008] In an exemplary embodiment, a comprehensive confidence score is determined based on the confidence weight coefficients of each agent and the target confidence score output by each agent, including: The average confidence score is determined based on the target confidence scores output by each agent. Based on the target confidence score and the average confidence score output by each agent, determine the confidence bias corresponding to each agent; Based on the confidence deviation of each agent, the consensus fit of each agent is determined; the consensus fit is used to measure the degree of matching between the agent's review conclusion and the group consensus. The overall confidence score is determined based on the confidence weight coefficient of each agent, the target confidence score output by each agent, and the consensus fit of each agent.

[0009] In an exemplary embodiment, if the maximum difference is greater than a preset threshold, the confidence scores output by each agent are corrected based on the collaborative debate results among the agents, resulting in corrected confidence scores for each agent, including: If the maximum difference is greater than a preset threshold, then the high-scoring agent and the low-scoring agent that generated the maximum difference are identified; the high-scoring agent is the agent with the highest confidence score; the low-scoring agent is the agent with the lowest confidence score. A high-scoring agent can ask a low-scoring agent questions related to the violation. Extracting evidence fragments from generative videos using low-scoring agents; Information calibration is performed based on evidence fragments using high-scoring and low-scoring agents, and the latest confidence scores output by the high-scoring and low-scoring agents are obtained. Recalculate the maximum difference based on the latest confidence scores of all agents' outputs; If the maximum difference is greater than the preset threshold but the preset maximum number of iterations has not been reached, return to the step of determining the high-scoring agent and the low-scoring agent that generated the maximum difference. If the maximum difference is less than or equal to a preset threshold or reaches a preset maximum number of iterations, the latest confidence score of each agent's output is used as the corrected confidence score of each agent's output.

[0010] In one exemplary embodiment, determining the violation risk of the generated video based on a comprehensive confidence score includes: If the overall confidence score is less than the first threshold, then the generated video is determined to have no risk of violation. If the overall confidence score is greater than or equal to the first threshold and less than the second threshold, then the generated video is determined to contain content with low risk of violation; the second threshold is greater than the first threshold. If the overall confidence score is greater than or equal to the second threshold and less than the third threshold, then the generated video is determined to contain content with a risk of violation; the third threshold is greater than the second threshold. If the overall confidence score is greater than or equal to the third threshold, then the generated video is determined to contain content with a high risk of violation.

[0011] In one exemplary embodiment, the intelligent agents include a visual semantic extraction agent, a cross-modal logic verification agent, a fact-checking agent, and an emotional metaphor analysis agent; the review results also include a review conclusion characterizing whether the generative video violates regulations; the generative video is reviewed by multiple intelligent agents, and the review results output by each agent are obtained, including: The visual semantic extraction agent extracts the visual semantic labels of each key frame in the generative video, and identifies whether there are violations in the generative video based on the visual semantic labels of each key frame; the visual semantic labels include at least event subject information, scene nature information, narrative meaning information, and violation risk tendency information. After extracting multimodal features from the generative video through a cross-modal logic verification agent, modal conflict detection is performed to obtain modal conflict detection results. Based on the modal conflict detection results, it is determined whether there are violations in the generative video. The modal conflict detection results include modal conflict type information, modal conflict location information, modal conflict intensity information, and modal conflict segment information. The fact-checking agent extracts the structured factual information corresponding to each event from the generative video and inputs it into a pre-built knowledge base for information comparison to obtain a fact-checking report. Based on the fact-checking report, it identifies whether there are violations in the generative video. The fact-checking report includes information on forgery points, distortion points, incorrect timelines, and incorrect citations corresponding to forged events. After extracting the emotional features of generative videos across multiple modalities using an emotional metaphor analysis agent, intermodal emotional contrast detection is performed to obtain the intermodal emotional contrast detection results. Based on these results, the presence of violations in the generative videos is identified. The intermodal emotional contrast detection results include information on implicit emotional risk types, implicit emotional expression techniques, and violation risk levels.

[0012] Secondly, this application provides a generative video violation risk identification device based on multi-agent collaboration, comprising: The acquisition module is used to acquire generated videos to be reviewed and to conduct a preliminary review of the generated videos based on preset review rules. The collaborative review module is used to review the generated video through multiple agents after the initial review has been passed, and to obtain the review results output by each agent. The review results include a confidence score, which represents the degree of certainty with which the agent determines that the generated video contains illegal content. The determination module is used to determine the overall confidence score based on the confidence scores output by each agent; The identification module is used to determine the risk of violations in the generated video based on the comprehensive confidence score.

[0013] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0014] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.

[0015] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.

[0016] As can be seen from the above technical solutions, the embodiments of this application have the following advantages: This application provides a method, apparatus, computer device, storage medium, and computer program product for identifying violations in generative videos based on multi-agent collaboration. The method acquires a generative video to be reviewed and performs a preliminary review based on preset review rules. If the generative video passes the preliminary review, multiple agents review the video, obtaining review results output by each agent. The review results include a confidence score, which represents the degree of certainty with which the agents determine that the generative video contains violations. A comprehensive confidence score is determined based on the confidence scores output by each agent. Finally, based on the comprehensive confidence score, the risk of violation in a generative video is further determined. The system identifies the violation risks of generative videos. Through initial review, explicit violations are quickly filtered out. Then, multiple agents conduct in-depth collaborative review of the generative videos from different dimensions, and the confidence scores of each agent are combined to obtain a comprehensive confidence score. This ultimately quantifies the violation risk, effectively solving the problem that existing technologies can only detect explicit violations but cannot identify implicit violations such as compliant visuals or audio but violating rules when combined. The division of labor and collaboration among multiple agents, along with the confidence fusion mechanism, enables the review system to accurately capture deep semantic conflicts and metaphorical risks that single-modal analysis cannot detect, thus achieving accurate detection of implicit violation risks in generative videos. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a generative video violation risk identification method based on multi-agent collaboration provided in this application embodiment; Figure 2 A data processing flowchart for a generative video violation risk identification method based on multi-agent collaboration provided in this application embodiment; Figure 3 A flowchart illustrating another generative video violation risk identification method based on multi-agent collaboration provided in this application embodiment; Figure 4 A structural block diagram of a generative video violation risk identification device based on multi-agent collaboration provided in an embodiment of this application; Figure 5 This is a schematic diagram of the internal structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] With the explosion of generative artificial intelligence (AIGC) and deep synthesis technology, the way content is produced has undergone fundamental changes. Existing technologies have certain shortcomings in facing new challenges of content violations, such as: (1) Lack of cross-modal semantic association capabilities. Existing technologies process video, audio, and text separately. In generative content, it is common to see "normal visuals (such as landscape videos) + normal voice-over (such as popular science tone) + normal subtitles (no sensitive words)," but the combination of the three expresses illegal semantics such as "distorting history, falsifying facts, and being sarcastic." For example, using AI-generated fake historical images with a serious narration, each modality may appear normal on its own, but the whole constitutes historical nihilism. Traditional technologies cannot identify such "implicit violations."

[0021] (2) The rule base is outdated and cannot cope with new types of violations. The matching method based on the word database and feature database is passive. For new types of violations such as metaphors, irony, and indirect criticism generated by AIGC, as well as the constantly evolving "slang", the rule base is difficult to update and cover in a timely manner.

[0022] (3) High false alarm rate and lack of contextual understanding. Traditional solutions are prone to false alarms due to the appearance of sensitive elements or slips of the tongue in the picture, because they cannot judge the true intention of the content by combining the context.

[0023] (4) Unable to handle multimodal adversarial attacks. Malicious content creators may intentionally exploit the differences between modalities to conduct adversarial attacks, such as overlaying high-noise music on sensitive speech or covering a large amount of text on sensitive images, causing the single-modal detection model to fail.

[0024] To address the aforementioned issues, this application provides a generative video violation risk identification method based on multi-agent collaboration, and specific descriptions of various embodiments are provided below.

[0025] In one exemplary embodiment, Figure 1 This is a flowchart illustrating a generative video violation risk identification method based on multi-agent collaboration, as provided in an embodiment of this application. Figure 1 As shown, a generative video violation risk identification method based on multi-agent collaboration is provided. Taking the application of this method to a server as an example, the method includes the following steps S102 to S108. Wherein: Step S102: Obtain the generated video to be reviewed, and conduct a preliminary review of the generated video based on the preset review rules.

[0026] Generative video refers to video content that is not actually filmed and is generated by artificial intelligence models (such as diffusion models, generative adversarial networks, etc.).

[0027] Among them, the preset review rules refer to the preliminary filtering rules formulated based on laws, policies or platform specifications, which are used to quickly screen out basic content that is obviously in violation of regulations.

[0028] The preliminary review refers to the first round of screening of videos using a simple keyword matching lightweight classification model.

[0029] Optionally, the server first obtains the generated video to be reviewed and calls a preset basic rule base to conduct a preliminary review.

[0030] Step S104: If the generative video passes the preliminary review, multiple agents review the generative video and obtain the review results output by each agent. The review results include a confidence score. The confidence score represents the degree of certainty that the agent determines that the generative video contains illegal content.

[0031] Among them, an intelligent agent refers to an independent reasoning model that is built based on a large language model or a multimodal large model and has specific dimension review capabilities.

[0032] The review result is a structured judgment output by each agent, which includes at least a confidence score.

[0033] The confidence score is a quantitative value representing the degree of certainty each agent has about the audit conclusion of its output (such as the existence of violations), and its value typically ranges from 0 to 1.

[0034] Optionally, if a video fails the initial review (e.g., contains obviously prohibited elements), it is directly marked as a violation and the process ends. If the video passes the initial review, multiple agents are invoked in parallel to conduct a deep review of the video. Each agent outputs its own review result, which includes a confidence score indicating its degree of certainty in determining that the video violates regulations.

[0035] Step S106: Determine the overall confidence score based on the confidence scores output by each agent.

[0036] The overall confidence score is a single total score obtained by weighting, fusing, or debate-correcting the confidence scores of all agents, and is used to quantify the overall risk of violation.

[0037] Optionally, the server merges the confidence scores of all agents to obtain a comprehensive confidence score.

[0038] Step S108: Determine the violation risk of the generated video based on the comprehensive confidence score.

[0039] Among them, the risk of violation refers to the risk level divided according to the comprehensive confidence score, which is used to determine whether the content is approved, reminded, manually reviewed or blocked.

[0040] Optionally, the server determines the violation risk level of the video (such as no risk, low risk, medium risk, high risk) based on the overall confidence score. Based on the violation risk level, it can determine whether the generated video can pass the review, whether a reminder is needed, and whether manual review or blocking is required.

[0041] The aforementioned generative video violation risk identification method based on multi-agent collaboration acquires the generative video to be reviewed and performs a preliminary review based on preset review rules. If the generative video passes the preliminary review, multiple agents review the video, obtaining review results from each agent. These results include a confidence score, which represents the degree of certainty with which the agent determines the presence of violation content in the generative video. A comprehensive confidence score is determined based on the confidence scores output by each agent. Finally, the violation risk of the generative video is determined based on the comprehensive confidence score. Therefore, by quickly filtering explicit violations through initial review, multiple agents are then used to conduct in-depth collaborative review of generative videos from different dimensions. The confidence scores of each agent are combined to obtain a comprehensive confidence score, which ultimately quantifies the violation risk. This effectively solves the problem that existing technologies can only detect explicit violations but cannot identify implicit violations such as compliant video or audio content that violates regulations when combined. The division of labor and collaboration among multiple agents and the confidence fusion mechanism enable the review system to accurately capture deep semantic conflicts and metaphorical risks that cannot be found by single-modal analysis, thereby achieving accurate detection of implicit violation risks in generative videos.

[0042] In an exemplary embodiment, determining a comprehensive confidence score based on the confidence scores output by each agent includes: determining the maximum difference between the confidence scores output by each agent; if the maximum difference is less than or equal to a preset threshold, then using the confidence scores output by each agent as the target confidence scores output by each agent; if the maximum difference is greater than the preset threshold, then correcting the confidence scores output by each agent through the collaborative debate results among the agents to obtain corrected confidence scores output by each agent, and using the corrected confidence scores output by each agent as the target confidence scores output by each agent; obtaining the confidence weight coefficients of each agent, and determining the comprehensive confidence score based on the confidence weight coefficients of each agent and the target confidence scores output by each agent.

[0043] The maximum difference refers to the difference between the maximum and minimum confidence scores of all agents. The preset threshold is a pre-defined upper limit for tolerance of divergence, such as 0.3.

[0044] The target confidence score refers to the confidence score used for final fusion after the divergence judgment. When the divergence is small, the target confidence score is the original confidence score, and the original confidence score of agent i can be represented as score_i. When the divergence is large, the target confidence score is the corrected confidence score, and the corrected confidence score of agent i can be represented as score_i'.

[0045] Among them, the corrected confidence score refers to the confidence score after being corrected through collaborative debate (questioning, evidence presentation, calibration) between agents.

[0046] Among them, the confidence weight coefficient is a preset importance weight for each agent, reflecting the relative credibility of its judgment in the overall fusion. The confidence weight coefficient can be pre-configured and dynamically updated. The confidence weight coefficient of agent i can be represented as weight_i.

[0047] Among them, the collaborative debate result refers to the consensus or revised conclusion reached by multiple agents after calibrating information on the content of disagreement through multiple rounds of questioning and answering.

[0048] Optionally, after obtaining the original confidence scores of each agent, the server first calculates the maximum difference between them. If the difference does not exceed a preset threshold (e.g., 0.3), the agents are considered to have reached a consensus, requiring no correction, and their original confidence scores are directly used as the target confidence score. If the maximum difference exceeds the threshold, a collaborative debate mechanism is triggered: each agent corrects its own confidence score through multiple rounds of interaction, obtaining a corrected confidence score. Subsequently, the confidence weight coefficient is obtained, and the confidence weight coefficient can be weighted and summed (or more complexly fused) with the target confidence score to calculate the final comprehensive confidence score.

[0049] In this embodiment, by introducing a disagreement detection and collaborative debate correction mechanism based on the maximum difference, the problem of inaccurate results caused by simple averaging or weighting when multiple agents make inconsistent judgments is solved. When the judgments of the agents are the same, they are directly fused, which is highly efficient. When there is a major disagreement, debate between agents is actively triggered, and the collective wisdom is used to eliminate ambiguity and correct biases, making the corrected confidence score more reliable. Then, the confidence weight coefficient is combined for differentiated fusion, which significantly improves the accuracy and robustness of the comprehensive confidence score, thereby more accurately reflecting the real violation risk of the generative video.

[0050] In an exemplary embodiment, determining the comprehensive confidence score based on the confidence weight coefficient of each agent and the target confidence score output by each agent includes: determining the average confidence score based on the target confidence score output by each agent; determining the confidence deviation for each agent based on the target confidence score and the average confidence score output by each agent; determining the consensus fit for each agent based on the confidence deviation; the consensus fit is used to measure the degree of matching between the agent's review conclusion and the group consensus; and determining the comprehensive confidence score based on the confidence weight coefficient of each agent, the target confidence score output by each agent, and the consensus fit for each agent.

[0051] The average confidence score refers to the arithmetic mean of the target confidence scores of all agents, which can be represented as avg_score.

[0052] Here, confidence bias is the difference between the target confidence score and the average confidence score of a single agent. The confidence bias corresponding to agent i can be expressed as deviation_i, where deviation_i = |score_i - avg_score|.

[0053] Among them, consensus fit measures the degree of matching between the agent's review conclusion and the group consensus. Consensus fit can be understood as consistent support. The consensus fit for agent i can be represented as support_i, where support_i = 1 - deviation_i. The larger the deviation_i, the lower the consensus fit support_i, indicating that agent i deviates further from the collective consensus.

[0054] In this embodiment, the overall confidence score can be calculated using the fusion formula Final Score = Σ The calculations show that the formula takes into account the inherent importance of the agent, its judgment value, and its fit with the group.

[0055] Optionally, after obtaining the target confidence scores of each agent, the server calculates their arithmetic mean to obtain the average confidence score. Then, for each agent, the server calculates the absolute difference between its score and the average to obtain the confidence bias. Next, the server calculates the consensus fit for each agent using the formula: consensus fit = 1 - confidence bias. Finally, the server merges the confidence weight coefficients, target confidence scores, and consensus fits of each agent to obtain the final comprehensive confidence score.

[0056] It should be noted that the fusion formula corresponding to the comprehensive confidence score in this application can be "Comprehensive Confidence Score = Σ(Weight i × Target Confidence i × Consensus Fit_i) / Σ(Weight_i × Consensus Fit_i)". By using the denominator Σ(Weight_i × Consensus Fit_i), Σ(Weight i × Target Confidence i × Consensus Fit_i) can be renormalized, making its value more reflective of the judgment of high-fit agents and reducing the interference of low-fit agents.

[0057] In this embodiment, by introducing a consensus fit factor into the comprehensive confidence score calculation, the undue influence of a few outlier agents (who make incorrect judgments or are subject to interference) on the final result is effectively suppressed. If an agent's conclusion deviates significantly from that of the majority of agents, its consensus fit is low, and its contribution to the final score will be automatically reduced. This avoids veto power or noise interference, and solves the problem that simple weighted averages are susceptible to extreme values ​​and cannot reflect collective wisdom. This makes the comprehensive confidence score more robust and closer to the real risk level, further improving the accuracy of violation risk detection.

[0058] In an exemplary embodiment, if the maximum difference is greater than a preset threshold, the confidence scores output by each agent are corrected based on the collaborative debate results among the agents to obtain corrected confidence scores. This includes: if the maximum difference is greater than the preset threshold, identifying the high-scoring agent and the low-scoring agent that generated the maximum difference; the high-scoring agent is the agent with the highest confidence score; the low-scoring agent is the agent with the lowest confidence score; the high-scoring agent raises a question related to the violation to the low-scoring agent; and the low-scoring agent extracts evidence from the generative video. The process involves: 1) calibrating information based on evidence fragments using high-scoring and low-scoring agents, and obtaining the latest confidence scores output by the high-scoring and low-scoring agents; 2) recalculating the maximum difference based on the latest confidence scores output by all agents; 3) returning to the step of determining the high-scoring and low-scoring agents that generated the maximum difference if the maximum difference is greater than a preset threshold but the preset maximum number of iterations has not been reached; and 4) using the latest confidence scores output by each agent as the corrected confidence scores output by each agent if the maximum difference is less than or equal to a preset threshold or the preset maximum number of iterations has been reached.

[0059] Among them, the high-scoring agent refers to the agent with the highest confidence score in the current round.

[0060] Among them, the low-scoring agent refers to the agent with the lowest confidence score in the current round.

[0061] Among them, the questioning problem is a follow-up question posed by a high-confidence agent to a low-confidence agent, requiring the agent to explain the basis of its judgment.

[0062] Among them, evidence fragments refer to video frames, text fragments, or audio information extracted from generative videos by low-scoring agents to support their judgments.

[0063] Information calibration refers to the process by which both parties re-examine their judgments based on evidence and correct any misunderstandings.

[0064] The preset maximum number of iterations refers to the maximum number of interactions (e.g., 3 rounds) set to avoid infinite loops.

[0065] The latest confidence score refers to the confidence value updated and output by each agent after one round of calibration.

[0066] Optionally, collaborative debate is initiated when the maximum difference exceeds a preset threshold. First, the highest-scoring agent and the lowest-scoring agent are identified. The high-scoring agent poses a challenge to the low-scoring agent, such as, "Why do you believe this segment contains factual distortions? Please provide evidence." The low-scoring agent responds by extracting relevant evidence fragments (such as keyframes and ASR text time points) from the generative video and returning them. Both parties perform information calibration based on the evidence fragments: the high-scoring agent may lower its confidence due to insufficient evidence, while the low-scoring agent may increase its confidence due to insufficient evidence. Both parties output their adjusted latest confidence scores. Then, the server recalculates the maximum difference between the latest confidence scores of all agents. If the maximum difference is still greater than the preset threshold and the current maximum iteration round has not been reached, the above process is repeated (a new high / low-scoring agent may be selected in the new round). If the maximum difference drops below the preset threshold or the current maximum iteration round is reached, the debate stops, and the latest confidence scores of each agent are used as the corrected confidence scores.

[0067] In this embodiment, a multi-round debate process of "high-score questioning - low-score evidence presentation - information calibration" is designed to provide a proactive disagreement resolution mechanism among agents. This addresses the problem that numerical calculations alone cannot eliminate cognitive conflicts when multiple agents make inconsistent judgments. The debate process simulates the discussion mode of a human expert panel, prompting agents to correct erroneous perceptions or reach a consensus through evidence exchange and rational persuasion, resulting in a more accurate and evidence-based final confidence score. This mechanism is applicable to complex violation scenarios in generative videos with blurred boundaries and requiring cross-verification of multi-source evidence (such as manipulation and metaphorical satire), significantly improving the effectiveness of correction.

[0068] In one exemplary embodiment, determining the violation risk of a generated video based on a comprehensive confidence score includes: if the comprehensive confidence score is less than a first threshold, determining that the generated video has no violation risk; if the comprehensive confidence score is greater than or equal to the first threshold and less than a second threshold, determining that the generated video contains content with low violation risk; the second threshold is greater than the first threshold; if the comprehensive confidence score is greater than or equal to the second threshold and less than a third threshold, determining that the generated video contains content with medium violation risk; the third threshold is greater than the second threshold; if the comprehensive confidence score is greater than or equal to the third threshold, determining that the generated video contains content with high violation risk.

[0069] The first, second, and third thresholds are three pre-defined, progressively increasing risk thresholds used to map the overall confidence score to discrete levels of violation risk. In practical applications, the first threshold can be set to 0.3, the second threshold to 0.5, and the third threshold to 0.7.

[0070] The statement "No risk of violation" indicates that the content is safe and can be automatically approved.

[0071] Among them, "low risk of violation" indicates a slight or suspicious tendency to violate regulations, and the system can issue a warning or conduct spot checks.

[0072] Among them, the "Middle Violation Risk" indicates a significant possibility of violation and requires manual review.

[0073] Among them, "high risk of violation" indicates a high degree of certainty of serious violation and should be automatically blocked.

[0074] Optionally, after calculating the overall confidence score, the server compares it with three preset thresholds: if the score is lower than the first threshold (e.g., 0.3), it is determined to have no risk of violation, and the server automatically approves the review; if the score is between the first and second thresholds (e.g., 0.3~0.5), it is determined to have low risk of violation, and the server may issue a reminder, suggesting manual spot checks or automatic approval and recording; if the score is between the second and third thresholds (e.g., 0.5~0.7), it is determined to have medium risk of violation, and the server automatically transfers the process to manual review; if the score is higher than the third threshold (e.g., 0.7), it is determined to have high risk of violation, and the server directly blocks the video and triggers an alarm.

[0075] In this embodiment, by setting multiple threshold levels, the comprehensive confidence score is mapped to multiple risk levels, and corresponding to different handling strategies (automatic approval, reminder, manual review, and blocking), which realizes refined hierarchical management of violation risks. This solves the problem that the binary decision of simply approving or blocking cannot meet the differentiated handling of different levels of risk in actual business. Low-risk content can be quickly approved to reduce the manual burden, while medium- and high-risk content triggers a more stringent review process to ensure that the risk is controllable. This hierarchical mechanism ensures security while taking into account review efficiency, enabling the review system to flexibly adapt to different business scenarios and security requirements.

[0076] In an exemplary embodiment, the intelligent agents include a visual semantic extraction agent, a cross-modal logic verification agent, a fact-checking agent, and an emotional metaphor analysis agent; the review result also includes a review conclusion characterizing whether the generative video violates regulations; the generative video is reviewed by multiple agents, and the review results output by each agent are obtained, including: extracting visual semantic tags of each key frame in the generative video through the visual semantic extraction agent, and identifying whether there are violations in the generative video based on the visual semantic tags of each key frame; the visual semantic tags include at least event subject information, scene nature information, narrative meaning information, and violation risk tendency information; after extracting multimodal features of the generative video through the cross-modal logic verification agent, performing modal conflict detection to obtain modal conflict detection results, and identifying whether there are violations in the generative video based on the modal conflict detection results; the modal conflict detection results... The results include information on modal conflict type, location, intensity, and fragments. A fact-checking agent extracts structured factual information corresponding to each event from the generative video and inputs it into a pre-built knowledge base for comparison, generating a fact-checking report. Based on this report, the system identifies whether violations exist in the generative video. The fact-checking report includes information on forged points, distortions, incorrect timelines, and incorrect citations corresponding to forged events. An emotional metaphor analysis agent extracts the emotional features of the generative video across multiple modalities and performs intermodal emotional contrast detection, generating intermodal emotional contrast detection results. Based on these results, the system identifies whether violations exist in the generative video. The intermodal emotional contrast detection results include information on implicit emotional risk type, implicit emotional expression techniques, and violation risk level.

[0077] Among them, the visual semantic extraction agent is a reasoning unit based on a multimodal large model, used to understand the deeper narrative meaning of the image.

[0078] The visual semantic tags include at least the following: event subject information, scene nature information, narrative meaning information, and violation risk information. Event subject information refers to the core participants or objects identified in the image, such as a historical figure, protesters, or equipment. Scene nature information refers to the qualitative description of the environment in which the image is set, such as ruins, a natural disaster site, a gathering, or a medical emergency room. Narrative meaning information refers to the deeper story theme or tendency expressed by the sequence of images, such as creating a tragic atmosphere, inciting confrontation, or vilifying a specific figure. Violation risk information is a potential risk direction judged based on the above information, such as the possibility of distorting history.

[0079] In practical applications, visual semantic agents not only identify objects in a scene but also understand its narrative meaning (e.g., "This is a scene that appears peaceful but actually depicts the ruins of war"). Their visual narrative understanding and reasoning is based on a multimodal large model to perform deep semantic reasoning on the scene, outputting the scene's implicit intentions and narrative tendencies (positive / negative / metaphorical / suggestive), and generating structured results with visual semantic labels {event subject, scene nature, narrative meaning, risk tendency}. Visual semantic agents do more than just identify "what it is"; they understand "what it expresses and what it implies," enabling a deep understanding of metaphors and negative narrative packaging in sensitive scenes.

[0080] Among them, the cross-modal logic verification agent is an agent specifically designed to detect whether there are logical conflicts between different modal information (images, audio, and subtitles).

[0081] The modal conflict detection results include modal conflict type information, conflict location information, conflict intensity information, and conflict segment information. Modal conflict type information refers to the specific category of inconsistency between different information channels (visual, audio, subtitles), such as misattribution (a scene of one event accompanied by narration of another), misidentification (a scene of person A accompanied by the voice of person B), temporal and spatial distortion (the scene shows 2020, but the narration says 2023), and splicing fraud (forcibly splicing together scenes from different time periods). Modal conflict location information refers to the specific timestamp of the conflict location (e.g., "00:12:30 - 00:12:35") or the area of ​​the scene (e.g., the "lower left corner subtitle area"). Modal conflict intensity information refers to the quantitative value of the obviousness of the conflict or the severity of the logical contradiction (e.g., a score between 0 and 1, with higher scores indicating more severe conflicts). Modal conflict segment information refers to the video clip, audio clip, or subtitle sentence containing the conflict, used for evidence preservation and manual review.

[0082] In practical applications, cross-modal logic verification agents (key agents) are specifically used to detect conflicts between modalities. For example, they can determine that "the screen shows event A, but the narration describes event B", thereby identifying violations such as misrepresentation and misinterpretation. The execution process includes: (1) decoupling and extraction of modal information, that is, extracting visual semantics (from the visual agent), text semantics (ASR speech transcription), and audio semantics (tone, background sound); (2) cross-modal entity alignment, that is, unifying and associating entities such as people, places, times, events, and objects; (3) modal conflict detection, judging consistency dimension by dimension, including whether the screen event is consistent with the text description, whether the screen time and place match the voice narration, and whether the screen subject and the subtitle refer to the same thing; (4) outputting the conflict type judgment result, such as misrepresentation, misinterpretation, misattribution, spatiotemporal confusion, and splicing forgery; (5) conflict confidence scoring, forming {conflict type, conflict location, conflict intensity, evidence fragment}. Cross-modal logic verification agents can directly identify difficult-to-distinguish types of violations such as "misrepresentation, misinterpretation, and fabricated rumors."

[0083] Among them, the fact-checking agent is an intelligent agent that compares entities and events extracted from videos with external knowledge bases.

[0084] The fact-checking report includes information on falsified details, distorted details, incorrect timelines, and incorrect citations related to the fabricated events. Structured factual information refers to breaking down the events presented in the video into standardized, comparable triplets, such as <subject, action, object / time / location>, like or <an event occurred in 2024>. Falsified details refer to sub-information presented in the video that is completely inconsistent with the facts in the knowledge base, such as "claiming that event A occurred in 2024, but actually occurred in 2020". Distorted details refer to sub-information in the video that selectively conceals, exaggerates, or distorts facts, such as "only exposing the accident while omitting subsequent rescue efforts". Incorrect timelines refer to erroneous descriptions in the video that do not match the actual historical timeline. Incorrect citations refer to sub-information in the video that incorrectly cites data, names of people, or names of organizations.

[0085] In practical applications, the fact-checking agent compares the extracted entities and events with external knowledge bases (such as historical databases and news databases) to identify the authenticity of content such as distorted facts and fabricated history. Its execution process includes: (1) Structural processing of factual statements to generate a triplet to be checked <subject, behavior, object / time / location>. (2) Knowledge base comparison, connecting with the knowledge base of the positive sample feedback and adaptive optimization module. This knowledge base has been refined based on a large amount of existing news, documentaries, films and TV dramas, including historical fact bases, authoritative news bases, official release bases, and geographical / time / person standard bases. (3) Fact authenticity determination. When the determination is consistent, it indicates that the fact is credible. When the determination is inconsistent, it indicates that the fact is distorted or fabricated. When the determination is without source, it indicates that the statement is unfounded. (4) Generate fact-checking report annotations, including forgery points, distortion points, incorrect timelines, and incorrect citations. The fact-checking intelligent agent can directly identify falsified history, distorted events, fabricated data, and false citations, enabling the review of generative videos in broadcast television and online audiovisual media to leap from "format compliance" to "factual accuracy." Among them, the emotional metaphor analysis agent is an agent that identifies hidden risks by analyzing the contrast between text sentiment, voice tone, and visual emotion.

[0086] The intermodal emotional contrast detection results include information on implicit emotional risk types, implicit emotional expression techniques, and violation risk levels. Implicit emotional risk types refer to the risk categories presented through emotional contrast or implicit expression, such as "implicitly inciting antagonism" or "using sarcasm to glorify an incident." Implicit emotional expression techniques refer to the specific artistic or technical methods used to achieve the aforementioned implicit emotions, such as reciting cheerful and festive texts in a "crying" (sad tone), narrating extremely terrifying texts at a calm pace, or pairing cheerful music with disaster scenes. Violation risk levels refer to the level comprehensively assessed based on the prominence and potential harm of the implicit emotions, such as "no risk," "low risk (requires warning)," and "high risk (requires interception)."

[0087] In practical applications, the emotional metaphor analysis agent identifies implicit attacks and sarcasm by analyzing the contrast between the tone of voice and the content of the text (such as "reciting cheerful words in a sad tone to express sarcasm"). Its execution process includes: (1) Multi-dimensional emotional feature extraction, extracting information such as text emotion (positive / negative / sarcasm / incitement), tone of voice (speech speed, emphasis, emotion, crying / laughing / mocking), and visual emotion (atmosphere, lighting, compositional cues). (2) Intermodal emotional contrast detection and judgment, if the tone of voice is sad and the text is cheerful, it indicates that there is sarcasm or irony; if the tone of voice is calm and the text is extreme, it indicates that there is implicit incitement; if the visual is positive and the subtitles are negative, it indicates that there is implicit attack. (3) Metaphor and suggestion recognition, identifying sarcasm, insinuation, implicit insults, metaphorical incitement, etc. (4) Implicit risk judgment output, the output content is {implicit risk type, expression method, risk level}. The emotional metaphor analysis agent identifies phenomena such as implicit sarcasm, irony, metaphorical attacks, and irony, breaking through the limitation of traditional review that can only identify explicit violations.

[0088] Optionally, the server can invoke the four agents mentioned above in parallel for review. The visual semantic extraction agent infers from the keyframes of the video, outputs visual semantic labels (such as event subject: a historical figure; scene nature: war ruins; narrative meaning: tragic; risk tendency: negative rendering), and judges violations accordingly. The cross-modal logic verification agent extracts on-screen events, audio narration, and subtitle text, performs entity alignment and consistency checks, outputs modal conflict detection results (e.g., conflict type: misrepresentation; conflict location: 00:12:30; conflict intensity: high; evidence fragment: ...), and determines a violation. The fact-checking agent compares the factual triples (subject, behavior, object / time) stated in the video with the knowledge base, outputs a fact-checking report (e.g., forgery point: claiming the event occurred in 2020, but actually in 2015; distortion point: fabricated data), and determines a violation. The sentiment metaphor analysis agent analyzes the contrast between text sentiment (positive / negative), voice tone (sad / calm), and on-screen emotion, outputs inter-modal sentiment contrast detection results (e.g., implicit risk type: irony; expression method: crying while reciting celebratory text; risk level: high), and determines a violation. Each agent ultimately outputs a review result including a confidence score and a review conclusion.

[0089] In this embodiment, by deploying four independent and complementary dedicated intelligent agents (visual semantics, cross-modal logic, fact-checking, and emotional metaphor), a comprehensive review system covering content understanding, logical consistency, factual authenticity, and implicit emotions is constructed. This solves the core problem of traditional review methods that can only identify explicit violations but cannot detect deep-seated risks. Each intelligent agent conducts in-depth analysis from the perspectives of narrative intent, modal conflict, factual authenticity, and emotional contrast, which can accurately discover complex violations in generative videos that cannot be detected by single-modal analysis, thereby greatly improving the comprehensiveness and accuracy of detection.

[0090] For the convenience of those skilled in the art, Figure 2 An exemplary data processing flowchart for a generative video violation risk identification method based on multi-agent collaboration is provided. The corresponding generative video violation risk identification process includes the following steps: Step 1: Video injection, inputting the generative video to be reviewed into the review system; Step 2: The review system performs a preliminary review of the video based on preset review rules: if the video fails the preliminary review, it is directly determined as a video that fails the review, and the process ends; if the video passes the preliminary review, it enters the subsequent in-depth review process; Step 3: The video that passes the preliminary review enters the multimodal data parsing and alignment module. This module is used to preprocess and fine-grained align the input audio and video content. It performs image feature extraction, audio feature extraction, and text feature extraction on the video, and fuses and aligns the extracted multimodal features to form a unified feature representation. Specifically, it extracts keyframes based on visual flow, extracts text from the screen using OCR, extracts entities from the screen using object detection, and performs surface meaning recognition (recognition). The process involves identifying objects, scenes, people, actions, and text, sampling keyframes of the input video to extract temporal relationships, and performing ASR (Audio Recognition) based on auditory streams to extract voiceprint features (emotions, intonation). The video content, audio content, and subtitle content are then aligned at the millisecond level along the timeline to construct spatiotemporal semantic units. Step 4: The aligned features are input to a multimodal content review agent cluster, where multiple agents (such as visual semantic agents, cross-modal logic verification agents, fact-checking agents, and sentiment metaphor analysis agents) conduct multi-dimensional in-depth review. Step 5: The outputs of each agent are fed into an agent coordination and confidence-based hybrid scoring module. Through collaborative debate and confidence fusion calculation, a comprehensive confidence score is generated, and the risk level is determined. If the video is deemed a suspected violation, it proceeds to the manual review stage; if it is deemed a video that passed machine review, it passes the review directly. Step 6: The manual review stage re-examines the suspected violation video to ultimately confirm whether it is a confirmed violation video. The above process constructs a complete automated review chain, from initial screening, multimodal feature analysis, in-depth review of intelligent agent clusters to confidence fusion and manual review.

[0091] This application also provides an agent adaptive optimization method based on human review results. Its core logic is to use the human reviewer's judgment of suspected violations as a feedback signal to the review system, serving as a reward signal for reinforcement learning (RLHF), and dynamically adjusting the weights of each agent to enable the system to adapt to emerging new violation methods. The method includes the following steps: Step 1: Sample Classification and Labeling. Historical audit results are divided into two categories: positive samples are those with correct audit conclusions, no complaints, and no missed violations; negative samples are those with misjudgments, missed judgments, incorrect judgments, failure to identify misrepresentations, or failure to detect factual errors.

[0092] Step 2: RLHF reward signal generation. A positive-negative differentiated reward mechanism is adopted: for positive samples, a positive reward R is given when the judgment is correct. + =1.0 (fixed reward, simple and stable), for negative samples, a negative penalty R is applied when a misclassification occurs. - =-1.0 (fixed penalty, simple and controllable).

[0093] Step 3: Distribute rewards or penalties based on agent contributions. For each agent i, calculate its adjusted reward value r. i =R×c i Where R is the global reward or penalty (R + Or R - ), c i This represents the agent's contribution to the judgment (value from 0 to 1). This mechanism ensures that agents who make correct judgments are rewarded, while agents who make incorrect judgments are punished.

[0094] Step 4: Dynamically adjust the agent weights. Use the weight update formula, where the new weight W... i_new =W i_old ×(1+r i The intuitive logic is: an agent r that provides a positive sample and makes a correct judgment. i For positive samples, the weight increases; for negative samples and incorrect judgments, the agent r... i If the value is negative, the weight decreases; the weight of agents that align with global judgments remains stable; the weight of agents that make abnormal judgments is automatically reduced.

[0095] Step 5: Weight Normalization. To ensure system stability, all weights are summed and then normalized to ensure that the sum of the weights of each agent is 1: W i =W i_new / ∑W i_new .

[0096] Through the aforementioned closed-loop feedback mechanism, the review system can continuously optimize the weights of each agent using the results of manual review, thereby continuously improving the accuracy and adaptability of detecting violations in generative videos.

[0097] Compared with existing technologies, the significant advantages of this invention include: (1) accurate identification of implicit violations: existing technologies can only detect explicit violations. This invention, through cross-modal logic verification agents and multimodal semantic fusion, can effectively identify new types of violations (such as distorted history and misrepresentation) that are compliant with visual and audio standards but violate regulations when combined, thus filling the gap in generative content review; (2) deep semantic understanding and reasoning capabilities: existing technologies are mostly based on keyword matching and lack understanding capabilities. This invention, based on large model agents, has common sense reasoning and contextual understanding capabilities, and can understand irony, metaphors, and indirect criticism, significantly reducing the false negative rate of adversarial examples; (3) strong interpretability of review results: traditional deep learning models are usually black boxes, while the agent collaboration process in this invention will output a complete reasoning chain (e.g., "It is determined to be a violation because the visual agent recognizes..."). The video shows the 2023 flood, but the audio agent identifies the narration as an event in 2024, and the logic agent judges it as a rumor. This provides solid evidence for manual review. (4) High dynamic adaptability and robustness: Compared with the existing technology with slow rule base updates, this invention adopts a confidence-based mixed scoring mechanism, which allows the agent weights to be adjusted to deal with new violation trends without retraining the model. At the same time, the multi-agent architecture enables the review system to make a comprehensive judgment based on other modal agents when a certain modality is disturbed (such as excessive noise), which has extremely high robustness. (5) Balancing efficiency and accuracy: Existing full-scale large model inference is often slow. This invention adopts a hierarchical strategy, that is, simple content is quickly filtered through lightweight rules, and complex content triggers multi-agent collaborative analysis, thereby effectively controlling the computational cost while ensuring high accuracy.

[0098] In another embodiment, such as Figure 3 As shown, a generative video violation risk identification method based on multi-agent collaboration is provided. Taking the application of this method to a server as an example, the method includes the following steps: Step S302: Obtain the generated video to be reviewed, and conduct a preliminary review of the generated video based on the preset review rules.

[0099] Step S304: If the generative video passes the preliminary review, multiple agents review the generative video and obtain the review results output by each agent. The review results include a confidence score. The confidence score represents the degree of certainty with which the agent determines that the generative video contains illegal content.

[0100] Step S306: Determine the maximum difference between the confidence scores output by each agent.

[0101] In step S308, if the maximum difference is less than or equal to a preset threshold, the confidence score output by each agent is taken as the target confidence score output by each agent.

[0102] Step S310: If the maximum difference is greater than a preset threshold, the confidence score output by each agent is corrected based on the collaborative debate results between the agents to obtain the corrected confidence score output by each agent, and the corrected confidence score output by each agent is used as the target confidence score output by each agent.

[0103] Step S312: Obtain the confidence weight coefficients of each agent, and determine the average confidence score based on the target confidence score output by each agent.

[0104] Step S314: Determine the confidence bias for each agent based on the target confidence score and the average confidence score output by each agent.

[0105] Step S316: Determine the consensus fit degree for each agent based on the confidence deviation for each agent; the consensus fit degree is used to measure the degree of matching between the agent's review conclusion and the group consensus.

[0106] Step S318: Determine the comprehensive confidence score based on the confidence weight coefficient of each agent, the target confidence score output by each agent, and the consensus fit of each agent.

[0107] Step S320: Determine the violation risk of the generated video based on the comprehensive confidence score.

[0108] It should be noted that the specific limitations of the above steps can be found in the above description of the specific limitations of a generative video violation risk identification method based on multi-agent collaboration.

[0109] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0110] The following describes the generative video violation risk identification device based on multi-agent collaboration provided in the embodiments of this application. The generative video violation risk identification device based on multi-agent collaboration has the same inventive concept as the generative video violation risk identification method based on multi-agent collaboration described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the generative video violation risk identification device based on multi-agent collaboration provided below can be found in the limitations of the generative video violation risk identification method based on multi-agent collaboration described above. The generative video violation risk identification device based on multi-agent collaboration described below can be referred to in correspondence with the generative video violation risk identification method based on multi-agent collaboration described above, and will not be repeated here.

[0111] In one exemplary embodiment, Figure 4 This application provides a schematic diagram of the structure of a generative video violation risk identification device based on multi-agent collaboration, as shown in the embodiments of this application. Figure 4 As shown, the generative video violation risk identification device based on multi-agent collaboration includes: an acquisition module 402, a collaborative review module 404, a determination module 406, and an identification module 408, wherein: The acquisition module 402 is used to acquire the generated video to be reviewed and to conduct a preliminary review of the generated video based on preset review rules. The collaborative review module 404 is used to review the generative video through multiple agents after the generative video has passed the initial review, and obtain the review results output by each agent; the review results include a confidence score; the confidence score represents the degree of certainty with which the agent determines that the generative video contains illegal content; The determination module 406 is used to determine the overall confidence score based on the confidence scores output by each agent; The identification module 408 is used to determine the risk of violations in the generated video based on the comprehensive confidence score.

[0112] In an exemplary embodiment, the determining module 406 is specifically used to determine the maximum difference between the confidence scores output by each agent; if the maximum difference is less than or equal to a preset threshold, the confidence score output by each agent is used as the target confidence score output by each agent; if the maximum difference is greater than the preset threshold, the confidence scores output by each agent are corrected based on the collaborative debate results between the agents to obtain the corrected confidence scores output by each agent, and the corrected confidence scores output by each agent are used as the target confidence scores output by each agent; the confidence weight coefficients of each agent are obtained, and the comprehensive confidence score is determined based on the confidence weight coefficients of each agent and the target confidence scores output by each agent.

[0113] In an exemplary embodiment, the determining module 406 is specifically configured to: determine the average confidence score based on the target confidence score output by each agent; determine the confidence deviation for each agent based on the target confidence score and the average confidence score output by each agent; determine the consensus fit for each agent based on the confidence deviation; the consensus fit is used to measure the degree of matching between the agent's review conclusion and the group consensus; and determine the comprehensive confidence score based on the confidence weight coefficient of each agent, the target confidence score output by each agent, and the consensus fit for each agent.

[0114] In an exemplary embodiment, the determining module 406 is specifically configured to: determine the high-scoring agent and the low-scoring agent that generated the maximum difference if the maximum difference is greater than a preset threshold; the high-scoring agent is the agent with the highest confidence score; the low-scoring agent is the agent with the lowest confidence score; raise questions related to the violation content to the low-scoring agent through the high-scoring agent; extract evidence fragments from the generative video through the low-scoring agent; perform information calibration based on the evidence fragments through the high-scoring agent and the low-scoring agent, and obtain the latest confidence scores output by the high-scoring agent and the low-scoring agent; recalculate the maximum difference based on the latest confidence scores output by all agents; if the maximum difference is greater than the preset threshold but the preset maximum iteration round has not been reached, return to the step of determining the high-scoring agent and the low-scoring agent that generated the maximum difference; if the maximum difference is less than or equal to the preset threshold or the preset maximum iteration round has been reached, use the latest confidence score output by each agent as the corrected confidence score output by each agent.

[0115] In an exemplary embodiment, the identification module 408 is configured to: determine that the generated video has no violation risk if the overall confidence score is less than a first threshold; determine that the generated video contains content with low violation risk if the overall confidence score is greater than or equal to the first threshold and less than a second threshold, where the second threshold is greater than the first threshold; determine that the generated video contains content with medium violation risk if the overall confidence score is greater than or equal to the second threshold and less than a third threshold, where the third threshold is greater than the second threshold; and determine that the generated video contains content with high violation risk if the overall confidence score is greater than or equal to the third threshold.

[0116] In an exemplary embodiment, the intelligent agent includes a visual semantic extraction agent, a cross-modal logic verification agent, a fact-checking agent, and an emotional metaphor analysis agent; the review result also includes a review conclusion characterizing whether the generative video violates regulations; the collaborative review module 404 is used to extract visual semantic tags of each key frame in the generative video through the visual semantic extraction agent, and to identify whether there are violations in the generative video based on the visual semantic tags of each key frame; the visual semantic tags include at least event subject information, scene nature information, narrative meaning information, and violation risk tendency information; after extracting multimodal features of the generative video through the cross-modal logic verification agent, modal conflict detection is performed to obtain modal conflict detection results, and based on the modal conflict detection results, the existence of violations in the generative video is identified; the modal conflict detection results include modal conflict type information. The system extracts information on modal conflict location, intensity, and fragments. A fact-checking agent extracts structured factual information corresponding to each event from the generative video and inputs it into a pre-built knowledge base for comparison, generating a fact-checking report. Based on this report, it identifies whether violations exist in the generative video. The fact-checking report includes information on forged points, distortions, incorrect timelines, and incorrect citations corresponding to forged events. An emotional metaphor analysis agent extracts emotional features from the generative video across multiple modalities and performs intermodal emotional contrast detection, generating intermodal emotional contrast detection results. Based on these results, it identifies whether violations exist in the generative video. The intermodal emotional contrast detection results include implicit emotional risk type information, implicit emotional expression techniques information, and violation risk level information.

[0117] In one exemplary embodiment, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the generative video violation risk identification methods based on multi-agent collaboration described in the above embodiments.

[0118] In one exemplary embodiment, this application also provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of any of the generative video violation risk identification methods based on multi-agent collaboration described in the above embodiments.

[0119] In one exemplary embodiment, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the generative video violation risk identification methods based on multi-agent collaboration as described in the above embodiments.

[0120] Indicatively, such as Figure 5 As shown, Figure 5This is a schematic diagram of the internal structure of a computer device 500 provided in an embodiment of this application. The computer device 500 can be provided as a server. (Refer to...) Figure 5 The computer device 500 includes a processing component 502, which further includes one or more processors, and memory resources represented by memory 501 for storing instructions, such as application programs, that can be executed by the processing component 502. The application programs stored in memory 501 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 502 is configured to execute instructions to perform the generative video violation risk identification method based on multi-agent cooperation of any of the above embodiments.

[0121] The computer device 500 may also include a power supply component 503 configured to perform power management of the computer device 500, a wired or wireless network interface 504 configured to connect the computer device 500 to a network, and an input / output (I / O) interface 505. The computer device 500 may operate on an operating system stored in memory 501, such as Windows Server™, Mac OS X™, Unix™, Linux™, Free BSD™, or similar.

[0122] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0123] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0124] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0125] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A generative video violation risk identification method based on multi-agent cooperation, characterized in that, The method includes: Obtain the generated video to be reviewed, and conduct a preliminary review of the generated video based on preset review rules; If the generative video passes the initial review, multiple agents review the generative video to obtain review results output by each agent; the review results include a confidence score; the confidence score represents the degree of certainty with which the agent determines that the generative video contains illegal content; Based on the confidence scores output by each of the aforementioned agents, a comprehensive confidence score is determined. The risk of violation of the generative video is determined based on the comprehensive confidence score.

2. The method according to claim 1, characterized in that, The step of determining the overall confidence score based on the confidence scores output by each of the intelligent agents includes: Determine the maximum difference between the confidence scores output by each of the agents; If the maximum difference is less than or equal to a preset threshold, then the confidence score output by each agent is taken as the target confidence score output by each agent. If the maximum difference is greater than the preset threshold, the confidence score output by each agent is corrected based on the collaborative debate results between the agents to obtain the corrected confidence score output by each agent, and the corrected confidence score output by each agent is used as the target confidence score output by each agent. Obtain the confidence weight coefficient of each agent, and determine the comprehensive confidence score based on the confidence weight coefficient of each agent and the target confidence score output by each agent.

3. The method according to claim 2, characterized in that, The step of determining the comprehensive confidence score based on the confidence weight coefficients of each agent and the target confidence score output by each agent includes: The average confidence score is determined based on the target confidence score output by each of the aforementioned agents. The confidence deviation for each agent is determined based on the target confidence score output by each agent and the average confidence score. Based on the confidence deviation of each agent, the consensus fit of each agent is determined; the consensus fit is used to measure the degree of matching between the agent's review conclusion and the group consensus. The overall confidence score is determined based on the confidence weight coefficient of each agent, the target confidence score output by each agent, and the consensus fit of each agent.

4. The method according to claim 2, characterized in that, If the maximum difference is greater than the preset threshold, then the confidence score output by each agent is corrected based on the collaborative debate results among the agents, resulting in a corrected confidence score output by each agent, including: If the maximum difference is greater than the preset threshold, then the high-scoring agent and the low-scoring agent that generated the maximum difference are determined; the high-scoring agent is the agent with the highest confidence score; the low-scoring agent is the agent with the lowest confidence score. The high-scoring agent raises questions related to the violation to the low-scoring agent. The low-scoring agent extracts evidence fragments from the generative video. The high-scoring agent and the low-scoring agent perform information calibration based on the evidence fragments, and obtain the latest confidence scores output by the high-scoring agent and the low-scoring agent. The maximum difference is recalculated based on the latest confidence scores output by all agents. If the maximum difference is greater than the preset threshold but the preset maximum number of iterations has not been reached, return to the step of determining the high-scoring agent and the low-scoring agent that generated the maximum difference; If the maximum difference is less than or equal to the preset threshold or reaches the preset maximum iteration round, the latest confidence score output by each agent is taken as the corrected confidence score output by each agent.

5. The method according to claim 1, characterized in that, The step of determining the violation risk of the generative video based on the comprehensive confidence score includes: If the overall confidence score is less than the first threshold, then the generative video is determined to have no risk of violation. If the overall confidence score is greater than or equal to the first threshold and less than the second threshold, then the generative video is determined to contain content with low violation risk; the second threshold is greater than the first threshold. If the overall confidence score is greater than or equal to the second threshold and less than the third threshold, then the generative video is determined to contain content with a risk of violation; the third threshold is greater than the second threshold. If the overall confidence score is greater than or equal to the third threshold, then the generative video is determined to contain content with a high risk of violation.

6. The method according to claim 1, characterized in that, The intelligent agents include a visual semantic extraction intelligent agent, a cross-modal logic verification intelligent agent, a fact-checking intelligent agent, and an emotional metaphor analysis intelligent agent; the review results also include a review conclusion characterizing whether the generative video violates regulations; The step of reviewing the generative video through multiple intelligent agents and obtaining the review results output by each intelligent agent includes: The visual semantic extraction agent extracts visual semantic tags for each key frame in the generative video, and identifies whether there are violations in the generative video based on the visual semantic tags of each key frame; the visual semantic tags include at least event subject information, scene nature information, narrative meaning information, and violation risk tendency information. After extracting the multimodal features of the generative video through the cross-modal logic verification agent, modal conflict detection is performed to obtain the modal conflict detection results. Based on the modal conflict detection results, it is determined whether there are violations in the generative video. The modal conflict detection results include modal conflict type information, modal conflict location information, modal conflict intensity information, and modal conflict segment information. The fact-checking agent extracts structured factual information corresponding to each event from the generative video and inputs it into a pre-built knowledge base for information comparison to obtain a fact-checking report. Based on the fact-checking report, it identifies whether there are violations in the generative video. The fact-checking report includes forgery point information, distortion point information, incorrect timeline information, and incorrect reference information corresponding to the forged event. After extracting the emotional features of the generative video in multiple modalities through the emotional metaphor analysis agent, intermodal emotional contrast detection is performed to obtain intermodal emotional contrast detection results. Based on the intermodal emotional contrast detection results, it is identified whether there are violations in the generative video. The intermodal emotional contrast detection results include implicit emotional risk type information, implicit emotional expression technique information, and violation risk level information.

7. A generative video violation risk identification device based on multi-agent collaboration, characterized in that, The device includes: The acquisition module is used to acquire the generated video to be reviewed and to conduct a preliminary review of the generated video based on preset review rules. The collaborative review module is used to review the generative video through multiple agents after the generative video passes the initial review, and obtain the review results output by each agent; the review results include a confidence score; the confidence score represents the degree of certainty with which the agent determines that the generative video contains illegal content; The determination module is used to determine the overall confidence score based on the confidence scores output by each of the intelligent agents; The identification module is used to determine the violation risk of the generative video based on the comprehensive confidence score.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.