An LLM-based Video Multimodal Detection Method and System
By projecting the visual, audio and text information in the video into a unified semantic space, calculating the modal consistency scores and coordinating cultural background and perspective differences, the problems of modal conflicts and cross-cultural understanding in video content detection are solved, and high-accurate multimodal detection is achieved.
Patent Information
- Application Number
- CN202510550460.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-29
AI Technical Summary
Existing video content detection technologies are unable to effectively identify modal conflicts, understand cross-cultural non-literal expressions, and coordinate different narrative perspectives, resulting in insufficient accuracy and robustness in dealing with complex semantic expressions and cross-cultural content.
Through multi-level neural networks, visual, audio and text information are projected into a unified semantic space, consistency scores between modals are calculated, potential conflict points are located, cultural background information is extracted, non-literal expressions are identified, and differentiated processing strategies are adopted to coordinate different narrative perspectives to generate multimodal detection results.
It significantly improves detection accuracy, especially the ability to understand complex semantics and cross-cultural content, reduces the rate of misjudgment, and achieves precise identification of modal conflicts and non-literal expressions and coordination of perspective differences.
Smart Images

Figure CN120071225B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and multimodal content analysis, and more specifically, it relates to a video multimodal detection method and system based on LLM. Background Art
[0002] With the popularization of social media and video platforms, modern video content has become increasingly complex, containing multimodal information such as visual images, audio voices, and text captions. These different modal information often do not simply coincide when conveying content. Especially in videos containing advanced semantic expressions such as irony, humor, and metaphor, the modalities may exhibit the characteristics of apparent conflict but semantic coherence.
[0003] Existing video content detection technologies mainly rely on unimodal analysis or simple multimodal fusion methods. These methods usually assume that the modal information is consistent and regard the inconsistency between modalities as noise or errors. For example, traditional rule-based methods rely on predefined modal consistency rules; statistical learning-based methods use shallow models to capture the correlations between modalities; although recent deep learning methods have improved multimodal fusion performance, they still mainly focus on the consistent information between modalities rather than conflict information.
[0004] These existing technologies have the following main defects: lack of an effective mechanism for identifying modal conflicts, inability to distinguish meaningful semantic conflicts (such as irony) from real content contradictions; limited ability to understand non-literal expressions that rely on specific social and cultural backgrounds, resulting in insufficient accuracy when dealing with internationalized content; ignoring the expression differences caused by different narrative perspectives and often misjudging perspective differences as content conflicts. These defects severely limit the accuracy and robustness of video content analysis systems when dealing with complex semantic expressions and cross-cultural content.
[0005] Therefore, a technical solution that can accurately understand and detect complex semantic expressions in videos is needed, especially a method and system that can accurately identify modal conflicts, understand cross-cultural non-literal expressions, and coordinate different narrative perspectives. Summary of the Invention
[0006] The present invention provides a video multimodal detection method and system based on LLM, which solves the technical problems of ignoring modal conflicts, limited cross-cultural understanding ability, and misjudgment of perspective differences in related technologies.
[0007] The present invention provides a video multimodal detection method based on LLM, including the following steps:
[0008] Extract visual, audio, and text information from the video, and project the different modal information into a unified semantic space through a multi-level neural network to generate a unified representation vector;
[0009] Based on the unified representation vectors, calculate the semantic consistency scores between different modal information, quantify the degree of inconsistency, locate potential conflict points, and construct a temporal graph of modal consistency;
[0010] Based on the temporal graph of modal consistency and video content, extract cultural background information, and calculate the deviation value between the expressed content and cultural expectations as an irony recognition feature;
[0011] Utilize the irony recognition feature to identify non-literal expressions in the video through hierarchical feature extraction and multi-modal cross-validation;
[0012] Based on the recognition results of non-literal expressions, identify different narrative perspectives in the video content, and evaluate the impact of perspectives on the judgment of modal consistency;
[0013] According to the modal consistency scores, the recognition results of non-literal expressions, and the perspective difference evaluation, adopt differential processing strategies for different types of conflicts identified, and output the detection results including modal relationship analysis and non-literal expression understanding.
[0014] In a preferred embodiment, in the step of projecting different modal information into a unified semantic space, a modality-specific transformation function is adopted, and the transformation function includes:
[0015] A visual transformation function for processing visual features of video key frames;
[0016] An audio transformation function for processing speech content and non-verbal sound features;
[0017] A text transformation function for processing subtitle text and scene text.
[0018] In a preferred embodiment, the step of calculating the semantic consistency scores between different modal information includes:
[0019] Use weighted cosine similarity to calculate the semantic consistency between different modal representation vectors;
[0020] Dynamically adjust the consistency judgment criterion through an adaptive threshold function;
[0021] Perform spatio-temporal clustering analysis on the detected conflict points, and merge conflict points that are close in time and semantics.
[0022] In a preferred embodiment, the step of extracting cultural background information includes:
[0023] Analyze language features, visual cultural markers, social norm features, and cultural references;
[0024] Query and dynamically expand the multi-cultural irony expression knowledge base;
[0025] Extract ironic signal features with distinctive characteristics for different cultural contexts.
[0026] In a preferred embodiment, the step of identifying non-literal expressions in the video through hierarchical feature extraction and multi-modal cross-validation includes:
[0027] Construct a three-level feature extraction system and analyze step by step from surface features, context features to deep semantic features;
[0028] Improve the recognition accuracy through cross-validation of language vision, language audio and three-modal integration verification;
[0029] Utilize a pre-trained large language model for deep semantic understanding and counterfactual reasoning.
[0030] In a preferred embodiment, the step of identifying different narrative perspectives in the video content includes:
[0031] Classify the narrative perspectives in the video content into subjective and objective perspectives, analyze the narrative person, and identify the narrative stance;
[0032] Establish a mapping relationship network of expression methods under different perspectives;
[0033] Evaluate the consistency of modal content based on different perspectives and calculate the consistency score after perspective adjustment.
[0034] In a preferred embodiment, the step of adopting a differential processing strategy includes:
[0035] Retain and interpretively describe irony, metaphor and meaningful semantic conflicts;
[0036] Evaluate the severity and compare the credibility of real contradiction conflicts;
[0037] Mark the perspective and coordinately represent the perspective difference conflicts;
[0038] Utilize a pre-trained large language model to make a comprehensive judgment on complex conflict situations.
[0039] In a preferred embodiment, the step of outputting the detection results including modal relationship analysis and non-literal expression understanding includes:
[0040] Construct a multi-level structured representation, including information at the video level, paragraph level and sentence level;
[0041] Add accurate timestamp and duration information to all detection results;
[0042] Construct a semantic association network among the detection results to reflect the logical and semantic connections between the contents.
[0043] In a preferred embodiment, cross-modal semantic alignment is performed in the unified semantic space to ensure that the representation vectors of the same semantic content in different modalities have high similarity, while preserving the characteristics of surface inconsistency but semantic relevance between modalities.
[0044] In a preferred embodiment, an LLM-based video multi-modal detection system is used to implement an LLM-based video multi-modal detection method, including:
[0045] A multi-modal information processing system for extracting visual, audio, and text modality information from a video and projecting different modality information into a unified semantic space;
[0046] A consistency evaluation system for calculating the semantic consistency score between different modality information, quantifying the degree of inconsistency, and locating potential conflict points;
[0047] A cultural context understanding system for extracting the cultural background information of video content and calculating the deviation value between the expressed content and cultural expectations;
[0048] A non-literal expression recognition system for identifying non-literal expressions such as irony and metaphor in a video through hierarchical feature extraction and multi-modal cross-validation;
[0049] A perspective coordination system for identifying different narrative perspectives in video content and evaluating the impact of perspectives on modality consistency judgment;
[0050] An intelligent conflict handling and output system for adopting different handling strategies for different types of conflicts and generating detection results including modality relationship analysis and non-literal expression understanding.
[0051] The beneficial effects of the present invention are as follows:
[0052] The detection accuracy is significantly improved: In videos containing modality conflicts, the detection accuracy is improved, and the false positive rate is reduced. It is particularly suitable for content containing complex expressions such as irony and humor, solving the technical problem that traditional methods cannot accurately identify modality conflicts.
[0053] The cross-cultural understanding ability is enhanced: Through the cultural context perception mechanism and the multi-cultural irony expression knowledge base, the system can accurately understand non-literal expressions under different cultural backgrounds, improving the understanding accuracy in cross-cultural scenarios and solving the technical problem of insufficient accuracy of traditional methods when dealing with internationalized content.
[0054] Multi-perspective coordinated understanding: Through the perspective difference coordination mechanism, the system can handle the expression differences caused by different narrative perspectives, reducing the misclassification rate and solving the technical problem that traditional methods misjudge perspective differences as content conflicts.
[0055] Unified Multimodal Representation and Processing: The system realizes the unified semantic space representation and processing of visual, audio, and text information. Compared with single-modal or simple modal splicing methods, it improves the information extraction efficiency and provides a solid technical foundation for the accurate analysis of complex video content. Brief Description of the Drawings
[0056] Figure 1 is a flowchart of a video multimodal detection method based on LLM according to the present invention;
[0057] Figure 2 is a detailed flowchart of generating a unified representation vector according to the present invention;
[0058] Figure 3 is a detailed flowchart of constructing a modal consistency temporal graph according to the present invention;
[0059] Figure 4 is a detailed flowchart of calculating the deviation value between the expression content and the cultural expectation according to the present invention;
[0060] Figure 5 is a detailed flowchart of identifying non-literal expressions in the video according to the present invention;
[0061] Figure 6 is a detailed flowchart of evaluating the influence of the perspective on the modal consistency judgment according to the present invention;
[0062] Figure 7 is a detailed flowchart of outputting the detection result according to the present invention. Detailed Description of the Embodiments
[0063] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.
[0064] In at least one embodiment of the present invention, a video multimodal detection method based on LLM is disclosed, as Figures 1 to 7 shown, including the following steps:
[0065] Step 1: Extract visual, audio, and text information from the video, and project the different modal information into a unified semantic space through a multi-layer neural network to generate a unified representation vector;
[0066] Specifically, it includes the following steps:
[0067] Step 1.1: Multimodal information extraction;
[0068] Extract the original information of three modalities, namely vision, audio, and text, from the input video respectively:
[0069] Vision information extraction: Sample the video frame by frame to obtain a key frame sequence, and use a pre-trained vision model to extract visual features;
[0070] Audio information extraction: Separate the audio track from the video and extract speech content and non-verbal sound features (such as pitch, speech rate, pauses, etc.);
[0071] Text information extraction: Identify the subtitle text, scene text in the video, and the text content transcribed from speech.
[0072] Step 1.2, modality-specific preprocessing;
[0073] Preprocess the original information of each modality extracted to make it suitable for subsequent modality conversion:
[0074] Vision preprocessing: Perform operations such as normalization and size adjustment on the image, and extract semantic-level features and emotional expression features;
[0075] Audio preprocessing: Perform noise filtering, separate speech content and emotional features, such as intonation changes, intensity, etc.;
[0076] Text preprocessing: Perform operations such as word segmentation and stop word removal, and extract semantic features and syntactic structure features.
[0077] Step 1.3, unified semantic space projection;
[0078] Adopt modality-specific conversion functions to convert the preprocessed information of each modality into representation vectors in a unified semantic space:
[0079] ;
[0080] Among them, represents the original information of the th modality, represents the specific conversion function for the th modality, represents the representation vector in the unified semantic space after conversion.
[0081] In specific implementation, adopt a combined structure of a multi-layer perceptron (MLP) and an attention mechanism, which can capture the complex features and relationships within the modality.
[0082] Step 1.4, cross-modal semantic alignment;
[0083] Through contrastive learning methods, semantic alignment between representation vectors of different modalities is achieved, ensuring that representation vectors of the same semantic content in different modalities have high similarity:
[0084] ;
[0085] in, represents the alignment loss function, which is used to measure the degree of semantic alignment between vectors of different modalities; For all modal pairs Perform summation; , Respectively represent and A vector that expresses specific semantic content in each modality; Indicates Expression in modality vectors of different semantic contents; Represents a function that calculates the similarity between two vectors; represents the margin parameter, which is used to control the distance difference between positive sample pairs and negative sample pairs; Indicates taking and The larger value of ensures that the loss function is non-negative.
[0086] By minimizing this loss function, the system can learn to bring different modal vectors expressing the same semantics closer together (increasing ), while pushing the modal vectors expressing different semantics further away (reducing ), thereby achieving cross-modal semantic alignment.
[0087] This step outputs a set of multimodal representation vectors in a unified semantic space, which provides a basis for subsequent consistency evaluation and conflict detection. The innovation of this step is that it uses a semantic alignment mechanism specifically for non-literal expression features, which can retain the superficial inconsistency but semantic relevance between modalities, avoiding the problem of traditional methods that simply regard meaningful surface differences as noise.
[0088] Step 2: Based on the unified representation vector, the semantic consistency scores between different modal information are calculated to quantify the degree of inconsistency, locate potential conflict points, and construct a modal consistency time series graph;
[0089] The specific steps include:
[0090] Step 2.1, modal pair consistency calculation;
[0091] For any two modes and , calculate its semantic consistency score:
[0092] ;
[0093] Among them, represents the semantic consistency score between modality and modality . The higher the value, the more consistent the semantics of the two modalities are; represents the representation vector of modality in the unified semantic space, which contains the semantic information of this modality; represents the representation vector of modality in the unified semantic space, which contains the semantic information of this modality; represents a function used to calculate the similarity between two modality vectors.
[0094] Similarity function Adopts weighted cosine similarity, considering the consistency in both the semantic content and sentiment expression dimensions:
[0095] ;
[0096] Among them, , respectively represent the semantic content vectors of modality and modality ; , respectively represent the sentiment expression vectors of modality and modality ; represents the cosine similarity function, with a range of [-1, 1]. The larger the value, the more similar the vector directions are; represents the weight coefficient of the semantic content similarity, with a value range of [0, 1], used to balance the importance of semantic content and sentiment expression in the consistency evaluation, and can be adjusted according to specific application scenarios.
[0097] Step 2.2, Conflict point detection and localization;
[0098] Based on the modality pair consistency score, detect and localize potential conflict points:
[0099] Calculate the consistency threshold:
[0100] ;
[0101] Among them, represents the dynamic threshold for determining modality conflict. When the consistency score is lower than this threshold, it is determined as a conflict point; represents the adaptive threshold function, which can dynamically adjust the threshold value according to historical data and the current context; represents the distribution of historical consistency scores, including the statistical information of the consistency scores of each modality pair in the previous time period; Represents the context feature vector of the current video segment, including information such as scene type, theme content, and expression style.
[0102] Identify conflict points: When this occurs, it is determined as a potential conflict point, and relevant modalities, timestamps, and conflict intensities are recorded.
[0103] Cluster analysis: Conduct spatio-temporal cluster analysis on the detected conflict points, merge conflict points that are similar in time and semantics, and reduce redundancy.
[0104] Step 2.3, preliminary classification of conflict types;
[0105] Conduct preliminary classification on the detected conflict points to distinguish common conflict types:
[0106] Irony type: The literal meaning is opposite to the potential semantics, but the emotional expression is consistent;
[0107] Hyperbole type: There is a degree difference between the literal meaning and the actual situation, but the direction is the same;
[0108] Metaphor type: Establish a mapping relationship between concepts in different domains;
[0109] True contradiction type: There are logical or factual conflicts in the information itself;
[0110] Through a method combining a pre-trained text classification model and rules, conduct preliminary classification of the conflicts.
[0111] Step 2.4, construction of a modality consistency time series graph;
[0112] Construct a modality consistency time series graph for the entire video to track the change rules of conflicts:
[0113] Use time as the horizontal axis, modality pairs as the vertical axis, and consistency scores as values to construct a time series heat map;
[0114] Extract key change points and mark the time nodes with significant changes in consistency;
[0115] Analyze the time series rules of conflict patterns, such as periodicity, gradual change, or mutation;
[0116] The output of this step includes the modality consistency evaluation result, conflict point location information, preliminary conflict type classification result, and modality consistency time series graph. The innovation of this step lies in the adoption of an adaptive threshold and a multi-dimensional similarity calculation method, which can flexibly adjust the conflict determination criteria according to the characteristics of different video contents. At the same time, by analyzing the time series graph, it captures the conflict evolution pattern, providing a basis for subsequent cultural context understanding and non-literal expression recognition.
[0117] Step 3: Based on the modality-consistent temporal graph and video content, extract cultural background information, and calculate the deviation value between the expressed content and cultural expectations as an irony recognition feature;
[0118] Specifically, it includes the following steps:
[0119] Step 3.1: Cultural context recognition;
[0120] Based on the multi-modal features of the video, identify the possible cultural context background of the video content:
[0121] Language feature analysis: Infer the cultural background through language recognition, dialect features, specific vocabulary, and expressions;
[0122] Visual cultural marker recognition: Extract culturally specific symbols, costumes, buildings, scenes, etc. from the visual content;
[0123] Social norm extraction: Analyze social norm features such as character interaction patterns, social etiquette, and expression taboos;
[0124] Cultural reference retrieval: Identify popular cultural elements, memes, historical references, etc. in a specific cultural background.
[0125] Through a weighted voting mechanism, integrate the above features to obtain the possible cultural context classification of the video, such as East Asian culture, Western culture, Middle Eastern culture, etc.
[0126] Step 3.2: Query and expansion of the culturally specific irony expression knowledge base;
[0127] For the identified cultural context, query and dynamically expand the culturally specific irony expression knowledge base:
[0128] Basic knowledge base query: Retrieve irony patterns and clues related to the identified cultural context from the pre-built multi-cultural irony expression knowledge base;
[0129] Context-related expansion: Based on the specific theme and content of the video, use a pre-trained language model to expand relevant irony expression patterns;
[0130] Expression pattern extraction: Extract common irony signals in this cultural background from the knowledge base, such as specific rhetorical devices, intonation changes, and facial expression combinations.
[0131] Step 3.3: Calculation of semantic expectation deviation;
[0132] Based on the cultural context and expressed content, calculate the deviation value between the literal meaning of the expression and the expected meaning in this cultural context:
[0133] ;
[0134] Among them, Represents the expression content to be analyzed, Represents the expression of the literal meaning vector, Represents in the cultural context the expected meaning vector of the expression, Represents the deviation calculation function, Represents the final irony intensity score.
[0135] Deviation function Adopts a weighted semantic sentiment difference measure:
[0136] ;
[0137] Among them, Represents the literal meaning vector and the expected meaning vector the semantic dimension difference measure between them, used to calculate the distance between the two vectors in the semantic space; Represents the literal meaning vector and the expected meaning vector the sentiment dimension difference measure between them, used to calculate the difference degree in the emotional expression of the two vectors; is the weight parameter, and the value range is , used to adjust the relative importance of semantic difference and sentiment difference in the overall deviation calculation. A larger value indicates more emphasis on semantic difference, and a smaller value indicates more emphasis on sentiment difference.
[0138] Step 3.4, Multi-cultural irony signal feature extraction;
[0139] For different cultural contexts, extract distinctive irony signal features:
[0140] East Asian cultural irony features: such as euphemistic expressions, modest statements, self-deprecation, etc.;
[0141] Western cultural irony features: such as exaggerated contrast, direct statements inconsistent with expressions, etc.;
[0142] Cross-cultural common irony features: such as intonation changes, pauses, rhetorical questions, etc.
[0143] This step outputs the cultural context recognition result, the irony intensity score of the expression content, and the culture-specific irony signal feature set. The innovation of this step lies in breaking through the limitation of traditional irony detection that ignores cultural differences. By introducing a cultural context perception mechanism and a dynamic knowledge base, it realizes the accurate recognition of non-literal expressions in different cultural backgrounds, and is especially suitable for the analysis of cross-cultural communication videos.
[0144] Step 4: Using irony recognition features, identify non-literal expressions in the video through hierarchical feature extraction and multimodal cross-validation;
[0145] Specifically, it includes the following steps:
[0146] Step 4.1: Hierarchical non-literal expression feature extraction;
[0147] Construct a three-level feature extraction system to analyze non-literal expression features step by step from the surface to the deep layer:
[0148] Surface features (primary features):
[0149] Linguistic markers: Extract obvious markers in the language, such as rhetorical questions, hyperbolic words, obvious contrasts, etc.;
[0150] Common pattern matching: Perform pattern matching based on a predefined expression pattern library;
[0151] Emotional markers: Identify the inconsistency between emotional words and the context.
[0152] Context features (secondary features):
[0153] Internal consistency of discourse: Analyze the semantic consistency within the expressed content;
[0154] Context coherence: Evaluate the semantic coherence degree of the expression with the context before and after;
[0155] Cross-modal consistency: Evaluate the consistency between language content and other modal information such as vision and audio.
[0156] Deep semantic features (tertiary features):
[0157] Pragmatic intention analysis: Infer the possible true intention of the speaker;
[0158] Social norm and expectation deviation: Analyze the deviation between the expressed content and social norms and expectations;
[0159] Discourse strategy recognition: Analyze the possible expression strategies and rhetorical devices adopted by the speaker.
[0160] Step 4.2: Multimodal cross-validation;
[0161] Improve the accuracy of non-literal expression recognition through cross-modal information cross-validation:
[0162] Language-vision cross-validation:
[0163] Consistency analysis of facial expressions and language content;
[0164] Coordination evaluation of body language and language content;
[0165] Analysis of the relevance between scene context and language content.
[0166] Cross-validation of language audio:
[0167] Analysis of the correlation between intonation changes and language content;
[0168] Correspondence between pauses, accents and semantic focuses;
[0169] Evaluation of the consistency between emotional expressions and language content.
[0170] Tri-modal integration verification:
[0171] Construct a weighted voting mechanism to integrate non-literal expression clues from three modalities;
[0172] Use Bayesian network to model the conditional dependence relationship between modalities;
[0173] Generate a comprehensive confidence score to reflect the reliability of non-literal expression recognition.
[0174] Step 4.3, Semantic understanding based on pre-trained language models;
[0175] Utilize the deep semantic understanding ability of large pre-trained language models to identify complex non-literal expressions:
[0176] Semantic representation extraction: Use pre-trained language models to extract deep semantic representations of the expressed content;
[0177] Literal-implicit meaning comparison: Calculate the difference between the literal meaning and the possible implicit meaning of the expression;
[0178] Counterfactual reasoning: Analyze the potential pragmatic intentions of the expression through counterfactual conditional reasoning;
[0179] Knowledge-enhanced understanding: Combine external knowledge bases to assist in understanding expressions specific to a particular domain or culture.
[0180] The model input format is:
[0181] [Text]{Text content};
[0182] [Visual description]{Visual feature description};
[0183] [Audio feature]{Audio feature description};
[0184] [Cultural context]{Cultural background information};
[0185] [Task] Identify possible non-literal expressions and their true meanings in the above content.
[0186] Step 4.4, Non-literal expression classification and confidence evaluation;
[0187] Fine - classify the identified non - literal expressions and evaluate the confidence of the recognition results:
[0188] Expression type classification: Classify the identified non - literal expressions into specific types such as irony, metaphor, hyperbole, pun, etc.;
[0189] Expression intensity scoring: Evaluate the intensity of non - literal expressions, from subtle implications to obvious expressions;
[0190] Confidence calculation: Calculate the confidence of the recognition results based on a comprehensive evaluation of multi - feature and multi - modal evidence;
[0191] Uncertainty handling: Identify situations with insufficient confidence, trigger further analysis or request manual confirmation;
[0192] The output of this step is the recognition results of non - literal expressions in the video content, including information such as expression type, location, intensity, and confidence. The innovation of this step lies in the adoption of hierarchical feature extraction and multi - modal cross - verification methods, combined with the deep semantic understanding ability of pre - trained language models, to achieve accurate recognition of complex non - literal expressions, especially for the precise understanding of cross - modal expressions.
[0193] Step 5: Based on the recognition results of non - literal expressions, identify different narrative perspectives in the video content and evaluate the impact of perspectives on modal consistency judgment;
[0194] Specifically, it includes the following steps:
[0195] Step 5.1: Narrative perspective recognition and classification;
[0196] Identify and classify the narrative perspectives in the video content:
[0197] Subjective - objective perspective classification:
[0198] Objective description: Content expression oriented towards neutral facts;
[0199] Subjective comment: Content expression containing personal opinions, emotions, and evaluations;
[0200] Mixed perspective: A composite expression in which objective description and subjective comment are intertwined;
[0201] Narrative person analysis:
[0202] First - person perspective: Content expression with the speaker as the reference point;
[0203] Third - person perspective: Content expression with the speaker as a bystander;
[0204] Second - person perspective: Content expression directly addressed to the audience;
[0205] Narrative stance recognition:
[0206] Supporting stance: Content expressing agreement and approval;
[0207] Opposing stance: Content expressing negation and criticism;
[0208] Neutral stance: Content that does not show a clear attitude;
[0209] Perspective recognition is carried out through a combination method of language markers, pronoun usage, emotional word distribution and other features, combined with rules and classification models.
[0210] Step 5.2, Construction of perspective-dependent expression mapping network;
[0211] Establish a mapping relationship network of expression methods under different perspectives:
[0212] Perspective conversion pattern extraction: Analyze the expression difference patterns of the same content under different perspectives;
[0213] Expression template learning: Learn expression templates and rules under specific perspectives from a large number of samples;
[0214] Perspective mapping relationship modeling: Construct a semantic mapping function from perspective A to perspective B;
[0215] The mapping network adopts a neural network structure enhanced by an attention mechanism, which can capture subtle perspective-dependent expression differences.
[0216] Step 5.3, Perspective consistency evaluation;
[0217] Evaluate the consistency of modal content based on different perspectives:
[0218] ;
[0219] Among them, The modal consistency score after considering perspective factors; 、 respectively represent the content representation vectors of modality and modality ; The basic similarity function; The perspective conversion function; and respectively represent the narrative perspective feature vectors of modality and modality ;
[0220] The perspective conversion function is defined as:
[0221] ;
[0222] Among them, Viewpoint difference metric function, with a value range of [0, 1]; Viewpoint influence adjustment parameter, which controls the degree of influence of viewpoint difference on consistency evaluation, with a value range of [0, 1];
[0223] When the two viewpoints are completely consistent, , so , no adjustment is made to the basic similarity;
[0224] When the viewpoint difference is large, is close to 1, decreases but is not zero, indicating that even though the viewpoints are different, the content may still be essentially consistent.
[0225] Step 5.4, generation of viewpoint - coordinated content representation;
[0226] Based on the identified viewpoint differences and mapping relationships, generate a coordinated and consistent content representation:
[0227] Viewpoint normalization: Convert the content from different viewpoints to a standard viewpoint for comparison;
[0228] Viewpoint weighted fusion: According to viewpoint reliability and relevance, perform weighted fusion on the content from different viewpoints;
[0229] Retention of viewpoint opposition representation: For intentional viewpoint opposition (such as debate content), retain the opposition representation and mark the relationship.
[0230] The output of this step is the viewpoint analysis result of the video content, the viewpoint mapping relationship, and the content representation after viewpoint coordination. The innovation of this step lies in solving the problem of misjudging viewpoint differences as content conflicts in traditional methods by identifying and coordinating the expression differences caused by different narrative viewpoints, and improving the system's understanding ability of complex content such as debate videos and multi - angle reports.
[0231] Step 6, according to the modal consistency score, non - literal expression recognition result, and viewpoint difference evaluation, adopt different processing strategies for different types of identified conflicts, and output the detection results including modal relationship analysis and non - literal expression understanding;
[0232] Specifically, it includes the following steps:
[0233] Step 6.1, differential processing of conflict types;
[0234] Adopt different processing strategies according to the conflict type:
[0235] Processing of meaningful semantic conflicts (such as irony, metaphor, etc.):
[0236] Retain conflict features: Regard such conflicts as important semantic expression features and retain them;
[0237] Explicitly mark non-literal expressions: Mark the type, strength, and credibility of non-literal expressions;
[0238] Generate explanatory descriptions: Provide possible explanations and true meanings of non-literal expressions;
[0239] Handling of real contradiction conflicts:
[0240] Contradiction severity assessment: Evaluate the impact of contradictions on the overall content understanding;
[0241] Credibility comparison: Compare the credibility of both sides of the conflict to determine the priority;
[0242] Consistency correction suggestions: Provide possible consistency correction suggestions for serious contradictions;
[0243] Handling of conflicts due to perspective differences:
[0244] Perspective marking: Clearly mark the expression inconsistencies caused by perspective differences;
[0245] Perspective coordination representation: Generate a unified representation after perspective coordination;
[0246] Multi-perspective preservation: Retain multi-perspective representations when necessary and mark the perspective relationships;
[0247] Step 6.2, Comprehensive judgment based on LLM;
[0248] Utilize the reasoning ability of pre-trained large language models to make comprehensive judgments on complex conflict situations:
[0249] Evidence integration: Integrate all modal evidence, conflict analysis results, and cultural context information into a structured input;
[0250] Key question generation: Automatically generate clarifying questions for uncertain conflict points;
[0251] Progressive reasoning: Adopt the Chain-of-Thought method to obtain comprehensive judgments through multi-step reasoning;
[0252] Multi-hypothesis generation and verification: Generate multiple possible explanatory hypotheses and verify them one by one in uncertain situations;
[0253] The model input format is:
[0254] [Modal conflict information]{Conflict description, type, location, strength, etc.};
[0255] [Cultural context]{Relevant cultural background information};
[0256] [Perspective information]{Relevant perspective difference information};
[0257] [Non-literal expression analysis] {Identified non-literal expression features};
[0258] [Task] Comprehensively judge the nature, cause, and most likely explanation of the above conflicts.
[0259] Step 6.3, Structured representation of detection results;
[0260] Organize the detection results into a structured representation:
[0261] Multi-level structured representation:
[0262] Video level: Macro information such as the overall content theme, style, cultural background, etc.;
[0263] Paragraph level: Meso information such as segmented information, paragraph themes, relationships between paragraphs, etc.;
[0264] Sentence / expression level: Micro information such as specific conflict points, non-literal expressions, etc.;
[0265] Temporal index association: Add accurate timestamp and duration information to all detection results for quick positioning;
[0266] Semantic association network: Construct a semantic association network between detection results to reflect the logical and semantic connections between contents;
[0267] Step 6.4, Result visualization and query interface;
[0268] Provide a visual representation of detection results and a flexible query interface:
[0269] Temporal visualization: Display the distribution of various detection results in the form of a timeline;
[0270] Conflict heat map: Intuitively display the conflict distribution and intensity through a heat map;
[0271] Semantic network graph: Visualize the semantic association network between contents;
[0272] Multi-modal parallel display: Simultaneously display visual, audio, and text contents at relevant moments;
[0273] Natural language query interface: Support querying and filtering of detection results based on natural language;
[0274] This step outputs the final video content detection results, including a structured representation of multi-dimensional information such as conflict analysis, non-literal expression understanding, and perspective coordination. The innovation of this step lies in adopting a differential processing strategy based on conflict types, combined with the reasoning ability of large language models, to achieve accurate understanding and processing of complex semantic expressions, while providing rich structured representations and visualization interfaces for subsequent content analysis and applications.
[0275] Application example of this embodiment
[0276] To verify the actual effect of this embodiment, the following takes the cross-cultural video content review system of social media as an example to show the implementation process and technical effect of this method in actual applications.
[0277] This example is applied to the video content review system of an international social media platform, which needs to process short video content with multi-language and multi-cultural backgrounds uploaded by users from all over the world. These video contents often contain complex expressions such as irony, humor, and culture-specific memes, resulting in frequent misjudgments by traditional review systems, which can neither effectively identify real harmful content nor often mislabel content containing legal non-literal expressions. The platform needs a more accurate multi-modal understanding system that can distinguish meaningful modal conflicts (such as ironic expressions) from real content contradictions.
[0278] Implementation process example:
[0279] Implementation of unified representation of multi-modal information:
[0280] The platform extracts three types of modal information from the uploaded videos: visual, audio, and text. The following table shows the process of multi-modal information extraction and unified representation conversion for a video clip containing ironic expressions:
[0281] Table 1: Example of multi-modal information extraction and conversion for an ironic video;
[0282]
[0283] In this example, the system converts different modal information into representation vectors in a unified semantic space through an optimized conversion function, which preserves the modal-specific semantics and emotional features and lays a foundation for subsequent conflict recognition.
[0284] Implementation of modal consistency evaluation and conflict localization:
[0285] Based on the representation vectors in the unified semantic space, the system calculates the consistency scores between different modalities and generates a modal consistency time series graph. The following table shows the modal consistency evaluation results of the above video clip:
[0286] Table 2: Modal consistency evaluation results for an ironic video;
[0287]
[0288] The system detected a significant inconsistency between the visual modality and the text / audio modality, but a high degree of consistency was maintained between the text and the audio. Through an adaptive threshold (set to 0.45 in this example), the system determined these points as conflict points and preliminarily classified them as irony-type conflicts.
[0289] Implementation of multi-level non-literal expression recognition:
[0290] The system performs multi-level feature extraction and multi-modal cross-validation on the video content to recognize complex non-literal expressions:
[0291] Table 3: Examples of multi-level non-literal expression recognition;
[0292]
[0293] Through hierarchical feature extraction and modal cross-validation, the system successfully recognized these non-literal expressions and provided an evaluation of the expression type and confidence level.
[0294] Implementation of perspective difference coordination and analysis:
[0295] The system identifies different narrative perspectives in the video content and evaluates the impact of perspectives on the judgment of modal consistency:
[0296] Table 4: Examples of perspective difference coordination and analysis;
[0297]
[0298] Through the perspective difference coordination mechanism, the system can adjust the evaluation results of modal consistency according to different narrative perspectives, avoiding misjudging perspective differences as content conflicts. Especially for the mixed content of objective descriptions and subjective comments, the consistency evaluation after perspective coordination is more accurate.
[0299] Implementation of intelligent conflict handling and result output:
[0300] For different types of conflicts identified, the system adopts differentiated handling strategies and generates structured detection results:
[0301] Table 5: Examples of intelligent conflict handling strategies;
[0302]
[0303] Through the intelligent conflict handling strategy, the system distinguishes meaningful semantic expressions (such as irony and exaggeration) from real content contradictions, effectively improving the accuracy and intelligence level of content review.
[0304] Verification of technical effects:
[0305] The detection accuracy is significantly improved:
[0306] This embodiment significantly improves the detection accuracy of complex video content in practical applications, especially for content containing modal conflicts and non-literal expressions. The following table shows the comparison of detection effects on different types of videos:
[0307] Table 6: Comparison of detection accuracies of different methods on videos with modal conflicts;
[0308]
[0309] As shown in Table 6, the accuracy of this embodiment has been increased by an average of 37.3% when processing various complex video contents, and it has been most significantly improved, reaching 41.1%, especially for videos containing culture-specific expressions.
[0310] Enhanced cross-cultural understanding ability:
[0311] This embodiment significantly enhances the system's understanding ability of non-literal expressions in different cultural backgrounds through cultural context recognition and a multi-cultural irony expression knowledge base. The following table shows the cross-cultural understanding accuracies on videos with different cultural backgrounds:
[0312] Table 7: Understanding accuracies of different methods on cross-cultural video contents;
[0313]
[0314] As shown in Table 7, this embodiment is significantly superior to traditional methods and conventional LLM methods in cross-cultural video content understanding, with an average accuracy improvement of 42.4%. Especially when processing cross-cultural mixed contents and videos with non-Western cultural backgrounds, the improvement is more significant, fully demonstrating the cultural adaptability and cross-cultural understanding ability of this method.
[0315] The above practical application examples and technical effect verifications prove that this embodiment has significant advantages in solving complex problems in video multi-modal detection, especially in dealing with modal conflicts, non-literal expressions, and cross-cultural contents, and can effectively improve the detection accuracy and enhance the system's cross-cultural understanding ability.
[0316] The embodiments of the present invention have been described above, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.
Claims
1. A video multi-modal detection method based on LLM, characterized in that, It includes the following steps: Extract visual, audio, and text information from the video, project different modality information into a unified semantic space through a multi-level neural network, and generate a unified representation vector; Based on the unified representation vector, calculate the semantic consistency scores between different modality information, quantify the degree of inconsistency, locate potential conflict points, and construct a modality consistency temporal graph; Among them, the step of calculating the semantic consistency scores between different modality information includes: Use weighted cosine similarity to calculate the semantic consistency between different modality representation vectors; Dynamically adjust the consistency judgment criterion through an adaptive threshold function; Conduct spatio-temporal clustering analysis on the detected conflict points, and merge conflict points that are close in time and semantics; Extract cultural background information based on the modality consistency temporal graph and video content, and calculate the deviation value between the expressed content and cultural expectations as an irony recognition feature; Utilize the irony recognition feature to identify non-literal expressions in the video through hierarchical feature extraction and multi-modal cross-validation, including: constructing a three-level feature extraction system to analyze step by step from surface features, context features to deep semantic features; Improve the recognition accuracy through language-vision, language-audio cross-validation and tri-modal integration verification; utilize a pre-trained large language model for deep semantic understanding and counterfactual reasoning; Based on the non-literal expression recognition result, identify different narrative perspectives in the video content, and evaluate the impact of the perspective on the modality consistency judgment; According to the modality consistency score, non-literal expression recognition result and perspective difference evaluation, adopt differential processing strategies for different types of identified conflicts, and output detection results including modality relationship analysis and non-literal expression understanding; Perform cross-modal semantic alignment in the unified semantic space to ensure that the representation vectors of the same semantic content in different modalities have high similarity, while retaining the characteristics of surface inconsistency but semantic relevance between modalities.
2. The video multi-modal detection method based on LLM according to claim 1, wherein, In the step of projecting different modality information into a unified semantic space, a modality-specific transformation function is adopted, and the transformation function includes: A visual transformation function for processing the visual features of video key frames; An audio transformation function for processing speech content and non-verbal sound features; A text transformation function for processing subtitle text and scene text.
3. A video multi-modal detection method based on LLM according to claim 1, characterized in that, The step of extracting cultural background information includes: Analyze language features, visual cultural markers, social norm features, and cultural references; Query and dynamically expand the multi-cultural irony expression knowledge base; Extract irony signal features with unique characteristics for different cultural contexts.
4. A video multi-modal detection method based on LLM according to claim 1, characterized in that The step of identifying different narrative perspectives in the video content includes: Conduct subjective and objective perspective classification, narrative person analysis, and narrative stance recognition on the narrative perspectives in the video content; Establish a mapping relationship network for expression methods under different perspectives; Evaluate the modality content consistency based on different perspectives, and calculate the consistency score after perspective adjustment.
5. A video multi-modal detection method based on LLM according to claim 1, characterized in that, The step of adopting differential processing strategies includes: Retain and descriptively explain irony, metaphor, and meaningful semantic conflicts; Conduct severity assessment and credibility comparison on real contradiction conflicts; Conduct perspective marking and coordinated representation on perspective difference conflicts; Utilize a pre-trained large language model to comprehensively judge complex conflict situations.
6. The video multi-modal detection method based on LLM according to claim 1, wherein The steps of the output including the detection results of modal relation parsing and non-literal expression understanding are as follows: Construct a multi-level structured representation, including information at the video level, paragraph level, and sentence level; Add accurate timestamp and duration information to all detection results; Construct a semantic association network among the detection results to reflect the logical and semantic connections between the contents.
7. A video multi-modal detection system based on LLM, which is used to implement a video multi-modal detection method based on LLM according to any one of claims 1-6, characterized in that, Include: A multi-modal information processing system for extracting three types of modal information, namely visual, audio, and text, from the video and projecting different modal information into a unified semantic space; A consistency evaluation system for calculating the semantic consistency score between different modal information, quantifying the degree of inconsistency, and locating potential conflict points; A cultural context understanding system for extracting the cultural background information of the video content and calculating the deviation value between the expressed content and the cultural expectation; A non-literal expression recognition system for identifying non-literal expressions such as irony and metaphor in the video through hierarchical feature extraction and multi-modal cross-validation; A perspective coordination system for identifying different narrative perspectives in the video content and evaluating the impact of the perspective on the modal consistency judgment; An intelligent conflict handling and output system for adopting differentiated handling strategies for different types of conflicts and generating detection results including modal relation parsing and non-literal expression understanding.
Citation Information
Patent Citations
Video target behavior anomaly detection method and system based on multi-modal feature fusion
CN114782882A
Brain-like cross-modal information interaction fusion method for video irony analysis
CN116071629A