A multi-modal risk event anomaly detection method and system based on fast-slow dual system and size model cooperation
Patent Information
- Application Number
- CN202610785107.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-28
AI Technical Summary
[0006]此外,现有基准数据集通常仅提供视频-文本输入与二值标签,并不显式暴露可验证声明、支持或反驳的外部证据以及中间推理信号,导致即便引入具备强推理能力的大模型,也缺乏从封闭世界检测迈向证据感知验证所必需的监督信息
[0019] Compared with existing technologies, the beneficial effects of this invention are as follows: First, it requires no training or fine-tuning of large models, making it a plug-and-play inference framework with low deployment costs and easy scalability; Second, it reconstructs detection into evidence-aware verification, introducing explicit statement decomposition and evidence grounding, which can effectively identify semantic-level risk events that are "perceptually consistent but factually incorrect," significantly improving detection accuracy; Third, through dynamic routing of semantic distillation and attentional cognitive signal injection, it efficiently integrates high-level inference signals from large models while retaining fine-grained perceptual cues, resulting in a small number of trainable parameters and high inference efficiency; Fourth, the framework outputs a traceable inference path from overall anomaly perception to statement-level evidence verification, significantly improving the interpretability of the judgment and providing reliable technical support for high-risk scenarios such as public opinion security.
Smart Images

Figure CN122654898A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and multimodal content security technology, specifically relating to a multimodal risk event anomaly detection method and system based on fast and slow dual systems and large and small model collaboration. Background Technology
[0002] Anomaly detection of risk events is a crucial task in areas such as public opinion security, emergency response, and social governance. Its goal is to automatically and accurately identify abnormal events that pose a risk to public perception and social order from massive amounts of multimodal content (such as short videos with audio, subtitles, and visuals). With the explosive growth of short video platforms, risk events appearing in the form of fabricated narratives, misleading situational splicing, and selective editing are spreading at an unprecedented speed. Because these events simultaneously intertwine visual images, audio narration, on-screen text, and chronological dynamics, they present a significant challenge to automated detection.
[0003] Most existing methods model this task as a perception-driven classification problem, mainly relying on deep neural networks to identify statistical inconsistencies or forgery traces (such as deepfakes, image-text mismatches, and fake audio) across visual, textual, and acoustic modalities. These methods rely on a fundamental implicit assumption that risky events can always be reliably identified through perceptual forgery traces or cross-modal inconsistencies.
[0004] However, this assumption is increasingly failing in real-world scenarios. Contemporary risky content often presents itself as "perceptually consistent but factually incorrect": the visuals are well-produced, the narration is fluent and natural, and there is a high degree of consistency across modalities, while its risk lies precisely hidden in the underlying factual statements, rather than in observable perceptual cues. In this situation, signals at the perceptual level are almost indistinguishable, making detection models that rely solely on perception structurally unable to differentiate between factual reporting and factually incorrect narratives that are difficult to distinguish at the audiovisual level.
[0005] This failure mode reveals a paradigmatic gap between detection and verification: detection-oriented methods operate under the closed-world assumption, learning only the correlation between input patterns and labels without accessing external reality; in contrast, real-world fact-checking requires explicit reasoning on verifiable statements and grounding them in authoritative external evidence.
[0006] Furthermore, existing benchmark datasets typically only provide video-text input and binary labels, without explicitly exposing external evidence supporting or refuting verifiable claims, or intermediate inference signals. This results in a lack of supervised information necessary to move from closed-world detection to evidence-aware verification, even when large models with strong inference capabilities are introduced. Directly performing end-to-end inference with large models faces problems such as high deployment costs, large inference latency, and difficulty in scaling; while simply performing knowledge distillation will lose the fine-grained perceptual cues crucial for short video risk detection. Therefore, how to efficiently introduce the high-level inference capabilities of large models into lightweight detection without fine-tuning them, while taking into account both perceptual and factual dimensions, has become an urgent technical problem to be solved. Summary of the Invention
[0007] To address the aforementioned problems, this invention proposes a multimodal risk event anomaly detection method and system based on the collaboration of a fast and slow dual system and a large-scale model. This invention reconstructs risk event anomaly detection from perception-driven classification to evidence-driven verification: the model no longer judges solely based on the inherent perception patterns of multimodal inputs, but instead jointly infers from statements extracted from content and evidence obtained through external retrieval.
[0008] Formally, traditional perception-driven detection optimizes on a given training set by maximizing the label probability: In the formula, θ represents the parameters of the trainable multimodal scoring function; N represents the number of training samples; This represents a trainable multimodal scoring function; This represents the true label of the i-th sample. ; , Let represent the video input and text input of the i-th sample, respectively.
[0009] This invention extends it to an evidence-aware verification objective that combines content, claims, and evidence for reasoning: In the formula, θ, N, , Same meaning as above; This represents the multimodal content of the i-th sample. ,in For visual frames, For audio features, Titles provided to users To capture optical character recognition text embedded in the image; This represents the set of verifiable claims extracted from the content. ; This refers to a collection of authoritative evidence retrieved from external sources. K represents the number of claims. The detector itself only operates on the raw short video input, and the claims and evidence are introduced as intermediate verification traces that can be generated online or as supervision and evaluation signals.
[0010] Based on the above reconstruction, this invention constructs a fast-slow dual-system framework inspired by the dual-process theory of cognitive science. This framework comprises five steps: data acquisition, fast system overall anomaly perception, slow system evidence reasoning, large-small model collaborative fusion, and risk assessment. The fast system simulates human "fast thinking," rapidly perceiving and summarizing multimodal inputs to generate fast cognitive signals. The slow system simulates human "slow thinking," decomposing content into verifiable atomic statements, performing multi-source evidence retrieval and consistency analysis, and generating slow cognitive signals. Furthermore, through dual encoder encoding, dynamic routing semantic distillation, and attention-based cognitive signal injection, the fast and slow cognitive signals are used as auxiliary semantic modalities to inject into a compact multimodal perception representation, achieving lightweight large-small model collaboration. The final detection result is output by a lightweight classifier.
[0011] Specifically, this invention provides a multimodal risk event anomaly detection method based on the collaboration of a fast and slow dual system and a large-scale model, comprising the following steps: Step 1, Data Acquisition: Acquire the multimodal input data to be detected, which includes video frames, audio waveforms, title text, optical character recognition text embedded in the video frame, and release time; Step 2, Fast System Overall Anomaly Perception: Using a large visual language model and guided by prompt templates, anomaly cues are detected in the multimodal input data in three dimensions: visual, text, and cross-modal. Then, the language model is used to reflect on and summarize the detected anomaly cues to obtain a compact fast cognitive signal. Step 3, Slow System Evidence Reasoning: The multimodal input data is decomposed into several independently verifiable atomic statements using a visual language big model and a global summary is generated. External evidence is obtained for each atomic statement through a multi-source evidence retrieval operator. Then, the language model is used to perform a consistency analysis of the statement and the evidence to obtain a slow cognitive signal composed of the judgment conclusion and the basis. Step 4, large and small model collaborative fusion: Multimodal encoding is performed on video, audio, text and the fast cognitive signal and slow cognitive signal by dual encoders. Dynamic routing semantic distillation is performed on the features of each modality to obtain semantic capsules of uniform length. Then, attention-based cognitive signal injection and hierarchical cross-modal fusion are performed, and adaptive weighted fusion is performed on the representations of each modality to obtain the final fused representation. Step 5, Risk Assessment: Input the final fused representation into a lightweight classifier and output the detection result of whether the multimodal input data to be detected is an abnormal risk event.
[0012] As a further aspect of the present invention, step two includes two nodes: anomaly detection and reflection summarization. Taking video frames, title text, optical character recognition text, and publication time as input, the visual language model, guided by anomaly detection prompt template, extracts anomaly cues in three dimensions: visual, text, and cross-modal. Then, guided by reflection summarization prompt template, the language model compresses the above-mentioned anomaly cues into compact summaries while preserving the original semantics. The three together constitute the fast cognitive signal.
[0013] As a further aspect of the present invention, step three sequentially includes three nodes: atomic statement decomposition, multi-source evidence retrieval, and consistency analysis. Guided by the atomic statement extraction prompt template, the visual language model decomposes the content to be detected into several independently verifiable atomic statements, each consisting of a standardized statement, a statement semantic type, and suggested verification methods, and generates a one-sentence global summary. For each atomic statement, the multi-source evidence retrieval operator directly uses its standardized statement as the retrieval query, returning the most relevant evidence fragments from multiple sources including public networks, authorized news, and structured knowledge. Then, guided by the verification prompt template, the language model performs a consistency analysis of the statements and evidence, integrating time, entity, and causal logic, outputting a judgment conclusion with values of support, contradiction, or uncertainty, along with its supporting explanation, and aggregating the verification results of all statements into the slow cognitive signal.
[0014] As a further aspect of the present invention, when performing multimodal encoding on each modality in step four, modal features are extracted for video, audio, text, and the fast cognitive signal and slow cognitive signal respectively using a dual encoder strategy that combines a semantic encoder and a perceptual encoder. The fast cognitive signal and slow cognitive signal are encoded using the same backbone encoder as the text, and each encoder used is a frozen pre-trained encoder that does not participate in training or fine-tuning.
[0015] As a further aspect of the present invention, the dynamic routing semantic distillation in step four maps each modality feature into a semantic capsule of uniform length. It includes two stages: main capsule generation and iterative consistent routing. In the main capsule generation stage, the input features are linearly projected and then projected through a Transformer layer to form a main capsule sequence carrying local temporal context. In the iterative consistent routing stage, the predicted vectors from the main capsule to each semantic capsule are calculated using a learnable routing transformation matrix. The coupling coefficient is obtained by flexibly maximizing and normalizing the routing log odds. The predicted vectors are weighted and summed using the coupling coefficients and nonlinearly squeezed to obtain the semantic capsule state. The routing log odds are then updated based on the consistency of the dot product between the predicted vector and the semantic capsule state. After iterating a preset number of times, a unified semantic capsule is output.
[0016] As a further aspect of the present invention, in step four, the attention-based cognitive signal injection and hierarchical cross-modal fusion first encodes the semantic capsules of each modality through a shared Transformer module using self-attention, and then performs hierarchical fusion in the following four steps using bidirectional collaborative attention: the first step is to synchronize audio and video to obtain audio and video representations; the second step is to focus on fast cognitive signals from slow cognitive signals to obtain reflective reasoning embeddings; the third step is to focus on the reflective reasoning embeddings from text to obtain cognitively enhanced text representations; and the fourth step is to finally fuse the audio and video representations with the cognitively enhanced text representations to obtain audio-video-text fused representations.
[0017] As a further aspect of the present invention, the adaptive weighted fusion in step four involves adaptively weighting and summing the six modal features—audio-video-text fusion representation, text representation, slow cognitive signal features and fast cognitive signal features after self-attention encoding, audio representation, and video representation—using weights obtained by flexible maximization normalization of learnable parameter vectors to obtain the final fused representation. In step five, a lightweight classifier consisting of one-dimensional convolution, average pooling, and a multilayer perceptron is used to output the probability distribution of each category through flexible maximization, and the category with the highest probability is taken as the anomaly detection result for the risk event.
[0018] Furthermore, the present invention also provides a multimodal risk event anomaly detection system based on the collaboration of a fast and slow dual system and a large and small model, comprising: Data acquisition module: used to acquire the multimodal input data to be detected, including video frames, audio waveforms, title text, optical character recognition text of embedded text in the video frame, and release time; The fast system overall anomaly perception module is used to detect anomalies in the multimodal input data in three dimensions: visual, text, and cross-modal, using a large visual language model guided by prompt templates. Then, the language model is used to reflect on and summarize the detected anomalies to obtain a compact fast cognitive signal. Slow system evidence reasoning module: It is used to decompose the multimodal input data into several independently verifiable atomic statements using a visual language big model and generate a global summary. For each atomic statement, external evidence is obtained through a multi-source evidence retrieval operator. Then, the language model is used to perform a consistency analysis of the statement and evidence to obtain a slow cognitive signal composed of the judgment conclusion and the basis. The large and small model collaborative fusion module is used to perform multimodal encoding on video, audio, text, and the fast and slow cognitive signals respectively through dual encoders, perform dynamic routing semantic distillation on the features of each modality to obtain semantic capsules of uniform length, and then perform attention-based cognitive signal injection and hierarchical cross-modal fusion, and perform adaptive weighted fusion on the representations of each modality to obtain the final fused representation. Risk assessment module: This module is used to input the final fused representation into a lightweight classifier and output the detection result of whether the multimodal input data to be detected is an abnormal risk event.
[0019] Compared with existing technologies, the beneficial effects of this invention are as follows: First, it requires no training or fine-tuning of large models, making it a plug-and-play inference framework with low deployment costs and easy scalability; Second, it reconstructs detection into evidence-aware verification, introducing explicit statement decomposition and evidence grounding, which can effectively identify semantic-level risk events that are "perceptually consistent but factually incorrect," significantly improving detection accuracy; Third, through dynamic routing of semantic distillation and attentional cognitive signal injection, it efficiently integrates high-level inference signals from large models while retaining fine-grained perceptual cues, resulting in a small number of trainable parameters and high inference efficiency; Fourth, the framework outputs a traceable inference path from overall anomaly perception to statement-level evidence verification, significantly improving the interpretability of the judgment and providing reliable technical support for high-risk scenarios such as public opinion security. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the overall process of the method of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] like Figure 1As shown, this invention proposes a multimodal risk event anomaly detection method based on the collaboration of fast and slow dual systems and large and small models. As a training-free, plug-and-play reasoning framework, its core idea is to use the fast system for overall anomaly perception, the slow system for evidence-based fact verification, and to inject high-level cognitive signals from both into a compact multimodal perceptual representation through lightweight large and small model collaboration, thereby forcing the final judgment to be based simultaneously on perceptual cues and external factual evidence. The method includes five steps: data acquisition, fast system overall anomaly perception, slow system evidence reasoning, large and small model collaborative fusion, and risk assessment. The implementation of each step is described in detail below with reference to the accompanying drawings.
[0023] Step 1: Data Acquisition.
[0024] This step marks the beginning of the method, where the system acquires the multimodal input data to be detected. The multimodal input data is a multimodal data set, including video frame V, audio waveform A, user-provided title text T, text O obtained through optical character recognition (OCR) of embedded text in the video frame, and the content's publication time. For example, in a risk event detection scenario, video frame V could be a keyframe sequence of a short video, title text T could be a title such as "Latest Developments in a Major Event," and O could be a specific statement embedded in the on-screen subtitles. For simplicity, this multimodal input will be uniformly represented as (V, A, T, O) in the following text.
[0025] Step two: Rapid system-wide anomaly detection.
[0026] This step corresponds to Figure 1 The fast system (System 1) aims to perform rapid overall anomaly detection on multimodal inputs, capturing inconsistencies between the perceptual and semantic levels. This step includes two nodes: anomaly cue detection and reflective summarization.
[0027] Node 1: Large-scale anomaly detection. In one embodiment, a visual language large-scale model is used to perform structured analysis of multimodal input through thought chain cues. Given video frame V and auxiliary text input (including title text)... Optical Character Recognition (OCR) Text With release time ), in the abnormal clue detection prompt template Guided by this, the visual language big data model extracts anomalous cues from three dimensions: visual, textual, and cross-modal. In the formula, , , These represent visual anomaly cues, textual anomaly cues, and cross-modal anomaly cues, which together constitute the cue set. ; Indicates in the prompt template Anomaly detection process of a large-scale visual language model under induced conditions; V represents a video frame. This indicates the title text. Indicates optical character recognition text, Indicates the publication time; This template represents an anomaly detection prompt. Its guiding model analyzes aspects such as visual realism, timeliness, character behavior, audio-visual synchronization, textual logical consistency and emotional tendency, cross-modal content consistency and information conflict.
[0028] Node Two: Language Model Reflection Summary. The anomalous cues generated by the large visual language model may be lengthy and redundant. To obtain a compact and complete representation, a language model is further employed in the reflection summary prompt template. Guided by this, the clues are distilled into a concise summary: In the formula, , , These represent compact summaries of visual, textual, and cross-modal anomalous cues after distillation, which together constitute rapid cognitive signals. , serving as a fast-perceived signal for subsequent reasoning and fusion; Indicates in the prompt template The reflexive summarization process of the induced language model compresses the three parts of video analysis, text analysis and cross-modal verification respectively while strictly preserving the original semantics; , , Same meaning as above; A template for indicating a reflection summary.
[0029] Step 3, Slow System Evidence Reasoning.
[0030] This step corresponds to Figure 1 The slow system (System 2) is used in this context. While the fast system can quickly screen suspicious content, it cannot establish factual accuracy. Therefore, the slow system performs prudent and evidence-based verification by explicitly reasoning about external knowledge. It consists of three nodes: atomic statement decomposition, multi-source evidence retrieval, and consistency analysis.
[0031] Node 1, Atomic Declaration Decomposition. To achieve accurate verification, a large visual language model is used to decompose the short video into a set of verifiable atomic declarations. Given video frame V, title text... With optical character recognition text Extract the hint template in the atomic declaration Decomposition is carried out under the guidance of: In the formula, C(2) represents the set of declarations consisting of K atomic declarations. The k-th atom is declared It consists of a triad of standardized statement, declared semantic type and suggested verification method, namely a complete, specific, self-contained statement that can uniquely identify the event and clearly include the subject, behavior and context; S(2) represents a one-sentence global summary of the core news facts, which is used to resolve ambiguity and provide global context in subsequent verification; Indicates in the prompt template Atomic declaration extraction process of a large-scale visual language model under induced conditions; V, , These represent video frames, title text, and optical character recognition text, respectively. This indicates the template for extracting atomic declarations; K represents the number of atomic declarations.
[0032] Node 2, Multi-source Evidence Retrieval. Given the extracted set of atomic statements C(2), for each atomic statement... External evidence is obtained through a multi-source evidence retrieval operator: In the formula, This refers to a multi-source evidence retrieval operator, which in one embodiment is implemented by a retrieval agent. It directly uses the standardized statements of atomic claims as retrieval queries without rewriting them, and can enable multiple content sources such as public network searches, authorized news articles and structured knowledge entries. For each query, it returns a single most relevant piece of evidence. This indicates the declaration of the k-th atom; This indicates the link to the source of the retrieved evidence. Indicate the type of evidence source. Indicates the time of evidence release. This indicates the evidence fragments returned by the retrieval.
[0033] Node 3, Consistency Analysis. For each statement-evidence pair... The language model checks the prompt template. Under the guidance of [unclear], a consistency analysis of evidence grounding was conducted: In the formula, Indicates in the prompt template The verification process of the language model under guidance comprehensively evaluates the consistency of statements and evidence in terms of time, entity and causal logic. When the evidence fragments themselves are insufficient, the global summary S(2) provides the global context. S(2) represents the declaration of the k-th atom, and S(2) represents the global summary. This indicates the search evidence corresponding to the statement; This is a template for displaying consistency verification prompts; This indicates the verification judgment conclusion for the k-th statement, and its value can be one of "support", "contradictory" or "uncertain". This provides a brief explanation of the basis for the judgment.
[0034] Aggregating the verification results of all claims yields the evidence-based reasoning signal, i.e., the slow cognitive signal: In the formula, V(2) represents the aggregated slow cognitive signal, which complements the overall abnormal cues of the fast system; This indicates that the verification judgment conclusion and the basis explanation for the k-th statement are correct; K represents the number of atomic statements.
[0035] Step 4: Collaborative fusion of large and small models.
[0036] This step fuses the overall anomaly summary S(1) from the fast system and the evidence verification result V(2) from the slow system with the perceptual representation obtained from the original short video input encoding. Since the two are heterogeneous in nature—fast and slow cognitive signals are semantic, sparse, and high-level, while the downstream classifier operates on dense multimodal perceptual features—forcing a small detector to mimic the output of a large model would result in the loss of crucial fine-grained perceptual cues. Therefore, this invention treats the cognitive output as an auxiliary semantic modality, injecting compact multimodal representations through attention-based aggregation to achieve efficient and lightweight collaboration between large and small models. This step sequentially includes four stages: dual-encoder multimodal encoding, dynamic routing semantic distillation, attention-based cognitive signal injection and hierarchical fusion, and adaptive weighted fusion.
[0037] The first step involves dual-encoder multimodal encoding. Short video risk detection requires sensitivity to both semantic content and manipulation-related forgery traces. Since a single encoder often favors only one aspect, this invention employs a dual-encoder strategy, combining a semantic encoder and a perceptual encoder for each modality. In the formula, , , Let f(1) represent the modal features of video, audio, and text, respectively, and f(2) represent the features extracted from the fast cognitive signal S(1) and f(2) represent the features extracted from the slow cognitive signal V(2). This indicates a concatenation operation along the feature dimension; This represents a 3D convolutional video encoder that models low-level motion and style patterns. This represents a cross-modal encoder for modeling high-level visual-linguistic semantics. An audio encoder that represents the modeling of speech content. An audio encoder representing the acoustic properties of a model. The text encoder represents the semantics of the modeling language; V, A, and T represent video frames, audio waveforms, and text input, respectively; S(1) and V(2) represent fast cognitive signals and slow cognitive signals, respectively; the fast and slow cognitive signals are encoded using the same backbone encoder as the text to ensure that they are in the same representation space as the perceptual features; the encoders mentioned above All are frozen pre-trained encoders.
[0038] The second step is dynamic routing semantic distillation. The modal features obtained from equation (9) have heterogeneous sequence lengths and feature dimensions. To map them to a unified latent space and suppress temporal noise, this invention proposes dynamic routing semantic distillation, denoted as: In the formula, This represents the unified semantic capsule obtained after semantic distillation of mode m. In one embodiment, the number of semantic capsules L is 32, and the capsule dimension d is 128; This represents a dynamic routing semantic distillation operator, which differs from static pooling. It uses iterative "routing by consistency" as a denoising filter to actively suppress temporal outliers and distill sparse cues into compact and aligned semantic capsules. The input features represent mode m. These correspond to video, audio, text, fast cognitive signals, and slow cognitive signals, respectively.
[0039] Dynamic routing semantic distillation operator The process comprises two stages: master capsule generation and iterative consistent routing. In the master capsule generation stage, the original features are first linearly projected and then projected through a Transformer layer to form the master capsule sequence, thereby capturing the local temporal context. In the formula, U represents the main capsule sequence. Its i-th row This corresponds to the main capsule at the original i-th time step; LayerNorm represents layer normalization; MHA represents multi-head attention, which is implemented here as a Transformer layer for modeling local temporal context; Indicates the input linear projection matrix; This represents the input feature sequence of mode m. , Let m be the sequence length of mode m.
[0040] In the iterative consensus routing phase, the variable-length sequence U is distilled into fixed-length semantic capsules. A dynamic routing algorithm is used. Let... ( ) as the main capsule, ( Let be the j-th target semantic capsule, and calculate the predicted vector using the learnable transformation matrix: In the formula, This represents the prediction vector of the j-th semantic capsule predicted by the i-th master capsule; This represents the learnable routing transformation matrix. ; This represents the i-th master capsule. ; Indicates the sequence number of the target semantic capsule.
[0041] The routing process is performed r times to refine the coupling coefficients, which are obtained by normalizing the log-probability of the routes through flexible maximization. In the formula, The coupling coefficient represents the degree of consistency between the i-th main capsule and the j-th semantic anchor. This represents the logarithmic probability of the route, initially set to 0. represents the natural exponential function; L represents the number of semantic capsules; i is the primary capsule index, and j and k are semantic capsule indices.
[0042] The candidate representation of the j-th semantic capsule is obtained by weighting the prediction vectors by the coupling coefficient, thus effectively filtering out noisy frames (which have smaller coupling coefficients): In the formula, Represents a candidate representation of the j-th semantic capsule; Represents the coupling coefficient; Represents the prediction vector; This indicates the length of the input main capsule sequence.
[0043] Candidate representations are normalized using a nonlinear squeezing function to obtain semantic capsule states, and the route log odds are updated based on dot product consistency. In the formula, Indicates the state of the j-th semantic capsule; This represents a nonlinear compression function that compresses the length of a vector to the interval (0,1), while preserving the vector direction and making the length a representation of the confidence level of its existence. Indicates a candidate representation; This indicates an assignment update; Indicates the logarithmic probability of the route; The dot product of the predicted vector and the semantic capsule state is used to measure their consistency. The routing iteration is performed a total of r times; in one embodiment, the number of routing iterations r is 3. After the iterations, a unified semantic capsule is obtained. This serves as the alignment input for the subsequent cognitive fusion module.
[0044] The third stage involves attention-based cognitive signal injection and hierarchical fusion. This invention does not force explicit distillation from a large model, but instead injects cognitive signals through attention-based interactions, allowing smaller detectors to selectively absorb evidence-based reasoning cues while maintaining sensitivity to low-level manipulation patterns. First, the semantic capsule for each modality undergoes self-attention encoding via a shared Transformer module: In the formula, z represents a certain modal semantic capsule of input (i.e., obtained from formula (10)). ); The features after self-attention encoding are represented as follows: for video, audio, text, fast cognitive signals, and slow cognitive signals, respectively... , , , , LN represents layer normalization; FFN represents feedforward network; MHA(z, z, z) represents multi-head self-attention where the query, key, and value are all z.
[0045] Subsequently, bidirectional collaborative attention is used to integrate the various modalities step by step. Given two representations x and y, the collaborative attention operation CoAttn(x, y) is defined as follows: In the formula, , Let x and y represent the representations of x and y after bidirectional collaborative attention interaction, respectively; MHA(x, y, y) represents multi-head attention with x as the query and y as the key and value, and MHA(y, x, x) represents the opposite; LN represents layer normalization. Pairwise fusion between the two representations is achieved by a lightweight projection ψ: In the formula, ψ(x, y) represents the pairwise fusion result of characterizing x and y; σ represents the activation function; Represents the projection weight matrix. [x || y] represents the bias vector; [x || y] represents the concatenation of x and y along the feature dimension.
[0046] Based on this, layered fusion is performed in the following four steps. The first step is audio-video synchronization, which begins with audio-video alignment based on the perceptual stream: In the formula, , This represents the representation of video and audio after collaborative attention interaction; This represents the fused audio and video representation; , These are the video and audio features encoded with self-attention; CoAttn and ψ have the same meaning as above. The second step involves the slow system focusing on the fast system to absorb reflective reasoning information: In the formula, This indicates the reflective reasoning embedding after being integrated into the fast system's reflective reasoning; This represents the slow cognitive signal characteristics after self-attention encoding. This represents the fast cognitive signal characteristics following self-attention encoding; Indicates For query, with Multi-head attention for keys and values; LN denotes layer normalization. The third step involves text focusing on the aforementioned cognitively enhanced representations: In the formula, This represents the text representation after cognitive signal enhancement; Represents text features after self-attention encoding; Indicates reflective reasoning embedding; Indicates For query, with Multi-head attention for keys and values; LN denotes layer normalization. The fourth step is the final audio / video-text fusion: In the formula, , Representation of audio and video With cognitively enhanced text representation Representation after collaborative attention interaction; This represents the final audio-video-text fusion representation; CoAttn and ψ have the same meaning as above. The above sequence is empirically stable for training: first, audio-video synchronization provides grounding for the perceptual stream; then, the slow system queries the fast system to obtain reflective reasoning embeddings, which in turn guides text refinement, and finally, it is fused with the audio-video stream.
[0047] The fourth step is adaptive weighted fusion. The final fused representation is obtained by adaptively weighting and summing the features from the six modalities: In the formula, r represents the final fusion representation; The adaptive fusion weights of the m-th modal features are represented by the weight vector α, which is obtained by normalizing the learnable parameter vector w using the softmax function; six modal features ( The values are sequentially taken as the final audio / video-text fusion representation. Text representation Slow cognitive signal characteristics after self-attention encoding Fast cognitive signal features after self-attention encoding Audio representation With video representation This allows for the joint encoding of perceptual and cognitive evidence; the adaptive weighting enables the detector to dynamically balance perceptual cues and cognitive inference signals while maintaining a compact model size.
[0048] Step 5: Risk assessment.
[0049] Input the final fused representation r obtained in step four into the lightweight classifier, and output the risk event anomaly detection result: In the formula, The predicted risk event anomaly detection result is represented as "normal" or "risk event anomaly"; r represents the final fused representation; Conv1D represents one-dimensional convolution, AvgPool represents average pooling, MLP represents multilayer perceptron, and softmax represents the soft maximization function, used to output the probability distribution of each category; the category with the highest probability is taken as the final decision. In one embodiment, the unified feature dimension d is 128, the semantic capsule length L is 32, and the fusion module can complete the cognitive signal injection with a small number of trainable parameters, thereby ensuring high efficiency of overall inference and being suitable for practical deployment.
[0050] In summary, without training or fine-tuning a large model, this invention unifies overall anomaly perception and evidence-based fact-checking into a single framework through the synergy of fast and slow dual systems and the fusion of large and small models. It can identify forgery traces at the perception level and semantic-level risk events that are "perceptually consistent but factually incorrect," and provides a traceable reasoning path from overall anomaly perception to declaration-level evidence verification, significantly improving detection accuracy and interpretability.
[0051] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This method of description is merely for clarity, and those skilled in the art should consider the specification as a whole. The technical solutions in the various embodiments can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multimodal risk event anomaly detection method based on the collaboration of a fast and slow dual system and a large-small model, characterized in that, Includes the following steps: Step 1, Data Acquisition: Acquire the multimodal input data to be detected, which includes video frames, audio waveforms, title text, optical character recognition text embedded in the video frame, and release time; Step 2, Fast System Overall Anomaly Perception: Using a large visual language model and guided by prompt templates, anomaly cues are detected in the multimodal input data in three dimensions: visual, text, and cross-modal. Then, the language model is used to reflect on and summarize the detected anomaly cues to obtain a compact fast cognitive signal. Step 3, Slow System Evidence Reasoning: The multimodal input data is decomposed into several independently verifiable atomic statements using a visual language big model and a global summary is generated. External evidence is obtained for each atomic statement through a multi-source evidence retrieval operator. Then, the language model is used to perform a consistency analysis of the statement and the evidence to obtain a slow cognitive signal composed of the judgment conclusion and the basis. Step 4, large and small model collaborative fusion: Multimodal encoding is performed on video, audio, text and the fast cognitive signal and slow cognitive signal by dual encoders. Dynamic routing semantic distillation is performed on the features of each modality to obtain semantic capsules of uniform length. Then, attention-based cognitive signal injection and hierarchical cross-modal fusion are performed, and adaptive weighted fusion is performed on the representations of each modality to obtain the final fused representation. Step 5, Risk Assessment: Input the final fused representation into a lightweight classifier and output the detection result of whether the multimodal input data to be detected is an abnormal risk event.
2. The multimodal risk event anomaly detection method based on fast and slow dual systems and large and small models in accordance with claim 1, characterized in that, Step two includes two nodes: anomaly detection and reflective summarization. Taking video frames, title text, optical character recognition text, and release time as input, the visual language model, guided by anomaly detection prompt templates, extracts anomaly cues in three dimensions: visual, text, and cross-modal. Then, guided by reflective summarization prompt templates, the language model compresses the above anomaly cues into compact summaries while preserving the original semantics. The three together constitute the fast cognitive signal.
3. The multimodal risk event anomaly detection method based on fast and slow dual systems and large and small models in accordance with claim 1, characterized in that, Step three includes three nodes in sequence: atomic statement decomposition, multi-source evidence retrieval, and consistency analysis. Under the guidance of the atomic statement extraction prompt template, the visual language big model decomposes the content to be detected into several independently verifiable atomic statements consisting of standardized statements, statement semantic types, and suggested verification methods, and generates a one-sentence global summary. For each atomic statement, the multi-source evidence retrieval operator directly uses its standardized statement as the retrieval query to return the most relevant evidence fragments from multiple sources such as public networks, authorized news, and structured knowledge. Guided by the verification prompt template, the language model then performs a consistency analysis on the statements and evidence by integrating time, entity, and causal logic, outputting a judgment conclusion with values of support, contradiction, or uncertainty, along with an explanation of the basis, and aggregating the verification results of all statements into the slow cognitive signal.
4. The multimodal risk event anomaly detection method based on fast and slow dual systems and large and small models in accordance with claim 1, characterized in that, In step four, when performing multimodal encoding on each modality, modal features are extracted for video, audio, text, and the fast cognitive signal and slow cognitive signal using a dual encoder strategy that combines a semantic encoder and a perceptual encoder. The fast cognitive signal and slow cognitive signal are encoded using the same backbone encoder as the text, and all encoders used are frozen pre-trained encoders that do not participate in training or fine-tuning.
5. The multimodal risk event anomaly detection method based on fast and slow dual systems and large and small models in accordance with claim 1, characterized in that, The dynamic routing semantic distillation in step four maps each modality feature into a semantic capsule of uniform length. It includes two stages: main capsule generation and iterative consistent routing. In the main capsule generation stage, the input features are linearly projected and then projected through a Transformer layer to form a main capsule sequence carrying local temporal context. In the iterative consistent routing phase, the predicted vectors from the main capsule to each semantic capsule are calculated using the learnable routing transformation matrix. The coupling coefficient is obtained by flexibly maximizing and normalizing the routing log odds. The semantic capsule state is obtained by weighted summation of the predicted vectors using the coupling coefficients and nonlinear compression. The routing log odds are then updated based on the consistency of the dot product between the predicted vectors and the semantic capsule states. After iterating a preset number of times, a unified semantic capsule is output.
6. The multimodal risk event anomaly detection method based on fast and slow dual systems and large and small models in accordance with claim 1, characterized in that, In step four, attention-based cognitive signal injection and hierarchical cross-modal fusion first encodes the semantic capsule of each modality through a shared Transformer module with self-attention, and then uses bidirectional collaborative attention to perform hierarchical fusion in the following four steps: the first step is to synchronize audio and video to obtain audio and video representations; the second step is to use slow cognitive signals to focus on fast cognitive signals to obtain reflective reasoning embeddings. The third step involves embedding the textual focus into the reflective reasoning to obtain a cognitively enhanced textual representation; the fourth step involves finally fusing the audio-visual representation with the cognitively enhanced textual representation to obtain an audio-visual-text fused representation.
7. The multimodal risk event anomaly detection method based on fast and slow dual systems and large and small models in accordance with claim 1, characterized in that, In step four, the adaptive weighted fusion process involves adaptively weighting and summing the six modal features—audio-video-text fusion representation, text representation, slow and fast cognitive signal features after self-attention encoding, audio representation, and video representation—using weights obtained from learnable parameter vectors through flexible maximization normalization, to obtain the final fused representation. In step five, a lightweight classifier consisting of one-dimensional convolution, average pooling, and a multilayer perceptron is used to output the probability distribution of each category through flexible maximization, and the category with the highest probability is taken as the anomaly detection result for the risk event.
8. A system employing the multimodal risk event anomaly detection method based on the collaboration of a fast-slow dual system and a large-small model as described in any one of claims 1-7, characterized in that, include: Data acquisition module: used to acquire the multimodal input data to be detected, including video frames, audio waveforms, title text, optical character recognition text of embedded text in the video frame, and release time; The fast system overall anomaly perception module is used to detect anomalies in the multimodal input data in three dimensions: visual, text, and cross-modal, using a large visual language model guided by prompt templates. Then, the language model is used to reflect on and summarize the detected anomalies to obtain a compact fast cognitive signal. Slow system evidence reasoning module: It is used to decompose the multimodal input data into several independently verifiable atomic statements using a visual language big model and generate a global summary. For each atomic statement, external evidence is obtained through a multi-source evidence retrieval operator. Then, the language model is used to perform a consistency analysis of the statement and evidence to obtain a slow cognitive signal composed of the judgment conclusion and the basis. The large and small model collaborative fusion module is used to perform multimodal encoding on video, audio, text, and the fast and slow cognitive signals respectively through dual encoders, perform dynamic routing semantic distillation on the features of each modality to obtain semantic capsules of uniform length, and then perform attention-based cognitive signal injection and hierarchical cross-modal fusion, and perform adaptive weighted fusion on the representations of each modality to obtain the final fused representation. Risk assessment module: This module takes the final fused representation as input to a lightweight classifier and outputs a detection result indicating whether the multimodal input data to be detected is an abnormal risk event.