Video detection method and device based on multi-modal thought chain and direct preference optimization
By employing a multimodal thinking chain and direct preference optimization video detection method, the lack of interpretability in existing technologies is addressed, achieving high efficiency, interpretability, and accuracy in negative video detection. This method generates logically rigorous explanatory text, thereby improving detection performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DALIAN UNIV OF TECH
- Filing Date
- 2026-03-27
- Publication Date
- 2026-07-03
AI Technical Summary
Existing negative video detection technologies cannot provide clear explanations, struggle to capture the obscure and ambiguous content in negative videos, and are prone to falling into false relevance, leading to misjudgments and a lack of interpretability, thus failing to meet the legal and regulatory requirements for algorithm interpretability.
We employ a video detection method based on multimodal thinking chain and direct preference optimization. By constructing enhanced input samples containing multi-source information, we utilize an open-source visual language large model for feature extraction and encoding. Combined with supervised fine-tuning loss and direct preference optimization algorithm, we generate target prediction sequences containing prediction labels and reasons.
It improves the interpretability and detection performance of video detection, generates logically coherent and well-supported explanatory texts, and significantly enhances the model's detection accuracy and persuasiveness on complex boundary samples.
Smart Images

Figure CN122336626A_ABST