Video detection method and device based on multi-modal thought chain and direct preference optimization

By employing a multimodal thinking chain and direct preference optimization video detection method, the lack of interpretability in existing technologies is addressed, achieving high efficiency, interpretability, and accuracy in negative video detection. This method generates logically rigorous explanatory text, thereby improving detection performance.

CN122336626APending Publication Date: 2026-07-03DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2026-03-27
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing negative video detection technologies cannot provide clear explanations, struggle to capture the obscure and ambiguous content in negative videos, and are prone to falling into false relevance, leading to misjudgments and a lack of interpretability, thus failing to meet the legal and regulatory requirements for algorithm interpretability.

Method used

We employ a video detection method based on multimodal thinking chain and direct preference optimization. By constructing enhanced input samples containing multi-source information, we utilize an open-source visual language large model for feature extraction and encoding. Combined with supervised fine-tuning loss and direct preference optimization algorithm, we generate target prediction sequences containing prediction labels and reasons.

Benefits of technology

It improves the interpretability and detection performance of video detection, generates logically coherent and well-supported explanatory texts, and significantly enhances the model's detection accuracy and persuasiveness on complex boundary samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122336626A_ABST
    Figure CN122336626A_ABST
Patent Text Reader

Abstract

This invention discloses a video detection method and apparatus based on multimodal thinking chain and direct preference optimization. The method includes: acquiring a set of video samples with known labels as a training set, constructing enhanced input samples containing multi-source information, and further optimizing an open-source visual language model to generate predicted labels and inferences for unknown videos. By constructing a two-stage framework including information enhancement and inference enhancement, combined with multimodal thinking chain and direct preference optimization algorithms, the method mathematically ensures the model maximizes mutual information utilization of multimodal contextual information and effectively suppresses the interference of spurious relevance through a contrastive learning mechanism. Experimental results show that this invention achieves good detection performance and exhibits significant advantages in interpreting the amount of generated information, logicality, and persuasiveness.
Need to check novelty before this filing date? Find Prior Art