基于事件一致性的视听事件检测方法及装置

CN115861879BActive Publication Date: 2026-07-17BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2022-11-25
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing audiovisual event detection methods have shortcomings in background fragment recognition and event semantic consistency, resulting in poor performance and ignoring the semantic consistency of events in the same video.

Method used

This paper proposes an audiovisual event detection method based on event consistency. It adopts audiovisual joint learning and semantic consistency modeling, including feature encoding at the segment level and semantic guidance at the video level. Semantic consistency modeling is performed using a cross-modal event representation extractor and GRU. It combines fully supervised and weakly supervised background class screening loss functions and inter-segment smoothing loss.

Benefits of technology

It improves the accuracy and discriminative power of audiovisual event detection, outperforming existing methods, especially in fully and weakly supervised tasks on the AVE dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115861879B_ABST
    Figure CN115861879B_ABST
Patent Text Reader

Abstract

本发明提出一种基于事件一致性的视听事件检测方法,包括:获取目标视频;将目标视频划分为N个不重叠的连续片段,获取图像流和音频流;对图像流和音频流进行特征提取,获取视听特征;通过视听联合学习将视听特征融合,其中,视听联合学习包括片段层面的特征编码以及视频层面的语义指导;将融合后的视听特征输入分类器中,得到目标视频的预测结果。本发明的方法利用事件的语义一致性来分别指导视觉和听觉模态的学习,可以确保模型更好地聚焦和定位发声对象。
Need to check novelty before this filing date? Find Prior Art