Audio-Visual Repetition Counting via Cross-Modal Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer vision solutions for automatically counting repetitive activities in videos fail in poor sight conditions such as low illumination, occlusion, and camera view changes, as they rely solely on visual content and assume periodic repetitions, making them less effective for non-stationary and 'in the wild' scenarios.
Innovation Solution
A method that processes both audio and video features using neural networks to predict repetitive actions, incorporating a temporal stride decision module and reliability estimation to select the best frame rate and modality-specific predictions, leveraging cross-modal temporal interaction for improved accuracy under challenging conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If vision-based repetition counting is used, then counting capability is provided, but accuracy deteriorates in poor sight conditions such as low illumination, occlusion, and camera view changes
Solution Approach 1:
The patent combines audio and video modalities into a unified repetition counting system. The audio stream processes sound signals while the video stream processes visual frames, and both streams are merged to produce a final repetition count. This multi-modal fusion allows the system to maintain accuracy when one modality degrades due to poor sight conditions.
Solution Approach 2:
The patent introduces an audio stream as an intermediary modality to compensate for visual degradation. The audio processing pipeline independently analyzes sound signals and provides complementary information that mediates the counting accuracy when visual conditions are poor, such as during occlusion or low illumination.
2Measurement precision
If existing audio-visual approaches are used, then some improvement in accuracy is achieved, but computational complexity increases due to iterative refinement processes
Solution Approach 1:
The patent performs preliminary action by determining the temporal stride (sampling rate) before the main repetition counting process. The system evaluates multiple candidate frame rates, selects the optimal one based on periodicity detection, and then uses this predetermined stride for the actual counting. This preliminary optimization reduces computational complexity by avoiding iterative refinement during the main counting process.
Solution Approach 2:
The patent changes the temporal sampling parameter (frame rate/stride) based on detected periodicity in the signal. By adapting the sampling rate to match the repetition frequency, the system achieves higher accuracy with fewer computations, avoiding the need for complex iterative refinement processes used in other approaches.
3Loss of information
If high frame rates are selected for video sampling, then more detailed temporal information is captured, but counting accuracy deteriorates due to omissions and computational burden
Solution Approach 1:
The patent makes the frame sampling rate dynamic by selecting different strides based on the detected periodicity of the repetition signal. Instead of using a fixed high frame rate, the system adapts the sampling interval to match the actual repetition frequency, capturing sufficient temporal information while avoiding the computational burden and potential omissions associated with uniformly high frame rates.
Data Source
AI summary
Repetitive activities can be captured in audio video content. The AV content can be processed in order to predict the number of repetitive activities present in the AV content. The accuracy of the predicted number may be improved, especially for AV content with challenging conditions, by basing the predictions on both the audio and video portions of the AV content.


