Audio-Visual Repetition Counting via Cross-Modal Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer vision solutions for automatically counting repetitive activities in videos fail in poor sight conditions such as low illumination, occlusion, and camera view changes, as they rely solely on visual content and assume periodic repetitions, making them less effective for non-stationary and 'in the wild' scenarios.

Innovation Solution

A method that processes both audio and video features using neural networks to predict repetitive actions, incorporating a temporal stride decision module and reliability estimation to select the best frame rate and modality-specific predictions, leveraging cross-modal temporal interaction for improved accuracy under challenging conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If vision-based repetition counting is used, then counting capability is provided, but accuracy deteriorates in poor sight conditions such as low illumination, occlusion, and camera view changes

Engineering Contradiction:
Improvecounting accuracyVSAvoidpoor sight conditions
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent combines audio and video modalities into a unified repetition counting system. The audio stream processes sound signals while the video stream processes visual frames, and both streams are merged to produce a final repetition count. This multi-modal fusion allows the system to maintain accuracy when one modality degrades due to poor sight conditions.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces an audio stream as an intermediary modality to compensate for visual degradation. The audio processing pipeline independently analyzes sound signals and provides complementary information that mediates the counting accuracy when visual conditions are poor, such as during occlusion or low illumination.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If existing audio-visual approaches are used, then some improvement in accuracy is achieved, but computational complexity increases due to iterative refinement processes

Engineering Contradiction:
Improverepetition counting accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary action by determining the temporal stride (sampling rate) before the main repetition counting process. The system evaluates multiple candidate frame rates, selects the optimal one based on periodicity detection, and then uses this predetermined stride for the actual counting. This preliminary optimization reduces computational complexity by avoiding iterative refinement during the main counting process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the temporal sampling parameter (frame rate/stride) based on detected periodicity in the signal. By adapting the sampling rate to match the repetition frequency, the system achieves higher accuracy with fewer computations, avoiding the need for complex iterative refinement processes used in other approaches.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If high frame rates are selected for video sampling, then more detailed temporal information is captured, but counting accuracy deteriorates due to omissions and computational burden

Engineering Contradiction:
Improvetemporal information coverageVSAvoidrepetition counting accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent makes the frame sampling rate dynamic by selecting different strides based on the detected periodicity of the repetition signal. Instead of using a fixed high frame rate, the system adapts the sampling interval to match the actual repetition frequency, capturing sufficient temporal information while avoiding the computational burden and potential omissions associated with uniformly high frame rates.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11694442B2Systems and methods for counting repetitive activity in audio video content
Publication Date: 2023.07.04 INCEPTION AI IP LTD
  • US11694442B2 patent drawing
  • US11694442B2 patent drawing
  • US11694442B2 patent drawing

AI summary

Repetitive activities can be captured in audio video content. The AV content can be processed in order to predict the number of repetitive activities present in the AV content. The accuracy of the predicted number may be improved, especially for AV content with challenging conditions, by basing the predictions on both the audio and video portions of the AV content.