Streaming User Activity Detection with Multimodal Speaker Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional systems struggle to accurately identify the speaker in video conferencing and streaming applications when the speaker is outside the camera's field-of-view, and they prematurely terminate sessions based on lack of physical input, even if users are still interacting.

Innovation Solution

Systems and methods that utilize image, audio, and location data to determine the position and identity of the speaker, providing spatial information alongside video and audio to identify the speaker, and monitor user interactions beyond physical inputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a camera with a limited field-of-view is used to capture video, then the device complexity is reduced, but the ability to identify speakers outside the field-of-view deteriorates

Engineering Contradiction:
Improvecamera systemVSAvoidspeaker identification
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple sensing modalities (audio microphones, image sensors, location sensors) into an integrated system that works together to identify speakers. The audio data from microphones is fused with image data and location data to determine speaker identity and position, even when the speaker is outside the camera's field-of-view.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces location sensors as an intermediary component that provides spatial information about users in the environment. This location data acts as a mediator between the audio data (which identifies who is speaking) and the visual data (which shows who is visible), enabling the system to identify and present information about speakers regardless of their position relative to the camera.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of energy

If session termination is based solely on lack of physical input, then computing resources are conserved, but user interaction monitoring accuracy deteriorates

Engineering Contradiction:
Improvecomputing resourcesVSAvoiduser interaction detection
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The system makes sensor data (image, audio, location) serve multiple functions: it is used both for speaker identification during communication and for monitoring user interaction status. This multi-functional use of the same data sources allows the system to detect various types of interactions (speaking, moving, presenting content) without requiring additional dedicated sensors or resources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses the device's own existing sensors (microphones, image sensors, location sensors) to monitor user interaction and determine session continuation. Rather than requiring external monitoring systems or additional input devices, the device self-monitors its own usage patterns through its built-in sensors, enabling automatic session management based on actual user engagement.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250247580A1User activity detection for content streaming systems and applications
Publication Date: 2025.07.31 NVIDIA CORP
  • US20250247580A1 patent drawing
  • US20250247580A1 patent drawing
  • US20250247580A1 patent drawing

AI summary

In various examples, monitoring user interactions for content streaming systems and applications is described herein. For instance, a system(s) that is providing content to a device, such as during an online session associated with an application, may receive sensor data from the device and then use the sensor data to determine whether a user is interacting with the application. The sensor data may include image data representing the user, audio data representing from the user, location data representing a location of the user, input data representing one or more inputs from the user, and/or the like. The system(s) may then determine whether to continue the session or terminate the session based at least on whether the user is interacting with the application. For example, the system(s) may determine to terminate the session based at least on the user not interacting with the application for a threshold amount of time.