Streaming User Activity Detection with Multimodal Speaker Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems struggle to accurately identify the speaker in video conferencing and streaming applications when the speaker is outside the camera's field-of-view, and they prematurely terminate sessions based on lack of physical input, even if users are still interacting.
Innovation Solution
Systems and methods that utilize image, audio, and location data to determine the position and identity of the speaker, providing spatial information alongside video and audio to identify the speaker, and monitor user interactions beyond physical inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a camera with a limited field-of-view is used to capture video, then the device complexity is reduced, but the ability to identify speakers outside the field-of-view deteriorates
Solution Approach 1:
The patent combines multiple sensing modalities (audio microphones, image sensors, location sensors) into an integrated system that works together to identify speakers. The audio data from microphones is fused with image data and location data to determine speaker identity and position, even when the speaker is outside the camera's field-of-view.
Solution Approach 2:
The system introduces location sensors as an intermediary component that provides spatial information about users in the environment. This location data acts as a mediator between the audio data (which identifies who is speaking) and the visual data (which shows who is visible), enabling the system to identify and present information about speakers regardless of their position relative to the camera.
2Loss of energy
If session termination is based solely on lack of physical input, then computing resources are conserved, but user interaction monitoring accuracy deteriorates
Solution Approach 1:
The system makes sensor data (image, audio, location) serve multiple functions: it is used both for speaker identification during communication and for monitoring user interaction status. This multi-functional use of the same data sources allows the system to detect various types of interactions (speaking, moving, presenting content) without requiring additional dedicated sensors or resources.
Solution Approach 2:
The system uses the device's own existing sensors (microphones, image sensors, location sensors) to monitor user interaction and determine session continuation. Rather than requiring external monitoring systems or additional input devices, the device self-monitors its own usage patterns through its built-in sensors, enabling automatic session management based on actual user engagement.
Data Source
AI summary
In various examples, monitoring user interactions for content streaming systems and applications is described herein. For instance, a system(s) that is providing content to a device, such as during an online session associated with an application, may receive sensor data from the device and then use the sensor data to determine whether a user is interacting with the application. The sensor data may include image data representing the user, audio data representing from the user, location data representing a location of the user, input data representing one or more inputs from the user, and/or the like. The system(s) may then determine whether to continue the session or terminate the session based at least on whether the user is interacting with the application. For example, the system(s) may determine to terminate the session based at least on the user not interacting with the application for a threshold amount of time.


