Content Streaming Speaker Identification via Spatial Sensor Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems struggle to accurately identify the current speaker in video and audio communications, especially when the speaker is outside the camera's field-of-view, and may prematurely terminate sessions based on lack of physical input, despite user interaction.
Innovation Solution
Systems determine spatial information and user interactions by processing sensor data from image, audio, and location sensors to identify the speaker's position and identity, and monitor user engagement beyond physical inputs, providing this information alongside video and audio to accurately identify speakers and manage session continuity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a camera with a limited field-of-view is used to capture video, then the device complexity is reduced, but the ability to identify speakers outside the field-of-view deteriorates
Solution Approach 1:
The patent combines multiple sensing modalities (audio sensors, image sensors, location sensors) into an integrated system that processes data from all sources to determine speaker identity and position. This merging allows the system to identify speakers even when they are outside the camera's field-of-view by using audio and location data to supplement visual information.
Solution Approach 2:
The patent introduces spatial information as an intermediary element that bridges the gap between limited camera coverage and comprehensive speaker identification. By processing location data and audio data alongside video, the system creates a more complete picture of who is speaking, even when the speaker is not visible in the video feed.
2Loss of energy
If session termination is based on lack of physical input, then computing resources are conserved, but user interaction detection accuracy deteriorates
Solution Approach 1:
The patent makes the interaction detection system universal by accepting multiple types of input signals beyond physical device interactions. The system processes audio data, image data, and location data to detect various forms of user engagement, such as speaking or presenting content, allowing it to accurately determine whether a user is still interacting with the application even without physical input.
Solution Approach 2:
The patent implements feedback mechanisms where the system continuously monitors multiple data streams (audio, image, location) to determine user engagement status. This feedback loop allows the system to adjust session termination decisions based on comprehensive evidence of user interaction, preventing premature termination while still conserving resources when users are truly inactive.
Data Source
AI summary
In various examples, providing spatial information for conversational systems and applications is described herein. Systems and methods are disclosed that determine information associated with users that are speaking, such as positions of the users with respect to devices and/or identifiers associated with the users, and then provide the information along with videos and/or audio captured using the devices. For instance, a first device may generate image data using one or more image sensors, audio data using one or more microphones, and/or location data using one or more location sensors. The image data, the audio data, and/or the location data may then be processed to determine the information associated with a user that is speaking. A second device that is presenting a video represented by the image data and/or outputting sound represented by the audio data may then further present content associated with the information.


