Selective Audio Playback Using Voice Profiles for Unwanted Sound Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional entertainment systems do not allow users to dynamically adjust audio features based on personal preferences, leading to distractions from unwanted sounds during media consumption, such as background noise or specific commentators, which disrupt the user's experience.
Innovation Solution
A system that separates audio and video streams, identifies and catalogs sounds, and adjusts output characteristics based on user preferences by muting, attenuating, or converting unwanted audio segments, using voice profiles and metadata to synchronize desired audio with the video stream.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If manual muting is used to prevent unwanted sounds, then unwanted sounds are blocked, but user convenience deteriorates due to constant remote inputs required
Solution Approach 1:
The system automatically identifies and mutes unwanted sounds without requiring continuous user input. The consumption device performs self-service by autonomously analyzing audio segments, comparing them against user preferences, and adjusting audio output characteristics dynamically during media playback.
Solution Approach 2:
The system dynamically adjusts audio output characteristics in real-time based on changing audio content and user preferences. Unlike static manual muting, the system continuously monitors audio segments and adapts the audio stream dynamically to maintain user convenience while blocking unwanted sounds.
2Adaptability or versatility
If conventional entertainment systems are used, then system simplicity is maintained, but user personalization capability deteriorates as all users must consume content the same way
Solution Approach 1:
The audio stream is segmented into discrete audio segments that can be individually analyzed and processed. Each segment is evaluated against user preferences, allowing selective adjustment of specific sounds while maintaining overall system simplicity. This segmentation enables personalization without requiring complete system redesign.
Solution Approach 2:
The system introduces an intermediary audio processing layer between the media asset and the user. This intermediary component handles the complexity of audio analysis and preference matching, allowing the core entertainment system to remain simple while adding advanced personalization capabilities through the mediating processing layer.
3Productivity
If audio and video are transmitted as separate segments and processed at client device, then transmission efficiency is improved, but processing load on consumption device increases
Solution Approach 1:
Audio segments are pre-processed and tagged with metadata identifying sound sources and characteristics before transmission to the consumption device. This preliminary action at the server side reduces the processing burden on the client device, as the heavy lifting of audio analysis has already been performed during encoding and segmentation.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Systems and methods are presented for providing to filter unwanted sounds from a media asset. Voice profiles of a first character and a second character are generated based on a first voice signal and a second voice signal received from the media device during a presentation. The user provides a selection to avoid a certain sound or voice in association with the second character. During a presentation of the media asset, a second audio segment is analyzed to determine, based on the voice profile of the second character, whether the second voice signal includes the voice of a second character. If so, the second voice signal output characteristics are adjusted to reduce the sound.