Hearing System Camera-Based Audio Source Identification in Complex Scenes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Hearing devices often fail to adequately support users in complex acoustic environments due to incorrect settings or misclassification of acoustic environments, making it challenging for users to adjust settings effectively.
Innovation Solution
A hearing system equipped with a camera and display, utilizing artificial intelligence or neural networks to identify audio sources visually, allowing users to easily select and modify sound inputs through a graphical interface, enabling quick adjustments in real-time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If automatic acoustic environment classification is used, then hearing device settings are adjusted automatically, but misclassification occurs in complex acoustic environments
Solution Approach 1:
The patent introduces a camera as an intermediary device that captures visual information about the acoustic environment. This visual data serves as a mediator between the complex acoustic scene and the classification algorithm, providing additional cues that help accurately identify audio sources and their locations without requiring perfect acoustic classification alone.
Solution Approach 2:
The patent adds a visual dimension to the traditional acoustic-only classification approach. By incorporating camera images that show the spatial location and visual appearance of audio sources, the system transitions from one-dimensional acoustic analysis to multi-dimensional sensing, improving classification reliability in complex environments.
2Adaptability or versatility
If manual adjustment options are provided, then users can customize settings, but the range of modifiers is overwhelming and difficult to navigate
Solution Approach 1:
The system automatically presents only the relevant adjustment options based on the detected acoustic environment and identified audio sources. Instead of requiring users to navigate through all possible modifiers, the system serves itself by intelligently filtering and presenting only the necessary controls for the current situation, simplifying the user interface while maintaining full adaptability.
Solution Approach 2:
The patent applies different levels of interface complexity to different situations. In simple acoustic environments, minimal controls are presented, while in complex environments, more advanced options become available. This local adaptation of interface quality ensures ease of operation in simple cases while maintaining versatility when needed.
3Productivity
If visual identification of audio sources is implemented, then users can quickly select and adjust specific sound inputs, but device complexity increases
Solution Approach 1:
The patent makes the camera serve multiple functions: it captures images for visual identification of audio sources, determines their spatial location, and provides contextual information for classification. This multi-functionality allows the same component to support multiple features (visual selection, automatic classification, spatial awareness) without proportionally increasing device complexity.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
A method for operating a hearing system (10) is provided. The hearing system (10) comprises a hearing device (12) configured to be worn at an ear of a user, a user device (14) communicatively coupled to the hearing device (12) and comprising a camera (36) and a display (30). The hearing device (12) comprises at least one sound input module (20) for generating an audio signal indicative of a sound detected in an environment of the hearing device (12), a first processing unit (40) for modifying the audio signal, and at least one sound output module (22) for outputting the modified audio signal. The method comprises: receiving image data from the camera (36), the image data being representative for a scene (80) in front of the camera (36); receiving an audio signal from the at least one sound input module (20), the audio signal being representative for the acoustic environment of the hearing device (12) substantially at a time the image data have been captured, wherein the acoustic environment comprises at least one audio source and wherein the audio signal is at least in part representative for a sound from the audio source; determining at least one visual object (88) as the audio source, within the scene (80) from the image data and the audio signal.