Audio Event Recognition Model Distribution for Smart Home Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current smart home devices face challenges in accurately recognizing and providing event designations for a wide range of sounds, especially those that occur simultaneously or are specific to individual users, due to the lack of distinctive audio data for events like door openings, animal presence, and traffic, which limits their effectiveness in diverse backgrounds and milieus.
Innovation Solution
A method and system where a processing node determines models associating audio data with event designations, allowing communication devices to identify events based on recorded sounds, using algorithms like principal component analysis and combined models, and shares these models across devices to enhance event recognition accuracy and scope.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If audio data is continuously recorded and compared to representative audio data for event recognition, then event designation can be provided, but accuracy is limited for non-distinctive sounds and simultaneous events
Solution Approach 1:
The patent segments the event recognition process into multiple stages: initial representative audio data comparison, followed by secondary audio data analysis when confidence is insufficient. This segmentation allows the system to handle both distinctive and non-distinctive sounds effectively, improving accuracy for diverse events without requiring all events to meet the same recognition threshold.
Solution Approach 2:
The patent introduces confidence scores as an intermediary mechanism between audio data comparison and event designation. The confidence score evaluates the quality of match between recorded audio and representative audio data, allowing the system to determine when additional analysis is needed. This intermediary enables accurate handling of both distinctive sounds (high confidence) and non-distinctive sounds (lower confidence requiring further processing).
2Measurement precision
If computing-intensive operations are performed in communication devices for model determination, then event recognition accuracy improves, but device complexity and energy consumption increase
Solution Approach 1:
The patent extracts the computationally intensive model determination operations from the communication device and relocates them to a remote server. The communication device retains only the lightweight functions of recording audio data, comparing with representative data, and transmitting results. This extraction maintains high event recognition accuracy through sophisticated modeling while significantly reducing device complexity and energy consumption.
Solution Approach 2:
The patent introduces a server as an intermediary between the communication device and the model determination process. The server receives audio data from multiple devices, performs comprehensive analysis to determine or update models, and distributes updated models back to devices. This intermediary architecture enables complex computations to be performed centrally while keeping individual communication devices simple and energy-efficient.
3Measurement precision
If representative audio data is obtained for each specific event, then event designation accuracy improves, but difficulty increases for obtaining data for non-distinctive sounds
Solution Approach 1:
The patent merges audio data from multiple communication devices to build representative audio data for events. Instead of requiring each device to independently capture and store representative audio for every possible event, the system combines audio recordings from multiple sources. This merging approach enables the system to obtain comprehensive representative data for non-distinctive sounds like door openings or animal presence by aggregating data across different devices and locations.
Solution Approach 2:
The patent creates a universal model determination system that serves multiple communication devices simultaneously. The server performs model determination for all devices, enabling each device to benefit from collectively gathered audio data. This universal approach allows the system to handle diverse events including non-distinctive sounds, as the aggregated data from multiple devices provides sufficient representation for events that would be difficult for a single device to capture independently.
Data Source
AI summary
A method performed by a processing node (10), comprising the steps of: i. obtaining (11), from at least one communication device (100), audio data (12) associated with a sound and storing (13) the audio data (12) in the processing node (10), ii. Obtaining (15) an event designation (16) associated with the sound and storing (17) the event designation (16) in the processing node (10), iii. determining (19) a model (20) which associates the audio data (12) with the event designation (16) and storing the model (21), and iv. Providing (23) the model (20) to the communication device (100). A method performed by the communication device (100), as well as a processing node (10), a communication device (100), a system (1000) and computer programs for performing the methods are also described.


