Acoustic Activity Recognition via Sound Effect Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current devices with microphones, such as smart speakers and smartwatches, are unable to recognize activities in their environment due to the lack of leveraging audio insights, relying heavily on large, domain-specific training datasets and requiring in-situ training, limiting their ability to classify sound events effectively in real-world applications.
Innovation Solution
A method and system that utilize a sound effects database to generate an augmented set of sound effects, projected into various audio domains, creating a training dataset for a model that recognizes features in audio signals without the need for in-situ training, allowing for ubiquitous acoustic activity sensing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large, domain-specific training datasets are used to train machine learning models for audio classification, then the model accuracy for specific sound events is improved, but the time required for training and the complexity of the system increases significantly
Solution Approach 1:
The patent applies preliminary action by pre-training a universal audio model on a large, diverse dataset of sound effects before deployment. This pre-trained model serves as a foundation that can be quickly adapted to specific domains without requiring lengthy retraining, thus resolving the contradiction between achieving high accuracy and minimizing training time.
Solution Approach 2:
The patent segments the training process into two stages: a general pre-training stage on diverse sound effects, and a specific fine-tuning stage for target domains. This segmentation allows the system to benefit from large datasets for general accuracy while keeping specific domain training time minimal, addressing the contradiction between precision and time loss.
2Adaptability or versatility
If in-situ training is performed in the user's environment, then the model adapts to specific acoustic contexts, but the requirement for large quantities of context-specific training data makes broad activity recognition unachievable
Solution Approach 1:
The patent implements universality by creating a universal audio recognition model that can operate across multiple domains and contexts without requiring domain-specific retraining. The model is trained on a diverse, universal dataset that encompasses various acoustic environments, enabling it to adapt to different contexts using a single, compact model rather than requiring large quantities of context-specific data.
Solution Approach 2:
The patent uses a pre-trained universal model as an intermediary between the raw audio input and the specific recognition task. This intermediary model provides a foundation that reduces the need for large context-specific training data, allowing the system to achieve broad adaptability without requiring substantial domain-specific datasets.
3Reliability
If constrained sets of recognized classes are used with domain-specific training data, then the model performance for those specific classes is improved, but the model cannot recognize a rich set of activities across different domains
Solution Approach 1:
The patent applies universality by training a single model on a diverse, multi-domain dataset that includes various sound effects from different contexts. This universal model maintains high reliability for specific classes while simultaneously expanding the range of recognizable activities across domains, eliminating the need to choose between specialized performance and broad versatility.
Data Source
AI summary
Embodiments are provided to recognize features and activities from an audio signal. In one embodiment, a model is generated from sound effect data, which is augmented and projected into an audio domain to form a training dataset efficiently. Sound effect data is data that has been artificially created or from enhanced sounds or sound processes to provide a more accurate baseline of sound data than traditional training data. The sound effect data is augmented to create multiple variants to broaden the sound effect data. The augmented sound effects are projected into various audio domains, such as indoor, outdoor, urban, based on mixing background sounds consistent with these audio domains. The model is installed on any computing device, such as a laptop, smartphone, or other device. Features and activities from an audio signal are then recognized by the computing device based on the model without the need for in-situ training.


