Acoustic Event Detection Using Natural Language Profiles

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current acoustic-event detection systems struggle to effectively identify custom acoustic events without pre-defined audio samples, relying on globally trained components that may not accurately detect unique sounds specific to individual environments or user-defined events.

Innovation Solution

A system that utilizes natural language descriptions to refine user-provided inputs, employing knowledge graphs and conversational mechanisms to identify audio samples corresponding to custom acoustic events, allowing for the detection of specific sounds like a particular dog bark or appliance beep by integrating text and audio embeddings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If globally trained acoustic event detection components are used, then the system can detect general acoustic events, but it cannot accurately detect custom or unique sounds specific to individual environments

Engineering Contradiction:
Improvedetection accuracy for custom soundsVSAvoidsystem configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by collecting audio samples during an enrollment phase before actual detection begins. Users provide sample recordings of custom sounds they want to detect, and the system processes these samples to create detection profiles. This preliminary data collection and processing enables the system to later detect these specific custom sounds accurately without requiring global retraining.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of audio data by generating synthetic training samples from user-provided sound recordings. These copies are used to train or fine-tune detection models specific to each user's custom sounds. The system can generate multiple variations and augmentations of the original sound samples, enabling accurate detection without needing extensive real-world data collection.

Inventive Principle:
Principle #26Copying

2Reliability

If the system uses pre-defined audio samples for training, then detection models can be trained, but it requires extensive pre-training data that may not be available for custom events

Engineering Contradiction:
Improvedetection reliabilityVSAvoidamount of training data required
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system enables self-service by allowing users to directly provide their own audio samples for training. Instead of requiring the system to collect and process large amounts of external training data, users themselves supply the relevant sound recordings from their own environments. The system then uses these user-provided samples to create detection models tailored to their specific needs, eliminating the dependency on extensive external pre-training datasets.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system applies parameter changes by transforming a small number of user-provided audio samples into multiple training variations through data augmentation techniques. By modifying parameters such as pitch, timing, volume, and adding background noise, the system generates diverse training data from limited original samples. This enables reliable model training with minimal initial data requirements.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If the system allows custom acoustic event detection, then it can recognize unique sounds, but it increases processing complexity and resource requirements

Engineering Contradiction:
Improvecustom event detection capabilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system applies segmentation by dividing the acoustic event detection task into separate specialized components: a general-purpose detector for common sounds and custom detectors for user-defined sounds. Each detector is optimized for its specific target events. This segmentation allows the system to handle custom detection requests efficiently without overwhelming the entire processing pipeline, reducing overall complexity while maintaining versatility.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12087320B1Acoustic event detection
Publication Date: 2024.09.10 AMAZON TECH INC
  • US12087320B1 patent drawing
  • US12087320B1 patent drawing
  • US12087320B1 patent drawing

AI summary

A system may be configured to detect custom acoustic events, where the system generates an acoustic event profile for the custom acoustic event based on a natural language description provided by a user and using an audio sample of the described acoustic event. For example, the user may describe the custom acoustic event as “dog bark.” The system may ask the user questions to refine the description (e.g., dog breed, dog gender, age, etc.). Using an audio sample of the refined description, the system may then determine that audio captured in the user's environment is a potential sample of the custom acoustic event. Such captured audio may be presented to the user for confirmation, and then may be used to detect future occurrences of the custom acoustic event in the user's environment.