On-Device Sound Detection Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio classification models struggle to detect and classify individual variations of general sound categories, particularly in home environments, due to the challenge of obtaining sufficient training data and the variability of sounds from devices like doorbells, appliances, and pets.

Innovation Solution

Electronic devices are equipped with on-device sound detection models trained using local audio samples, employing a two-stage inference architecture with a trigger model and a detection model, allowing for customizable sound classification and notification systems without relying on extensive off-device data collection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If general audio classification models trained on vast datasets are used, then general sound category classification is improved, but detection of individual variations of specific sounds deteriorates

Engineering Contradiction:
Improvesound detection accuracyVSAvoidindividual sound variation detection
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the sound detection task into two distinct stages: a first pass using a general audio classification model for broad category identification, and a second pass using a specialized detection model for individual sound variation detection. This segmentation allows each model to be optimized for its specific function, resolving the contradiction between general classification accuracy and individual variation detection capability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism where the first pass model acts as a filter that pre-processes audio inputs before they reach the second pass detection model. This intermediary stage reduces the computational burden on the specialized detection model while ensuring that only relevant audio segments are analyzed in detail, thereby improving both efficiency and detection accuracy for individual sound variations

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If extensive off-device data collection is performed, then training data quantity is improved, but data privacy concerns and user burden worsen

Engineering Contradiction:
Improvetraining data quantityVSAvoiddata privacy concerns
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent implements self-service by enabling the detection model to be trained directly on the user's device using locally captured audio samples. This eliminates the need to upload sensitive audio data to external servers, allowing the system to accumulate sufficient training data while maintaining data privacy and reducing user burden. The model learns from individual variations in the user's specific environment without requiring extensive off-device data collection

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies preliminary action by collecting and processing training data locally on the device before the actual deployment of the detection model. This preliminary local data collection and model training process ensures that sufficient training examples are available while maintaining privacy, as the data never leaves the device. The system prepares everything in advance locally, eliminating the need for subsequent data uploads

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12182674B2Sound detection for electronic devices
Publication Date: 2024.12.31 APPLE INC
  • US12182674B2 patent drawing
  • US12182674B2 patent drawing
  • US12182674B2 patent drawing

AI summary

The subject disclosure provides systems and methods for providing locally trained models for detecting individual sounds using electronic devices. Local detection of individual sounds with a detection model at an electronic device can be provided by obtaining training samples for the detection model with the electronic device, and generating additional negative and positive training samples based on the obtained training samples. A two-stage detection process may be provided, in which a trigger model at a device compares an audio input to a reference sound to trigger a detection model at the device. The detection of individual sounds with a detection model at an electronic device can also leverage audio capture capabilities of multiple devices in an acoustic scene to capture multiple concurrent training samples.