Dilated Convolutional Neural Network for Keyword Spotting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional keyword spotting methods require significant memory and computational resources, making them unsuitable for low-resource devices and failing to provide real-time, high-accuracy results.

Innovation Solution

A dilated convolutional neural network (DCNN) is implemented on a low-power device for keyword detection in continuous audio streams, utilizing gated activation units and skip-connections to reduce computational requirements while maintaining accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional keyword spotting methods are used, then keyword detection accuracy can be maintained, but memory resources and computational resources are significantly consumed

Engineering Contradiction:
Improvekeyword detection accuracyVSAvoidmemory resources
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the keyword spotting task into multiple dilation layers, each processing different temporal contexts with varying dilation rates. This divides the computational workload across specialized layers, reducing the memory footprint required for any single layer while maintaining overall detection accuracy through hierarchical feature extraction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dilation layers that operate in the temporal dimension by expanding the receptive field without increasing the number of parameters proportionally. This dimensional approach allows the model to capture long-range dependencies in audio signals while keeping memory consumption manageable through efficient parameter sharing across time steps.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If traditional keyword spotting methods are used, then keyword detection can be performed, but computational resources are significantly consumed making them unsuitable for low-resource devices

Engineering Contradiction:
Improvekeyword detection capabilityVSAvoidcomputational resources
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The computational task is segmented into multiple dilation layers with increasing dilation rates, allowing progressive feature extraction. Each layer processes information at different temporal scales, reducing the computational burden on any single layer and enabling efficient processing on low-power devices while maintaining detection capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The dilation layers perform preliminary feature extraction and filtering before final keyword detection. By pre-processing the audio signal through multiple dilation stages that capture temporal patterns at different scales, the system reduces the computational complexity of the final detection step, making it suitable for resource-constrained devices.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If traditional keyword spotting methods are used, then detection can be performed, but real-time response cannot be provided

Engineering Contradiction:
Improvedetection capabilityVSAvoidreal-time response
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The detection process is segmented into parallel dilation layers that can process different temporal contexts simultaneously. This segmentation enables real-time processing by distributing the computational load across multiple layers that operate in parallel, rather than sequentially processing the entire signal through a single deep network.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The dilation layers dynamically adjust the receptive field size through varying dilation rates, allowing the model to adaptively focus on different temporal scales. This dynamic approach enables real-time response by efficiently capturing relevant temporal patterns without requiring processing of the entire historical signal, thus improving productivity while maintaining detection reliability.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250174225A1Dilated convolutions and gating for efficient keyword spotting
Publication Date: 2025.05.29 SNIPS
  • US20250174225A1 patent drawing
  • US20250174225A1 patent drawing

AI summary

A method for detection of a keyword in a continuous stream of audio signal, by using a dilated convolutional neural network, implemented by one or more computers embedded on a device, the dilated convolutional network comprising a plurality of dilation layers, including an input layer and an output layer, each layer of the plurality of dilation layers comprising gated activation units, and skip-connections to the output layer, the dilated convolutional network being configured to generate an output detection signal when a predetermined keyword is present in the continuous stream of audio signal, the generation of the output detection signal being based on a sequence of successive measurements provided to the input layer, each successive measurement of the sequence being measured on a corresponding frame from a sequence of successive frames extracted from the continuous stream of audio signal, at a plurality of successive time steps.