Spiking Neural Network Auditory Source Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in effectively separating mixed audio signals into their individual auditory sources, particularly in monaural, unsupervised, and online auditory source segregation for applications like speech enhancement and activity detection.

Innovation Solution

A spiking neural network-based method that selects an audio attribute, represents it as a spiking event, and determines coincidence with a single source by processing audio signals through a network of neurons with synaptic connections and memristor elements, utilizing spike-timing-dependent plasticity for adaptive learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional computational techniques are used for auditory source separation, then the system can process audio signals, but the complexity of the task and data makes the design of the function burdensome and impractical

Engineering Contradiction:
Improveability to infer function from observationsVSAvoidcomplexity of function design
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The spiking neural network performs unsupervised learning, allowing the system to automatically infer source separation functions from audio observations without requiring manual function design or training data. The network self-organizes to separate mixed audio sources based on temporal coherence principles, eliminating the need for complex pre-programmed separation algorithms.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If conventional audio processing methods are used, then processing can be performed, but it is cumbersome and inadequate for monaural unsupervised source segregation

Engineering Contradiction:
Improveease of source segregationVSAvoidprocessing complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces conventional mechanical/audio processing methods with a biologically-inspired spiking neural network that operates on temporal coincidence detection. Instead of using traditional signal processing algorithms, the system uses networks of spiking neurons that naturally perform source separation through their temporal response properties, simplifying the operational complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If prior knowledge of target source is required, then source separation can be achieved, but it limits applicability to unsupervised scenarios

Engineering Contradiction:
Improveapplicability to unsupervised scenariosVSAvoidaccuracy of source identification
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The spiking neural network is pre-configured with temporal coincidence detection capabilities and connectivity structures that enable it to automatically perform source separation when presented with mixed audio signals. This preliminary structural configuration allows the network to achieve accurate source identification without requiring prior knowledge or training on specific target sources.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9269045B2Auditory source separation in a spiking neural network
Publication Date: 2016.02.23 QUALCOMM INC
  • US9269045B2 patent drawing
  • US9269045B2 patent drawing
  • US9269045B2 patent drawing

AI summary

A method of audio source segregation includes selecting an audio attribute of an audio signal. The method also includes representing a portion of the audio attribute that is dominated by a single source as a source spiking event. In addition, the method includes representing a remaining portion of the audio signal as an audio signal spiking event. The method further includes determining whether the remaining portion coincides with the single source based on coincidence of the source spiking event and audio signal spiking event.