Multi-Microphone Sound Identification via Power Level Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sound identification systems face challenges in distinguishing sound from a source of interest from background noise, particularly in low power or low cost applications, due to complexity and resource-intensive computations, and reliance on statistical models or heuristics developed through machine learning or template matching.

Innovation Solution

A sound processing system utilizing multiple microphones, where one microphone is closer to the source of interest, processes audio feeds by time synchronizing and filtering frequencies between them to identify sound originating from the point of interest, thereby reducing processing and power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If statistical models or heuristics developed through machine learning or template matching are used to distinguish sound from source of interest, then sound identification accuracy is improved, but system complexity and computational resource requirements increase

Engineering Contradiction:
Improvesound identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the complex statistical models and machine learning components from the voice activity detection system. Instead of using sophisticated algorithms, the invention employs simple signal processing techniques such as power level difference ratio calculations and basic filtering operations that eliminate the need for resource-intensive computational models while maintaining effective noise suppression.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces expensive, complex computational models with simple, computationally inexpensive operations. The system uses basic arithmetic operations (power level differences, ratios) and straightforward filtering techniques that require minimal processing resources, making the system suitable for low-power and low-cost applications without relying on sophisticated statistical models.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Measurement precision

If statistical models or heuristics developed through machine learning or template matching are used to distinguish sound from source of interest, then sound identification accuracy is improved, but computational resource consumption increases

Engineering Contradiction:
Improvesound identification accuracyVSAvoidcomputational resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts and removes the complex statistical models and machine learning components from the voice activity detection system. Instead of using sophisticated algorithms, the invention employs simple signal processing techniques such as power level difference ratio calculations and basic filtering operations that eliminate the need for resource-intensive computational models while maintaining effective noise suppression.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces expensive, complex computational models with simple, computationally inexpensive operations. The system uses basic arithmetic operations (power level differences, ratios) and straightforward filtering techniques that require minimal processing resources, making the system suitable for low-power and low-cost applications without relying on sophisticated statistical models.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Measurement precision

If complex mechanisms are used to distinguish sound from source of interest, then sound identification capability is improved, but processing time and computational load increase

Engineering Contradiction:
Improvesound identification capabilityVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent extracts and removes the complex statistical models and machine learning components from the voice activity detection system. Instead of using sophisticated algorithms, the invention employs simple signal processing techniques such as power level difference ratio calculations and basic filtering operations that eliminate the need for resource-intensive computational models while maintaining effective noise suppression.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the approach from complex statistical parameter analysis to simple power level parameter comparisons. By focusing on fundamental acoustic parameters (power levels, frequency differences) rather than complex statistical features, the system achieves rapid processing with minimal computational overhead while maintaining effective voice activity detection capability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3360137B1Identifying sound from a source of interest based on multiple audio feeds
Publication Date: 2019.07.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3360137B1 patent drawingFigure 1
  • EP3360137B1 patent drawingFigure 2A~2C
  • EP3360137B1 patent drawingFigure 3A~4

AI summary

Methods and systems for identifying sound from a source of interest are provided for herein. In some embodiments, a first audio feed is captured by a first microphone and a second audio feed is captured by a second microphone. The first microphone may be located closer in proximity to the source of interest than the second microphone. The first audio feed can be processed utilizing the second audio feed to produce a first processed audio feed that can enable identification of sound originating from the source of interest. In some embodiments, the second audio feed can be additionally processed utilizing the first audio feed to produce a second processed audio feed. In such embodiments, frequencies from the first processed audio feed can be compared against frequencies of the second processed audio feed to identify sound originating from the source of interest. Other embodiments may be described and/or claimed herein.