Optical Microphone Array for Low-Noise Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems on mobile devices face challenges due to high computational intensity and sensitivity to background noise, which is exacerbated by the limitations of conventional microphones, including high self-noise and size constraints that affect signal-to-noise ratio.

Innovation Solution

An optical microphone array with closely spaced microphones and a dual-processor system, where the first processor identifies speech presence and issues a wake-up signal to activate a more powerful second processor for advanced speech recognition, leveraging low self-noise and small size benefits of optical microphones to enhance signal processing and reduce power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Volume of moving object

If conventional MEMS condenser microphones are used, then the device can be made small, but the self-noise increases and signal-to-noise ratio deteriorates

Engineering Contradiction:
Improvemicrophone sizeVSAvoidself-noise
Core Design Contradiction:
Volume of moving objectVSObject-generated harmful factors

Solution Approach 1:

The patent replaces the mechanical/electrical condenser microphone system with an optical microphone system. The optical microphone uses a light source, membrane, and photodetector to convert sound waves into electrical signals through optical means rather than electrical capacitance measurement. This substitution eliminates the high self-noise inherent in small MEMS condenser microphones while maintaining small form factor, as the optical system does not suffer from the same impedance and noise limitations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If multiple microphones are used to improve speech recognition performance, then noise rejection improves, but device complexity and hardware accommodation become problematic

Engineering Contradiction:
Improvespeech recognition performanceVSAvoidhardware complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple microphone functions into a single integrated optical microphone array system. Rather than using multiple separate conventional microphones that would increase hardware complexity and require multiple preamplifier circuits, the system employs multiple optical microphone elements that can be closely spaced and integrated on a single substrate, reducing overall hardware complexity while maintaining the noise rejection benefits of multiple sensors.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The optical microphone array serves multiple functions simultaneously: it provides speech detection, noise rejection through beamforming, wake-word recognition, and full speech recognition capabilities. The dual-processor architecture enables the same hardware to handle both low-power wake-word detection and high-performance speech recognition, making the system universally applicable to various speech processing tasks without requiring separate dedicated hardware for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If powerful processors are used for speech recognition, then recognition accuracy improves, but power consumption increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the speech recognition processing into two distinct stages handled by separate processors: a first processor performs initial wake-word detection on continuously captured audio, and only when a wake-word is detected does the system activate a second, more powerful processor for full speech recognition processing. This segmentation allows the system to maintain high recognition accuracy when needed while minimizing power consumption during idle periods by keeping the powerful processor in sleep mode.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system employs periodic action by continuously monitoring audio input with a low-power wake-word detector and only activating the high-power speech recognition processor periodically when speech is actually detected. This approach transforms the continuous high-power operation into periodic high-power bursts, dramatically reducing average power consumption while maintaining the capability for accurate speech recognition when required.

Inventive Principle:
Principle #19Periodic action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables efficient and accurate speech recognition with improved signal-to-noise ratio and power management, allowing for implementation in small form factor devices like smartphones and smart watches with reduced battery impact.

Implementation Method 1

an array of optical microphones on a substrate, wherein the optical microphones are arranged at a mutual spacing of less than 5 mm, each of said optical microphones providing a signal indicative of displacement of a respective membrane as a result of an incoming audible sound

Methodology Applied
Scientific EffectOptical detection of membrane displacement: Optical Tweezers

Data Source

PatentEP3281200B1Speech recognition
Publication Date: 2020.12.16 SINTEF TTO AS
  • EP3281200B1 patent drawingFigure 1
  • EP3281200B1 patent drawingFigure 2
  • EP3281200B1 patent drawingFigure 3~4

AI summary

An optical microphone arrangement comprises: an array of optical microphones (4) on a substrate (8), each of said optical microphones FIG. 2 (4) providing a signal indicative of displacement of a respective membrane (24) as a result of an incoming audible sound; at first processor (12) arranged to receive said signals from said optical microphones (4) and to perform a first processing step on said signals to produce a first output; and a second processor (14) arranged to receive at least one of said signals or said first output; wherein at least said second processor (14) determines presence of at least one element of human speech from said audible sound.