Wearable Voice Recognition via Bone Conduction Signal Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

User devices with voice recognition capabilities often respond to unintentional or incorrect commands due to their inability to differentiate between intended and unintended speakers, especially in noisy environments, and tend to prioritize the loudest source of speech rather than the specific desired person.

Innovation Solution

A wearable terminal device equipped with a bone conduction sensor and microphones that concurrently record audio and bone conduction signals, calculating a similarity metric to determine if the signals originate from the intended user, allowing for accurate voice command detection and adjustment of audio focusing parameters to enhance signal clarity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the device responds to all spoken questions/commands equally, then it can capture all potential user inputs, but it also responds to unintentional commands or commands from wrong persons

Engineering Contradiction:
Improvevoice command recognition capabilityVSAvoidfalse activation rate
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the voice recognition process into two distinct signal acquisition paths: bone conduction signal path and air-conducted sound signal path. The bone conduction sensor captures signals through the skull bone, while microphones capture air-conducted sounds. By comparing these two segmented signal paths, the system can reliably distinguish between intentional commands (present in both paths) and unintentional sounds (present in only one path), thereby reducing false activations while maintaining versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces bone conduction signals as an intermediary verification mechanism. The bone conduction sensor acts as an intermediary that captures only sounds produced by the wearer's own vocal cords, serving as a trusted reference. By using this intermediary signal to validate air-conducted sound commands, the system ensures that only genuine user intentions are recognized, reducing responses to wrong persons or unintentional commands.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the device prioritizes the loudest source of speech, then it can capture clear audio signals, but it responds to wrong persons instead of the intended user

Engineering Contradiction:
Improveaudio signal clarityVSAvoiduser-specific recognition capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The bone conduction sensor serves as an intermediary that provides user-specific identification. Since bone conduction signals are mechanically coupled to the wearer's skull, they inherently identify the specific user. This intermediary signal allows the system to prioritize the intended user's voice over other loud speakers in the environment, resolving the contradiction between signal clarity and user-specific recognition.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent uses partial action by focusing only on the bone conduction signal component for user identification purposes, rather than analyzing all audio signals equally. By selectively processing the bone conduction path to identify the wearer, the system can ignore loud external speech sources and focus only on commands from the intended user, thus maintaining both signal clarity and user specificity.

Inventive Principle:
Principle #16Partial or excessive action

3Device complexity

If the device uses only air-conducted microphones, then the structure is simple, but it cannot differentiate between intended and unintended speakers in noisy environments

Engineering Contradiction:
Improvesensor configurationVSAvoidspeaker differentiation capability
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the audio sensing function into two independent subsystems: bone conduction sensing (via sensors coupled to the skull) and air-conducted sound sensing (via microphones). This segmentation allows each subsystem to perform its specialized function - bone conduction for user identification and air-conducted sound for command recognition - thereby achieving speaker differentiation capability without requiring an overly complex single-system design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges two different sensing modalities (bone conduction and air-conducted sound) into a unified voice recognition system. By combining these complementary sensing approaches, the system achieves both simple individual component designs and sophisticated overall speaker differentiation capability, resolving the contradiction between device complexity and measurement precision.

Inventive Principle:
Principle #5Merging (Combining)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution enables the wearable terminal device to accurately recognize voice commands from the intended user amidst noise and multiple speakers, reducing false activations and improving voice recognition accuracy.

Implementation Method 1

receiving, via a bone conduction sensor, a bone conduction signal

Methodology Applied
Scientific EffectBone conduction:

Data Source

PatentUS12177623B2Bone conduction confirmation
Publication Date: 2024.12.24 NOKIA TECHNOLOGIES OY
  • US12177623B2 patent drawing
  • US12177623B2 patent drawing
  • US12177623B2 patent drawing

AI summary

According to an aspect, there is provided an apparatus for a wearable terminal device including circuitry configured for performing the following. The apparatus receives, via a bone conduction sensor, a bone conduction signal and, via at least one microphone over the air, an audio signal. The bone conduction signal and the audio signal are, at least in part, substantially concurrently recorded signals. The apparatus calculates a value of a similarity metric for evaluating an extent of similarity between the bone conduction and audio signals. In response to the value exceeding a pre-defined threshold, the apparatus causes performing one or more actions including executing, in response to detecting a voice command, the voice command and/or modifying, if the audio signal is received via a plurality of microphones over the air, one or more audio focusing parameters of the plurality of microphones for increasing the value of the similarity metric.