Own Voice Detection via Air and Bone Conduction Signal Comparison

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in determining whether speech detected by a device, such as a smartphone, was spoken by the user wearing a wearable accessory like earphones, which is crucial for accurate speech recognition and preventing spoof attacks.

Innovation Solution

A method and system utilizing both air-conducted and bone-conducted speech signals, filtered to extract components at the speech articulation rate, compare these components, and determine if the speech was generated by the user based on a threshold difference, enabling accurate own voice detection and preventing spoof attacks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition functionality is enabled on a device, then the device can process and act on spoken commands, but the device cannot distinguish whether the speech was spoken by the user wearing the accessory or by someone else

Engineering Contradiction:
Improvespeech recognition functionalityVSAvoidspeaker identification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the speech detection process into two distinct signal paths: air-conducted speech detection and bone-conducted speech detection. By separating these detection modes, the system can independently analyze each signal type and compare their characteristics to determine whether the speech originates from the user wearing the accessory, thereby resolving the speaker identification accuracy issue while maintaining speech recognition functionality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces bone-conducted speech as an intermediary detection mechanism. The bone conduction sensor acts as a mediator that directly detects vibrations from the user's vocal cords through the skull, providing a reliable indicator of user-generated speech that complements the air-conducted speech detection and enables accurate speaker identification

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If the device accepts any detected speech for processing, then speech recognition can be performed on all input, but unauthorized or spoofed speech can trigger false positive recognition and unauthorized access

Engineering Contradiction:
Improvespeech processing throughputVSAvoidsecurity against spoof attacks
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements preliminary verification by comparing air-conducted and bone-conducted speech signals before the speech recognition process is initiated. This pre-check mechanism filters out non-user speech early in the processing pipeline, ensuring that only speech originating from the user wearing the accessory is passed to the speech recognition engine, thereby preventing spoof attacks while maintaining productivity for legitimate user input

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs a feedback mechanism where the comparison result between air-conducted and bone-conducted signals is used to control whether speech recognition should be performed. The system continuously monitors the consistency between the two signal types and adjusts its acceptance of speech input accordingly, providing real-time security validation that blocks unauthorized access attempts

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Effectively differentiates between user-generated speech and non-user-generated speech, enhancing the reliability of speech recognition systems and preventing unauthorized access by accurately determining if the device is being worn or held by the user.

Implementation Method 1

detecting a second signal representing bone-conducted speech using a bone-conduction sensor of the device

Methodology Applied
Scientific EffectBone conduction: Vibration

Data Source

PatentUS11842725B2Detection of speech
Publication Date: 2023.12.12 CIRRUS LOGIC INC
  • US11842725B2 patent drawing
  • US11842725B2 patent drawing
  • US11842725B2 patent drawing

AI summary

A method of own voice detection is provided for a user of a device. A first signal is detected, representing air-conducted speech using a first microphone of the device. A second signal is detected, representing bone-conducted speech using a bone-conduction sensor of the device. The first signal is filtered to obtain a component of the first signal at a speech articulation rate, and the second signal is filtered to obtain a component of the second signal at the speech articulation rate. The component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate are compared, and it is determined that the speech has not been generated by the user of the device, if a difference between the component of the first signal at the speech articulation rate and the component of the second signal at the speech articulation rate exceeds a threshold value.