Multi-modal Activity Detection for AR Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing augmented reality systems face challenges in accurately detecting user speech in noisy environments, leading to false positives and distorted user speech input, which can result in incorrect classification of user activity.

Innovation Solution

A computing device equipped with a suite of sensors, including body-oriented and environment-oriented sensors, processes multi-modal data to accurately classify user activities such as talking, whispering, or shouting, and adjusts audio playback accordingly, using machine learning models to distinguish between user speech and environmental noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional single-modality sensors are used for speech detection, then the device complexity is low, but the speech detection accuracy deteriorates in noisy environments leading to false positives

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidsensor suite complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the speech detection task across multiple sensor modalities (microphone for air conduction, bone conduction transducer for bone conduction, accelerometer for body vibrations). Each sensor captures different aspects of speech, and the neural network integrates these segmented signals to achieve accurate detection in noisy environments without requiring a single complex sensor

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The wearable device incorporates multi-functional sensors that serve multiple purposes: the microphone captures environmental sound, the bone conduction transducer detects bone vibrations, and the accelerometer monitors body movements. This universal sensor suite enables the device to perform speech detection, noise cancellation, and activity recognition simultaneously

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If audio playback is continuously provided, then the user experience for media consumption is improved, but the user's ability to hear environmental noise and converse deteriorates

Engineering Contradiction:
Improveuser experience for media consumptionVSAvoidenvironmental noise awareness
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system continuously monitors speech detection status and provides feedback to control audio playback. When speech is detected, the system automatically pauses or mutes audio playback, allowing the user to hear environmental noise and converse naturally. This feedback loop creates a responsive system that adapts to user needs in real-time

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The audio playback system dynamically adjusts its state based on detected user activity. The playback transitions between playing, pausing, and muting states according to speech detection results, creating a flexible system that prioritizes either media consumption or environmental awareness based on current user needs

Inventive Principle:
Principle #15Dynamics

3Loss of information

If speech detection is performed in crowded and noisy environments, then the device can detect user speech, but the detection accuracy deteriorates leading to distorted or lost user speech input

Engineering Contradiction:
Improveuser speech input qualityVSAvoidenvironmental noise interference
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The system uses bone conduction transducers as an intermediary to capture speech vibrations directly through the skull, bypassing the air conduction path that is contaminated by environmental noise. This intermediary sensing path provides a cleaner signal from the user's voice that can be processed more accurately even in crowded environments

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system combines signals from multiple sensor modalities (air conduction microphone, bone conduction transducer, accelerometer) to create a composite speech signal. This composite approach leverages the strengths of each sensor type while compensating for their individual weaknesses in noisy environments, resulting in more accurate speech recognition

Inventive Principle:
Principle #40Composite materials

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system effectively reduces false positives and improves user experience by accurately identifying user speech and adjusting audio playback in real-time, enhancing interaction with augmented reality devices in noisy conditions.

Implementation Method 1

a first sensor configured to detect body-oriented data from the body of a user wearing the computing device... the sensor may measure body vibrations generated by a user of the computing device

Methodology Applied
Scientific EffectBone conduction:

Implementation Method 2

the sensor may measure body vibrations generated by a user of the computing device while the user moves and speaks

Methodology Applied
Scientific EffectVibration detection: Vibration

Implementation Method 3

one or more second sensors... receive second sensor data... environment-oriented data, comprising air vibration data representing vibrations measured through air

Methodology Applied
Scientific EffectAir conduction: Sound

Data Source

PatentUS11895474B2Activity detection on devices with multi-modal sensing
Publication Date: 2024.02.06 GOOGLE LLC
  • US11895474B2 patent drawing
  • US11895474B2 patent drawing
  • US11895474B2 patent drawing

AI summary

Methods, systems, devices, and computer-readable storage media for activity detection of a user of a computing device, using multi-modal sensing. A device can be configured to receive sensor data corresponding to multiple modalities and process the sensor data to predict an activity performed by a user of a computing device. The device in response to the detected activity can perform a response action, such as muting or pausing audio playback from the computing device. Different modalities can be combined, such as body vibration data, air vibration data, and image data, which can be processed to distinguish user activity, e.g., speaking versus not speaking, to allow the computing device to perform the correct corresponding action.