Multi-modal Activity Detection for AR Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality systems face challenges in accurately detecting user speech in noisy environments, leading to false positives and distorted user speech input, which can result in incorrect classification of user activity.
Innovation Solution
A computing device equipped with a suite of sensors, including body-oriented and environment-oriented sensors, processes multi-modal data to accurately classify user activities such as talking, whispering, or shouting, and adjusts audio playback accordingly, using machine learning models to distinguish between user speech and environmental noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional single-modality sensors are used for speech detection, then the device complexity is low, but the speech detection accuracy deteriorates in noisy environments leading to false positives
Solution Approach 1:
The system segments the speech detection task across multiple sensor modalities (microphone for air conduction, bone conduction transducer for bone conduction, accelerometer for body vibrations). Each sensor captures different aspects of speech, and the neural network integrates these segmented signals to achieve accurate detection in noisy environments without requiring a single complex sensor
Solution Approach 2:
The wearable device incorporates multi-functional sensors that serve multiple purposes: the microphone captures environmental sound, the bone conduction transducer detects bone vibrations, and the accelerometer monitors body movements. This universal sensor suite enables the device to perform speech detection, noise cancellation, and activity recognition simultaneously
2Ease of operation
If audio playback is continuously provided, then the user experience for media consumption is improved, but the user's ability to hear environmental noise and converse deteriorates
Solution Approach 1:
The system continuously monitors speech detection status and provides feedback to control audio playback. When speech is detected, the system automatically pauses or mutes audio playback, allowing the user to hear environmental noise and converse naturally. This feedback loop creates a responsive system that adapts to user needs in real-time
Solution Approach 2:
The audio playback system dynamically adjusts its state based on detected user activity. The playback transitions between playing, pausing, and muting states according to speech detection results, creating a flexible system that prioritizes either media consumption or environmental awareness based on current user needs
3Loss of information
If speech detection is performed in crowded and noisy environments, then the device can detect user speech, but the detection accuracy deteriorates leading to distorted or lost user speech input
Solution Approach 1:
The system uses bone conduction transducers as an intermediary to capture speech vibrations directly through the skull, bypassing the air conduction path that is contaminated by environmental noise. This intermediary sensing path provides a cleaner signal from the user's voice that can be processed more accurately even in crowded environments
Solution Approach 2:
The system combines signals from multiple sensor modalities (air conduction microphone, bone conduction transducer, accelerometer) to create a composite speech signal. This composite approach leverages the strengths of each sensor type while compensating for their individual weaknesses in noisy environments, resulting in more accurate speech recognition
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system effectively reduces false positives and improves user experience by accurately identifying user speech and adjusting audio playback in real-time, enhancing interaction with augmented reality devices in noisy conditions.
Implementation Method 1
a first sensor configured to detect body-oriented data from the body of a user wearing the computing device... the sensor may measure body vibrations generated by a user of the computing device
Implementation Method 2
the sensor may measure body vibrations generated by a user of the computing device while the user moves and speaks
Implementation Method 3
one or more second sensors... receive second sensor data... environment-oriented data, comprising air vibration data representing vibrations measured through air
Data Source
AI summary
Methods, systems, devices, and computer-readable storage media for activity detection of a user of a computing device, using multi-modal sensing. A device can be configured to receive sensor data corresponding to multiple modalities and process the sensor data to predict an activity performed by a user of a computing device. The device in response to the detected activity can perform a response action, such as muting or pausing audio playback from the computing device. Different modalities can be combined, such as body vibration data, air vibration data, and image data, which can be processed to distinguish user activity, e.g., speaking versus not speaking, to allow the computing device to perform the correct corresponding action.


