Multi-Sensor Voice Control for High-Frequency Voiceprint Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voiceprint recognition using a bone vibration sensor loses high-frequency components, leading to inaccurate recognition due to the sensor's limited frequency capture range, which is typically less than 1 kHz.
Innovation Solution
Employing a combination of in-ear, out-of-ear, and bone vibration sensors to capture multiple voice components, with feature extraction and dynamic fusion coefficients to enhance voiceprint recognition accuracy and robustness, particularly compensating for high-frequency signal loss.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a bone vibration sensor is used to capture voice signals, then the sensor can detect vibrations through bone conduction, but high-frequency components (above 1 kHz) are lost, leading to inaccurate voiceprint recognition
Solution Approach 1:
The patent combines multiple voice sensors (in-ear voice sensor, out-of-ear voice sensor, and bone vibration sensor) to capture voice signals through different pathways. The in-ear and out-of-ear sensors capture air-conducted sound across the full frequency range, while the bone vibration sensor captures bone-conducted vibrations. By merging the outputs of these sensors, the system recovers high-frequency components that would be lost if only bone conduction were used, thereby improving voiceprint recognition accuracy.
Solution Approach 2:
The wearable device is designed to perform multiple voice capture functions using different sensors. The in-ear voice sensor captures sound through the ear canal, the out-of-ear voice sensor captures ambient sound, and the bone vibration sensor captures bone-conducted vibrations. This multi-functional approach allows the system to adapt to different acoustic environments and compensate for the limitations of individual sensors, particularly recovering high-frequency information.
2Measurement precision
If multiple voice sensors are used to capture different voice components, then voiceprint recognition accuracy is improved, but device complexity increases
Solution Approach 1:
The patent integrates multiple voice sensors (in-ear voice sensor, out-of-ear voice sensor, and bone vibration sensor) into a single wearable device. These sensors work together to capture different voice components, with their outputs being processed and fused to improve voiceprint recognition accuracy while managing the complexity through unified device architecture.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Improves voiceprint recognition accuracy and robustness by leveraging multiple sensors and dynamic fusion coefficients, enhancing user experience through efficient and accurate identity authentication.
Implementation Method 1
After the user wears a wearable device, an external auditory canal and a middle auditory canal form a closed cavity, and there is amplification effect, that is, cavity effect, for a sound in the cavity. Therefore, a sound captured by the in-ear voice sensor is clearer, and especially, there is obvious enhancement effect for a high-frequency sound signal.
Implementation Method 2
A bone vibration sensor is a common voice sensor. When a sound is transmitted through a bone, the bone vibrates. The bone vibration sensor senses vibration of the bone, and converts a vibration signal into an electrical signal to capture the sound.
Data Source
AI summary
This application provides a voice control method and apparatus, a wearable device, and a terminal. The method includes: obtaining voice information of a user; obtaining identity information of the user based on a first voiceprint recognition result of a first voice component of the voice information, a second voiceprint recognition result of a second voice component of the voice information, and a third voiceprint recognition result of a third voice component of the voice information, where the first voice component is captured by an in-ear voice sensor of a wearable device, the second voice component is captured by an out-of-ear voice sensor of the wearable device, and the third voice component is captured by a bone vibration sensor of the wearable device; and executing an operation instruction when the identity information of the user matches the preset identity information.


