Speech Interaction Device Selection Based on Acoustic Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In multi-space smart home scenarios, existing speech processing methods fail to efficiently select the most appropriate speech interaction device to process user inputs based on varying speech features such as loudness, sound pressure, or the presence of wake-up words, leading to suboptimal interaction and potential misinterpretation of user instructions.
Innovation Solution
A method and apparatus that acquire speech features from input speech and select the most suitable speech interaction device based on these features, such as selecting devices in descending order of loudness or sound pressure, or waking up specific devices with wake-up words, to ensure accurate processing and execution of user commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple speech interaction devices process the input speech, then the coverage of speech reception is improved, but the accuracy of selecting the most appropriate device decreases
Solution Approach 1:
The patent changes the parameter of device selection from static (any device can process) to dynamic (selection based on speech features). By introducing parameters such as loudness, sound pressure, and wake-up word detection, the system dynamically selects the most appropriate device, resolving the contradiction between broad coverage and accurate selection.
Solution Approach 2:
The system implements feedback by having multiple devices report their speech reception status and features back to a selection mechanism. This feedback loop enables the system to compare speech features from different devices and select the optimal one, maintaining both broad coverage and high selection accuracy.
2Measurement precision
If speech features are analyzed to select the best device, then the accuracy of speech processing is improved, but the complexity of the system increases
Solution Approach 1:
The patent segments the speech processing function into two parts: speech feature analysis (performed by multiple devices) and device selection (performed by a central or distributed selector). This segmentation allows the system to maintain high accuracy while managing complexity through functional division.
Solution Approach 2:
Multiple speech interaction devices are designed with universal functionality to detect and report speech features. This multi-functionality allows the same hardware to serve both as a speech receiver and a feature analyzer, reducing overall system complexity while maintaining high processing accuracy.
3Measurement precision
If devices are selected based on loudness and sound pressure, then the accuracy of capturing user intent is improved, but the time required for feature analysis increases
Solution Approach 1:
The system performs preliminary action by continuously monitoring speech features (loudness, sound pressure) even before a complete speech input is received. This allows the selection process to begin in advance, reducing the time penalty associated with feature analysis while maintaining high accuracy in capturing user intent.
Data Source
AI summary
Embodiments of a method and apparatus for processing a speech are provided. The method can include: acquiring, in response to determining at least one speech interaction device in a target speech interaction device set receiving an input speech, a speech feature of the input speech received by a speech interaction device of the at least one speech interaction device; and selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech. Some embodiments realize the selection of a targeted speech interaction device.


