Speech Interaction Device Selection Based on Acoustic Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-space smart home scenarios, existing speech processing methods fail to efficiently select the most appropriate speech interaction device to process user inputs based on varying speech features such as loudness, sound pressure, or the presence of wake-up words, leading to suboptimal interaction and potential misinterpretation of user instructions.

Innovation Solution

A method and apparatus that acquire speech features from input speech and select the most suitable speech interaction device based on these features, such as selecting devices in descending order of loudness or sound pressure, or waking up specific devices with wake-up words, to ensure accurate processing and execution of user commands.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple speech interaction devices process the input speech, then the coverage of speech reception is improved, but the accuracy of selecting the most appropriate device decreases

Engineering Contradiction:
Improvecoverage of speech receptionVSAvoidaccuracy of device selection
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter of device selection from static (any device can process) to dynamic (selection based on speech features). By introducing parameters such as loudness, sound pressure, and wake-up word detection, the system dynamically selects the most appropriate device, resolving the contradiction between broad coverage and accurate selection.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system implements feedback by having multiple devices report their speech reception status and features back to a selection mechanism. This feedback loop enables the system to compare speech features from different devices and select the optimal one, maintaining both broad coverage and high selection accuracy.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If speech features are analyzed to select the best device, then the accuracy of speech processing is improved, but the complexity of the system increases

Engineering Contradiction:
Improveaccuracy of speech processingVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the speech processing function into two parts: speech feature analysis (performed by multiple devices) and device selection (performed by a central or distributed selector). This segmentation allows the system to maintain high accuracy while managing complexity through functional division.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Multiple speech interaction devices are designed with universal functionality to detect and report speech features. This multi-functionality allows the same hardware to serve both as a speech receiver and a feature analyzer, reducing overall system complexity while maintaining high processing accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If devices are selected based on loudness and sound pressure, then the accuracy of capturing user intent is improved, but the time required for feature analysis increases

Engineering Contradiction:
Improveaccuracy of user intent captureVSAvoidtime for feature analysis
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by continuously monitoring speech features (loudness, sound pressure) even before a complete speech input is received. This allows the selection process to begin in advance, reducing the time penalty associated with feature analysis while maintaining high accuracy in capturing user intent.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11244686B2Method and apparatus for processing speech
Publication Date: 2022.02.08 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11244686B2 patent drawing
  • US11244686B2 patent drawing
  • US11244686B2 patent drawing

AI summary

Embodiments of a method and apparatus for processing a speech are provided. The method can include: acquiring, in response to determining at least one speech interaction device in a target speech interaction device set receiving an input speech, a speech feature of the input speech received by a speech interaction device of the at least one speech interaction device; and selecting, based on the speech feature of the input speech received by the speech interaction device in the at least one speech interaction device, a first speech interaction device from the at least one speech interaction device to process the input speech. Some embodiments realize the selection of a targeted speech interaction device.