Voice Interaction Device Response Accuracy via Confidence Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In voice interaction technologies, there is a challenge in accurately distinguishing between human-machine interaction sounds and non-human-machine interaction sounds, particularly in one-wakeup-successive-interaction scenarios, leading to potential misapplication of recognized information and a decrease in user experience due to incorrect device responses.

Innovation Solution

An interactive voice-control method and apparatus that determine interaction confidence and matching status based on acoustic and semantic feature representations of sound signals, using machine learning models to differentiate between speech intended for interaction and non-interaction sounds, thereby improving response accuracy and user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition is performed on all collected sound signals, then the device can respond to more potential interactions, but the device may incorrectly respond to non-interaction sounds leading to decreased user experience

Engineering Contradiction:
Improveresponse coverageVSAvoidresponse accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an interaction confidence determination module as an intermediary between speech recognition and device response. This module evaluates acoustic features, semantic features, and matching status to generate an interaction confidence score, acting as a mediator that filters sound signals before triggering device responses, thus resolving the contradiction between comprehensive response coverage and response accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by using the determined interaction confidence and matching status to control whether the device responds to the sound signal. This feedback mechanism allows the system to learn from recognition results and adjust its response behavior, improving reliability while maintaining adaptability through iterative optimization

Inventive Principle:
Principle #23Feedback

2Speed

If the device responds to all recognized information, then interaction responsiveness is improved, but misinterpretation of non-interaction sounds increases

Engineering Contradiction:
Improveinteraction responsivenessVSAvoidsound signal discrimination accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent changes the parameter of sound signal evaluation by introducing multiple dimensions: acoustic feature representation, semantic feature representation, and matching status. These parameter changes enable more precise discrimination of sound signals while maintaining fast interaction responsiveness through efficient feature extraction and comparison

Inventive Principle:
Principle #35Parameter changes

3Productivity

If traditional speech recognition is used without interaction confidence assessment, then processing speed is faster, but discrimination between interaction and non-interaction sounds is less accurate

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidinteraction discrimination accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the sound signal processing into distinct components: acoustic feature extraction, semantic feature extraction, and matching status determination. This segmentation allows each component to be optimized independently, maintaining processing efficiency while improving interaction discrimination accuracy through specialized feature analysis

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11503155B2Interactive voice-control method and apparatus, device and medium
Publication Date: 2022.11.15 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11503155B2 patent drawing
  • US11503155B2 patent drawing
  • US11503155B2 patent drawing

AI summary

The present disclosure discloses an interactive voice-control method and apparatus, a device and a medium. The method includes: obtaining a sound signal at a voice interaction device and recognized information that is recognized from the sound signal; determining an interaction confidence of the sound signal based at least on at least one of an acoustic feature representation of the sound signal and a semantic feature representation associated with the recognized information; determining a matching status between the recognized information and the sound signal; and providing the interaction confidence and the matching status for controlling a response of the voice interaction device to the sound signal.