Full-Duplex Voice Dialogue Duration Matching for Topic Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Full-duplex voice dialogue systems struggle with accurately determining user topics due to network instability and varying speech speeds, leading to incorrect interactions and poor user experience.
Innovation Solution
A method and system where a voice dialogue terminal records and uploads audio to a cloud server, which determines a reply content and duration, and the terminal presents the content only when the durations match, discarding mismatched content to ensure accurate responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the system interacts in full-duplex mode with instant response, then the interaction freedom is improved, but the topic consistency deteriorates
Solution Approach 1:
The patent introduces a feedback mechanism where the terminal compares the duration of audio uploaded to the cloud server with the duration of the returned reply content. This feedback loop ensures that only replies corresponding to the actual user input are presented, maintaining topic consistency while preserving full-duplex interaction freedom.
Solution Approach 2:
The system dynamically adjusts its response behavior based on duration matching. When durations match, the reply is presented; when they don't match, the system waits for updated content. This dynamic adjustment resolves the contradiction between instant response and topic consistency.
2Speed
If the system processes audio in real-time, then the response speed is improved, but the measurement precision deteriorates
Solution Approach 1:
The terminal performs preliminary action by uploading audio to the cloud server and obtaining reply content with duration information before presenting it to the user. This preliminary processing allows real-time response while ensuring accurate topic identification through duration verification.
Data Source
Figure 1~2
Figure 3
Figure 4~5
AI summary
Disclosed is a full-duplex voice dialogue method applied to a voice dialogue terminal and including recording and uploading by an awakened voice dialogue terminal audio to a cloud server for determining a reply content and a first duration of the audio analyzed for determining the reply content; receiving by the voice dialogue terminal the reply content and the first duration sent by the cloud server; determining whether the first duration is equal to a duration from the moment awakening the voice dialogue terminal to the current moment of uploading the audio; and presenting the reply content to a user if consistent. Both the reply content determined by the cloud server and the duration of the audio is acquired, and the reply content is presented to the user only when the first duration and the second duration are determined as consistent, thereby ensuring proper reply content.