Adaptive Interjection Timing for Voice Response Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition terminal devices do not account for increased server response time due to complex user utterances, leading to excessively long interjection-to-response times and user discomfort.
Innovation Solution
An information processing device with a processor that acquires and recognizes user voice data, determines the timing of an interjection based on the time from voice data acquisition to response output, and outputs the interjection and response accordingly, delaying the interjection as needed to minimize the interjection-to-response time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If the system uses a fixed bridge word timing regardless of utterance complexity, then the system structure remains simple, but the interjection-to-response time becomes excessively long for complex utterances causing user discomfort
Solution Approach 1:
The patent applies dynamics by making the interjection timing adaptive rather than fixed. The determination unit dynamically adjusts the interjection timing based on the predicted response time, which varies with utterance complexity. This allows the system to optimize the interjection-to-response time for each specific utterance while maintaining a relatively simple overall system structure.
Solution Approach 2:
The patent changes the timing parameter of the interjection based on the predicted response time. Instead of using a constant bridge word duration, the system adjusts the interjection timing parameter according to the complexity of the user's utterance and the expected server response time, thereby reducing user discomfort for complex utterances.
2Ease of operation
If the interjection is output immediately after voice acquisition, then the dialogue feels responsive, but the interjection-to-response time becomes too long when server processing takes time
Solution Approach 1:
The patent applies preliminary action by outputting the interjection before the response is fully prepared, but timing it based on the predicted response time. This allows the system to maintain dialogue responsiveness by having the interjection appear early, while avoiding excessively long interjection-to-response times by pre-calculating the optimal timing based on server processing expectations.
Solution Approach 2:
The system uses feedback from the predicted response time determination to adjust the interjection timing. The determination unit predicts how long the server will take to process the utterance and uses this feedback to set the appropriate interjection timing, ensuring the interjection-to-response time remains acceptable while maintaining overall dialogue responsiveness.
Data Source
AI summary
An information processing device includes a processor configured to acquire voice data of a voice uttered by a user, recognize the acquired voice, determine a timing of an interjection in accordance with time from completion of the voice data acquisition to start of output of a response generated based on a result of the voice recognition, output the interjection at the determined timing of the interjection, and output the response at the time of start of the output of the response.


