Multi-User Voice Response Management via Dynamic Occupation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice UI systems are designed for one-to-one conversations and struggle to efficiently manage multiple users, leading to system occupation and difficulty in responding to multiple users simultaneously.
Innovation Solution
An information processing device with a response generation unit, decision unit, and output control unit that prioritizes and outputs responses to multiple users based on speech order, using a combination of voice and display responses to manage concurrent conversations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the voice UI system is used in a house or public space with multiple users, then the system can serve more users, but a certain user is likely to occupy the system, preventing others from using it
Solution Approach 1:
The patent implements dynamic response occupation management where the system automatically transitions voice response occupation between users based on speech detection. When a new user speaks, the system detects the speech, determines the user is different from the current occupant, and transitions occupation to the new user. This dynamic adjustment allows multiple users to access the system sequentially without manual intervention, resolving the contradiction between multi-user capability and system accessibility.
2Ease of operation
If the system responds to multiple users simultaneously, then user convenience is improved, but the system cannot efficiently manage multiple conversations at once
Solution Approach 1:
The patent segments the voice response occupation into distinct user-specific occupations. Each user can occupy the voice response channel independently, and the system manages these occupations separately through user identification and transition control. This segmentation allows the system to handle multiple users without creating complex intertwined conversation management, as each user's interaction is treated as a separate, manageable unit.
3Reliability
If the system uses only voice output for responses, then the response quality is high, but the system cannot respond to multiple users at the same time
Solution Approach 1:
The patent implements multi-functionality by allowing the system to output responses through multiple channels: voice output and display output. The display can show text responses simultaneously while voice responds to another user, enabling the system to maintain high response quality through voice while increasing overall response throughput by utilizing display as an additional output function. This multi-functional approach resolves the contradiction between response quality and response throughput.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
To provide an information processing device, control method, and program that can improve convenience of a speech recognition system by outputting appropriate responses to respective users when the plurality of users are talking. The information processing device includes: a response generation unit configured to generate responses to speeches from a plurality of users; a decision unit configured to decide methods of outputting the responses to the respective users on the basis of priorities according to order of the speeches from the plurality of users; and an output control unit configured to perform control such that the generated responses are output by using the decided methods of outputting the responses.