Dialog Processing Controller for Distributed Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional client-server type speech recognition systems face limitations in handling vast vocabularies due to device capabilities and network response speed issues, leading to suboptimal execution of user tasks.
Innovation Solution
An information processing device and method that employs a dialog processing controller to distribute tasks based on priority, executing speech recognition and dialog processing between devices and a server to optimize task execution, ensuring that tasks with high priority are executed promptly even under network constraints.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition processing is performed at cloud servers to handle vast vocabularies, then the vocabulary capacity is improved, but the response speed deteriorates due to network communication delays
Solution Approach 1:
The system segments speech recognition processing into two parts: device-side processing for quick response to device operation tasks, and cloud server processing for comprehensive information search tasks. This segmentation allows the system to simultaneously achieve fast response for critical operations and vast vocabulary coverage for information searches.
Solution Approach 2:
The system dynamically determines whether to perform speech recognition at the device or cloud server based on the task type and current network conditions. High-priority tasks with limited vocabulary requirements are processed locally, while low-priority tasks benefiting from cloud computing power are processed remotely, optimizing both speed and vocabulary capacity adaptively.
2Speed
If speech recognition processing is performed entirely at the device, then the response speed is improved, but the vocabulary capacity deteriorates due to device limitations
Solution Approach 1:
The system segments speech recognition processing into two parts: device-side processing for quick response to device operation tasks, and cloud server processing for comprehensive information search tasks. This segmentation allows the system to simultaneously achieve fast response for critical operations and vast vocabulary coverage for information searches.
Solution Approach 2:
The system dynamically determines whether to perform speech recognition at the device or cloud server based on the task type and current network conditions. High-priority tasks with limited vocabulary requirements are processed locally, while low-priority tasks benefiting from cloud computing power are processed remotely, optimizing both speed and vocabulary capacity adaptively.
3Adaptability or versatility
If multiple devices access the cloud server simultaneously, then the vocabulary capacity is improved, but the response speed deteriorates due to network traffic congestion
Solution Approach 1:
The system segments speech recognition processing into two parts: device-side processing for quick response to device operation tasks, and cloud server processing for comprehensive information search tasks. This segmentation allows the system to simultaneously achieve fast response for critical operations and vast vocabulary coverage for information searches.
Solution Approach 2:
The system dynamically determines whether to perform speech recognition at the device or cloud server based on the task type and current network conditions. High-priority tasks with limited vocabulary requirements are processed locally, while low-priority tasks benefiting from cloud computing power are processed remotely, optimizing both speed and vocabulary capacity adaptively.
4Adaptability or versatility
If device operation tasks are processed at the cloud server, then the vocabulary capacity is improved, but the task execution accuracy deteriorates due to network dependencies
Solution Approach 1:
The system segments speech recognition processing into two parts: device-side processing for quick response to device operation tasks, and cloud server processing for comprehensive information search tasks. This segmentation allows the system to simultaneously achieve fast response for critical operations and vast vocabulary coverage for information searches.
Solution Approach 2:
The system dynamically determines whether to perform speech recognition at the device or cloud server based on the task type and current network conditions. High-priority tasks with limited vocabulary requirements are processed locally, while low-priority tasks benefiting from cloud computing power are processed remotely, optimizing both speed and vocabulary capacity adaptively.
Data Source
AI summary
Included are a speech recognition result obtainer that obtains a speech recognition result, which is text data obtained by speech recognition processing, a priority obtainer that obtains priority corresponding to each of a plurality of tasks that are each identified by a plurality of dialog processing based on the speech recognition result; and a dialog processing controller that causes a plurality of devices to perform the distributed execution of the plurality of dialog processing mutually different from each other. The dialog processing controller provides, based on the priority, control information in accordance with a task identified by the distributed execution to an executer that operates based on the control information.


