Dialog Processing Controller for Distributed Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional client-server type speech recognition systems face limitations in handling vast vocabularies due to device capabilities and network response speed issues, leading to suboptimal execution of user tasks.

Innovation Solution

An information processing device and method that employs a dialog processing controller to distribute tasks based on priority, executing speech recognition and dialog processing between devices and a server to optimize task execution, ensuring that tasks with high priority are executed promptly even under network constraints.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition processing is performed at cloud servers to handle vast vocabularies, then the vocabulary capacity is improved, but the response speed deteriorates due to network communication delays

Engineering Contradiction:
Improvevocabulary capacityVSAvoidresponse speed
Core Design Contradiction:
Adaptability or versatilityVSSpeed

Solution Approach 1:

The system segments speech recognition processing into two parts: device-side processing for quick response to device operation tasks, and cloud server processing for comprehensive information search tasks. This segmentation allows the system to simultaneously achieve fast response for critical operations and vast vocabulary coverage for information searches.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically determines whether to perform speech recognition at the device or cloud server based on the task type and current network conditions. High-priority tasks with limited vocabulary requirements are processed locally, while low-priority tasks benefiting from cloud computing power are processed remotely, optimizing both speed and vocabulary capacity adaptively.

Inventive Principle:
Principle #15Dynamics

2Speed

If speech recognition processing is performed entirely at the device, then the response speed is improved, but the vocabulary capacity deteriorates due to device limitations

Engineering Contradiction:
Improveresponse speedVSAvoidvocabulary capacity
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system segments speech recognition processing into two parts: device-side processing for quick response to device operation tasks, and cloud server processing for comprehensive information search tasks. This segmentation allows the system to simultaneously achieve fast response for critical operations and vast vocabulary coverage for information searches.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically determines whether to perform speech recognition at the device or cloud server based on the task type and current network conditions. High-priority tasks with limited vocabulary requirements are processed locally, while low-priority tasks benefiting from cloud computing power are processed remotely, optimizing both speed and vocabulary capacity adaptively.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If multiple devices access the cloud server simultaneously, then the vocabulary capacity is improved, but the response speed deteriorates due to network traffic congestion

Engineering Contradiction:
Improvevocabulary capacityVSAvoidresponse delay
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system segments speech recognition processing into two parts: device-side processing for quick response to device operation tasks, and cloud server processing for comprehensive information search tasks. This segmentation allows the system to simultaneously achieve fast response for critical operations and vast vocabulary coverage for information searches.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically determines whether to perform speech recognition at the device or cloud server based on the task type and current network conditions. High-priority tasks with limited vocabulary requirements are processed locally, while low-priority tasks benefiting from cloud computing power are processed remotely, optimizing both speed and vocabulary capacity adaptively.

Inventive Principle:
Principle #15Dynamics

4Adaptability or versatility

If device operation tasks are processed at the cloud server, then the vocabulary capacity is improved, but the task execution accuracy deteriorates due to network dependencies

Engineering Contradiction:
Improvevocabulary capacityVSAvoidtask execution accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system segments speech recognition processing into two parts: device-side processing for quick response to device operation tasks, and cloud server processing for comprehensive information search tasks. This segmentation allows the system to simultaneously achieve fast response for critical operations and vast vocabulary coverage for information searches.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically determines whether to perform speech recognition at the device or cloud server based on the task type and current network conditions. High-priority tasks with limited vocabulary requirements are processed locally, while low-priority tasks benefiting from cloud computing power are processed remotely, optimizing both speed and vocabulary capacity adaptively.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10950230B2Information processing device and information processing method
Publication Date: 2021.03.16 PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
  • US10950230B2 patent drawing
  • US10950230B2 patent drawing
  • US10950230B2 patent drawing

AI summary

Included are a speech recognition result obtainer that obtains a speech recognition result, which is text data obtained by speech recognition processing, a priority obtainer that obtains priority corresponding to each of a plurality of tasks that are each identified by a plurality of dialog processing based on the speech recognition result; and a dialog processing controller that causes a plurality of devices to perform the distributed execution of the plurality of dialog processing mutually different from each other. The dialog processing controller provides, based on the priority, control information in accordance with a task identified by the distributed execution to an executer that operates based on the control information.