Voice Agent Processing Device With Local–Cloud Workload Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-based agent systems face challenges in efficiently processing user interactions and providing personalized responses across multiple devices and platforms, particularly in home environments, due to high-load processing demands and the need for seamless integration of voice recognition, semantic analysis, and response synthesis.
Innovation Solution
An information processing device and method that integrates a local agent device with a cloud-based system, utilizing a TV agent and external agent services to manage user interactions, perform voice recognition, semantic analysis, and response synthesis, while ensuring personalized and efficient service delivery through account management and profile customization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If high-load processing including voice recognition and semantic analysis is executed on the agent service side, then processing capability is improved, but system complexity and load increase
Solution Approach 1:
The patent segments the agent system into multiple components: agent devices (local processing units) and agent service (cloud-based processing). Voice recognition, semantic analysis, and information retrieval are divided between local execution on agent devices and cloud-based execution on the agent service, distributing processing load and reducing system complexity at any single point.
Solution Approach 2:
The patent introduces an agent device as an intermediary between the user and the agent service. The agent device handles local voice input, preliminary processing, and coordination with the cloud-based agent service, mediating the interaction and managing the complexity of high-load processing tasks.
2Measurement precision
If voice recognition and semantic analysis are performed on the cloud, then accuracy is improved, but response time increases
Solution Approach 1:
The patent implements preliminary voice recognition and processing on the agent device before transmitting to the cloud-based agent service. This preliminary action prepares and pre-processes voice inputs locally, reducing the complexity and time required for cloud-based semantic analysis while maintaining accurate recognition through the two-stage approach.
3Adaptability or versatility
If multiple agent services are integrated across various devices, then versatility is improved, but system complexity increases
Solution Approach 1:
The patent creates a universal agent service architecture that can be accessed by multiple types of agent devices (television receivers, audio output devices, information processing devices). The agent service provides multi-functional capabilities including voice recognition, semantic analysis, information retrieval, and control operations across diverse devices, standardizing integration and reducing complexity through a common interface.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention provides an information processing device that processes a voice-based agent interaction, and an information processing method, and provides an information processing system. The information processing device is provided with: a communication unit that receives information related to an interaction with a user through an agent residing in a first apparatus; and a control unit that controls an external agent service. The control unit collects the information that includes at least one among an image or a voice of the user, information related to operation of the first apparatus by the user, and sensor information detected by a sensor with which the first apparatus is equipped. The control unit controls calling of the external agent service.