Voice Agent Processing Device With Local–Cloud Workload Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-based agent systems face challenges in efficiently processing user interactions and providing personalized responses across multiple devices and platforms, particularly in home environments, due to high-load processing demands and the need for seamless integration of voice recognition, semantic analysis, and response synthesis.

Innovation Solution

An information processing device and method that integrates a local agent device with a cloud-based system, utilizing a TV agent and external agent services to manage user interactions, perform voice recognition, semantic analysis, and response synthesis, while ensuring personalized and efficient service delivery through account management and profile customization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If high-load processing including voice recognition and semantic analysis is executed on the agent service side, then processing capability is improved, but system complexity and load increase

Engineering Contradiction:
Improveprocessing capabilityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the agent system into multiple components: agent devices (local processing units) and agent service (cloud-based processing). Voice recognition, semantic analysis, and information retrieval are divided between local execution on agent devices and cloud-based execution on the agent service, distributing processing load and reducing system complexity at any single point.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an agent device as an intermediary between the user and the agent service. The agent device handles local voice input, preliminary processing, and coordination with the cloud-based agent service, mediating the interaction and managing the complexity of high-load processing tasks.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If voice recognition and semantic analysis are performed on the cloud, then accuracy is improved, but response time increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidresponse time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary voice recognition and processing on the agent device before transmitting to the cloud-based agent service. This preliminary action prepares and pre-processes voice inputs locally, reducing the complexity and time required for cloud-based semantic analysis while maintaining accurate recognition through the two-stage approach.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple agent services are integrated across various devices, then versatility is improved, but system complexity increases

Engineering Contradiction:
Improveservice integrationVSAvoidintegration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal agent service architecture that can be accessed by multiple types of agent devices (television receivers, audio output devices, information processing devices). The agent service provides multi-functional capabilities including voice recognition, semantic analysis, information retrieval, and control operations across diverse devices, standardizing integration and reducing complexity through a common interface.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3805914B1Information processing device, information processing method, and information processing system
Publication Date: 2025.10.29 SONY GROUP CORP
  • EP3805914B1 patent drawingFigure 1
  • EP3805914B1 patent drawingFigure 2
  • EP3805914B1 patent drawingFigure 3

AI summary

The present invention provides an information processing device that processes a voice-based agent interaction, and an information processing method, and provides an information processing system. The information processing device is provided with: a communication unit that receives information related to an interaction with a user through an agent residing in a first apparatus; and a control unit that controls an external agent service. The control unit collects the information that includes at least one among an image or a voice of the user, information related to operation of the first apparatus by the user, and sensor information detected by a sensor with which the first apparatus is equipped. The control unit controls calling of the external agent service.