Voice Interaction Mediator for Cross-Agent Service Coordination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice interaction agent systems struggle to enable cooperation between similar services provided by different voice interaction agents, leading to difficulties in controlling these agents to perform unified tasks, such as playing consecutive songs across different providers.

Innovation Solution

An information processing system comprising a first device that acquires and transfers user voice inputs to multiple voice interaction agents, converting control signals to match each agent's capabilities, allowing for coordinated recognition and response across multiple servers, including a main VPA server and sub VPA servers, to facilitate seamless integration and control of services.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If similar services are provided independently by different voice interaction agents, then each agent can function autonomously, but the agents cannot cooperate with each other to perform unified tasks

Engineering Contradiction:
Improveservice independenceVSAvoidcooperation capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces a first device (mediator) that receives user voice inputs and intelligently routes them to appropriate voice interaction agents (second device or third device). This intermediary coordinates between multiple independent agents, enabling cooperation without compromising their autonomy. The mediator converts control signals between different agents, allowing them to work together on unified tasks while maintaining their independent service capabilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If multiple voice interaction agents operate independently, then system complexity is reduced, but control coordination between agents becomes difficult

Engineering Contradiction:
Improvesystem structureVSAvoidcontrol coordination
Core Design Contradiction:
Device complexityVSEase of operation

Solution Approach 1:

The first device serves multiple functions: it acts as a voice input interface, a routing controller, a signal converter, and a coordination hub. By giving the first device universal functionality, the system avoids creating separate complex control structures for each agent, thereby reducing overall system complexity while improving control coordination ease.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If each voice interaction agent processes voice inputs independently, then service autonomy is maintained, but unified task execution across agents fails

Engineering Contradiction:
Improveservice autonomyVSAvoidunified task execution
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent merges the control and coordination functions of multiple independent agents into a unified system managed by the first device. While the second device and third device maintain their service autonomy for voice recognition and response generation, the first device combines their capabilities to execute unified tasks. For example, when a user requests music playback, the system can coordinate between agents to play songs across different services seamlessly.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11646034B2Information processing system, information processing apparatus, and computer readable recording medium
Publication Date: 2023.05.09 TOYOTA JIDOSHA KK
  • US11646034B2 patent drawing
  • US11646034B2 patent drawing
  • US11646034B2 patent drawing

AI summary

An information processing system includes: a first device configured to acquire a user's uttered voice, transfer the user's uttered voice to at least one of a second and a third devices each actualizing a voice interaction agent, when a control command is acquired, convert a control signal based on the acquired control command to a control signal that matches the second device, and transmit the converted control signal to the second device; a second device configured to recognize the uttered voice transferred from the first device, and output, to the first device, a control command regarding a recognition result obtained by recognizing the uttered voice and response data based on the control signal; and a third device configured to recognize the uttered voice transferred from the first device, and output, to the first device, a control command regarding a recognition result obtained by recognizing the uttered voice.