Audio Token Insertion for Voice AI Coordination

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice AI assistance services lack a technical system to enhance convenience when used in coordination with content, such as television or mobile receivers, and there is a need for improved convenience in this context.

Innovation Solution

An information processing apparatus and method that inserts or detects tokens related to voice AI assistance services within audio streams of content, allowing for better coordination and control of voice recognition processes, including the use of audio watermarks to manage voice recognition process prohibition or service delivery parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice AI assistance service is used in coordination with content, then user convenience is improved, but unintended voice recognition may occur

Engineering Contradiction:
Improveuser convenienceVSAvoidunintended voice recognition
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system inserts tokens into the audio stream before content playback to pre-establish control mechanisms. These tokens are detected by the voice AI system to determine whether voice recognition should be permitted or prohibited during specific content segments, preventing unintended recognition before it occurs.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Tokens serve as intermediary elements between the content audio stream and the voice AI assistance service. The detection unit identifies these tokens to mediate the interaction, allowing the system to control whether voice recognition commands are processed based on the token's indication of permission or prohibition.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If extensive blacklists or whitelists are used to manage voice recognition, then recognition accuracy is improved, but operational costs increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidoperational costs
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system extracts the voice recognition control functionality from complex blacklist/whitelist management. Instead of maintaining extensive lists of permitted or prohibited phrases, the invention uses compact tokens embedded in the audio stream to directly indicate whether recognition should occur, significantly reducing management complexity and operational costs.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the control parameter from maintaining extensive lists of recognized phrases to using simple presence/absence indicators (tokens) in the audio stream. This parameter change transforms the management approach from complex categorical control to simple binary control, reducing operational burden while maintaining recognition accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3683792B1Information processing device and information processing method
Publication Date: 2024.07.03 SONY GROUP CORP
  • EP3683792B1 patent drawingFigure 1
  • EP3683792B1 patent drawingFigure 2
  • EP3683792B1 patent drawingFigure 3

AI summary

The present technique relates to an information processing apparatus and an information processing method that can improve the convenience of a voice AI assistance service used in coordination with content. A first information processing apparatus including an insertion unit that inserts a token, which is related to use of a voice AI assistance service in coordination with content, into an audio stream of the content and a second information processing apparatus including a detection unit that detects the inserted token from the audio stream of the content can be provided to improve the convenience of the voice AI assistance service used in coordination with the content. The present technique can be applied to, for example, a system in coordination with the voice AI assistance service.