Voice System Utterance Segmentation for Partial Repetition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice recognition systems fail to accurately respond to user requests for reuttering specific parts of system utterances, leading to unnecessary repetition and incorrect information being provided when users ask for clarification or specific details.

Innovation Solution

An information processing device and method that determine the type of user utterance, such as requesting all or part of the previous system utterance, to generate appropriate responses, including reuttering all or a portion of the previous system utterance based on the user's request.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the system reutters all system utterances from the beginning when a user requests reutterance, then the user can hear the complete information again, but the user has to repeatedly hear unnecessary information, resulting in wasted time

Engineering Contradiction:
Improvecompleteness of reutterance informationVSAvoidtime wasted hearing unnecessary information
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the system utterance into multiple parts or sections, allowing the system to identify and reutter only the specific portion that the user wants to hear again, rather than reuttering the entire utterance from the beginning. This segmentation enables precise control over what information is reuttered, eliminating unnecessary repetition while maintaining completeness of relevant information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts or identifies the specific portion of the system utterance that the user wants to reutter by analyzing the user's request. The system takes out only the relevant segment from the complete system utterance and reutters that extracted portion, avoiding the waste of time from hearing unnecessary information while ensuring the complete relevant information is provided.

Inventive Principle:
Principle #2Taking out (Extraction)

2Ease of manufacture

If the system uses predefined keywords to detect reutterance requests, then the system can process specific user utterances, but the system cannot respond to cases where these specific words are not included

Engineering Contradiction:
Improvesimplicity of detection mechanismVSAvoidability to respond to various user utterances
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent changes the detection parameter from relying on specific predefined keywords to analyzing the temporal relationship and semantic content between the user utterance and the system utterance. By changing how the system detects reutterance requests - focusing on the relationship between utterances rather than specific words - the system gains versatility to respond to various user utterances while maintaining a relatively simple detection mechanism based on timing and content analysis.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If the system answers questions using a general document set, then the system can provide information, but the system may output an answer different from what the user wanted to hear

Engineering Contradiction:
Improveability to answer various questionsVSAvoidaccuracy of matching user intent
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by focusing the information retrieval on the specific local context of the system utterance that the user wants to reutter, rather than searching through a general document set. The system retrieves information locally from the recent system utterance with matching semantic content, ensuring high accuracy in matching user intent while maintaining the ability to answer various questions about the system's own output.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12002460B2Information processing device, information processing system, and information processing method, and program
Publication Date: 2024.06.04 SONY GROUP CORP
  • US12002460B2 patent drawing
  • US12002460B2 patent drawing
  • US12002460B2 patent drawing

AI summary

A device and a method that determine an utterance type of a user utterance and generate a system response according to a determination result are achieved. A user utterance type determination unit that determines an utterance type of a user utterance, and a system response generation unit that generates a system response according to a type determination result determined by the user utterance type determination unit are included. The user utterance type determination unit determines whether the user utterance is of type A that requests all reutterances of a system utterance immediately before the user utterance, or type B that requests a reutterance of a part of the system utterance immediately before the user utterance. The system response generation unit generates a system response to reutter all system utterances immediately before the user utterance in a case where the user utterance is of type A, and generates a system response to reutter a part of system utterances immediately before the user utterance in a case where the user utterance is of type B.