Speech Intention Estimation via Breakpoint Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech dialogue systems struggle to accurately estimate the intention of users, particularly when complex sentences are spoken, as they fail to correctly divide expressions into units of user intention, leading to inaccurate interpretation.

Innovation Solution

An information processing device and method that detect breakpoints in user speech through recognition results, allowing for semantic analysis of divided speech sentences to estimate user intention more accurately, using a combination of voice, image, and sensor recognition to identify pauses, intonation boundaries, and other speech properties.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech sentences are divided using language grammar, then sentence structure can be analyzed, but various expressions including user intentions fail to be correctly divided into intention units

Engineering Contradiction:
Improveintention estimation accuracyVSAvoidintention boundary information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments speech sentences into multiple divided speech sentences based on detected breakpoints during speech recognition. This segmentation allows the system to analyze each segment's intention separately and combine results, improving overall intention estimation accuracy while preserving intention boundary information that would be lost in holistic analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses feedback from speech recognition results to dynamically detect breakpoints and adjust speech sentence division. The recognition results provide feedback about what has been understood so far, allowing the system to identify appropriate breakpoints where intention boundaries likely occur, thereby maintaining information integrity while enabling structured analysis.

Inventive Principle:
Principle #23Feedback

2Reliability

If semantic analysis is performed on entire long sentences, then comprehensive meaning can be captured, but user intentions embedded in complex sentences fail to be accurately estimated

Engineering Contradiction:
Improveintention estimation reliabilityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides complex speech sentences into multiple smaller divided speech sentences at detected breakpoints before performing semantic analysis. This segmentation reduces processing complexity by breaking down large analytical tasks into manageable segments while improving reliability by enabling precise intention detection in each segment, which is then aggregated to form the complete intention understanding.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of performing complete semantic analysis on the entire speech sentence at once, the system performs partial semantic analysis on individual divided speech sentences. This partial action approach reduces computational complexity while maintaining or improving intention estimation reliability through cumulative analysis of segments.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11335334B2Information processing device and information processing method
Publication Date: 2022.05.17 SONY GROUP CORP
  • US11335334B2 patent drawing
  • US11335334B2 patent drawing
  • US11335334B2 patent drawing

AI summary

There is provided an information processing device and an information processing method that enable the intention of a speech of a user to be estimated more accurately. The information processing device includes: a detection unit configured to detect a breakpoint of a speech of a user on the basis of a result of recognition that is to be obtained during the speech of the user; and an estimation unit configured to estimate an intention of the speech of the user on the basis of a result of semantic analysis of a divided speech sentence obtained by dividing a speech sentence at the detected breakpoint of the speech. The present technology can be applied, for example, to a speech dialogue system.