Voice Interaction Response Generation via Prosodic Emotion Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice interaction systems fail to generate response sentences that accurately reflect the user's interest and emotion, as they rely on matching predefined expressions, potentially missing emotional content in speech.

Innovation Solution

A response sentence generation apparatus that converts user voice to text, extracts prosodic information to specify emotion occurrence words, and generates response sentences incorporating these words, using conversion, extraction, and generation means to create tailored responses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If expressions are extracted based on matching with predefined special expression lists, then the system can generate responses using simple text analysis, but the system fails to capture emotional content and user interest accurately

Engineering Contradiction:
Improveemotion detection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines text-based expression analysis with prosodic (voice tone, rhythm, stress) analysis to comprehensively detect user emotions. By merging multiple analysis dimensions (text content + voice characteristics), the system achieves more accurate emotion detection than text-only approaches while managing complexity through integrated processing.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system introduces prosodic information as an intermediary element that bridges text analysis and emotion detection. Prosodic features serve as a mediator that captures emotional nuances in voice that are not reflected in text alone, enabling more accurate emotion recognition without requiring direct complex text interpretation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If the system uses predefined expression lists for response generation, then the response generation process is simple and fast, but the responses do not reflect actual user interest or emotional state

Engineering Contradiction:
Improveresponse generation speedVSAvoiduser emotion information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system performs preliminary extraction of prosodic information from user input before response generation. By pre-processing and storing voice characteristic data alongside text analysis results, the system prepares emotion-related information in advance, enabling fast response generation that incorporates emotional context without adding significant processing delay.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the system only analyzes text content without considering prosodic information, then the processing is simpler and faster, but emotional expressions and emphasized content are missed

Engineering Contradiction:
Improveemotion recognition accuracyVSAvoidanalysis complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the analysis process into distinct components: text analysis, prosodic analysis (extracting voice tone, stress, rhythm), and integrated emotion determination. By dividing the complex analysis task into separable modules that can be processed independently and then combined, the system achieves comprehensive emotion recognition while managing analytical complexity through structured segmentation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10861458B2Response sentence generation apparatus, method and program, and voice interaction system
Publication Date: 2020.12.08 TOYOTA JIDOSHA KK
  • US10861458B2 patent drawing
  • US10861458B2 patent drawing
  • US10861458B2 patent drawing

AI summary

A response sentence generation apparatus includes a conversion device for converting an input voice of a user into text information, an extraction device for extracting prosodic information from the input voice, a specifying device for specifying an emotion occurrence word indicating an occurrence of an emotion of the user based on the text information and the prosodic information, and a generation device for selecting a character string including the specified emotion occurrence word from the text information and generating a response sentence by performing predetermined processing on the selected character string.