AI Virtual Assistant Markup Language Text Output Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional AI virtual assistants struggle to effectively convey emotions through text responses, limiting the emotional bond with users as they often display entire texts simultaneously, which fails to express intended emotions and feelings.

Innovation Solution

An electronic device and method that processes user speech to obtain intent and emotion information, generating a response text with markup language that controls the output unit, such as phoneme, consonant and vowel, syllable, or word units, to enhance emotional expression.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the entire response text is displayed simultaneously on the screen, then the information is conveyed efficiently, but the emotional expression and user engagement are reduced

Engineering Contradiction:
Improveinformation conveyance efficiencyVSAvoiduser engagement and emotional experience
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The response text is segmented into multiple display units (e.g., words, phrases, or sentences) that are displayed sequentially rather than all at once. This segmentation allows the system to control the timing and rhythm of information presentation, creating a more engaging user experience while maintaining efficient communication. The markup language enables fine-grained control over which segments appear when and how they appear.

Inventive Principle:
Principle #1Segmentation

2Productivity

If the text response is displayed without markup language, then the display is simple and fast, but the emotional bond and nuance are lost

Engineering Contradiction:
Improvedisplay speedVSAvoidemotional expression capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The markup language introduces additional parameters for controlling text display characteristics, such as display timing, pacing, emphasis, and sequencing. These parameters can be adjusted based on the intended emotion or message nuance, allowing the same text to convey different emotional tones through controlled parameter variations rather than changing the text itself.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If the AI assistant uses standard text responses, then the system is simple and efficient, but the emotional bond with users is weak

Engineering Contradiction:
Improvesystem simplicityVSAvoidemotional bonding capability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The markup language acts as an intermediary layer between the simple text response generation and the rich emotional expression requirements. Rather than complicating the core text generation system, the markup language provides a separate mechanism for controlling presentation style, allowing emotional nuance to be added without changing the fundamental simplicity of the AI assistant's response generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11922127B2Method for outputting text in artificial intelligence virtual assistant service and electronic device for supporting the same
Publication Date: 2024.03.05 SAMSUNG ELECTRONICS CO LTD
  • US11922127B2 patent drawing
  • US11922127B2 patent drawing
  • US11922127B2 patent drawing

AI summary

According to an embodiment, an electronic device comprises: a memory, a communication module comprising communication circuitry, and a processor operatively connected with the memory and the communication module. The processor is configured to control the electronic device to: obtain a utterance text corresponding to utterance speech, obtain an intent of the utterance text and emotion information based on the utterance speech and the utterance text, obtain a response text for the utterance text based on the intent of the utterance text and the emotion information, obtain a markup language including information about an output unit of text of the response text based on at least one of the intent of the utterance text, the emotion information, or the response text, and add the markup language to the response text and provide the response text. The text output unit is at least one selected from among a phoneme unit, a consonant and vowel unit, a syllable unit, or a word unit. A text output method in an artificial intelligence virtual assistant service of an electronic device may be performed using an artificial intelligence model.