Conversational Processing Apparatus for Multi-Modal Referential Expression Handling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional conversational processing systems in multi-modal environments are limited by their inability to utilize non-verbal information such as motions, gestures, and facial expressions, requiring explicit indications for referential expressions, which hinders real-time and daily-life interactions.

Innovation Solution

A conversational processing method and apparatus that extracts referential expressions from input sentences, generates intermediate expressions to represent modifying relations among words, and searches for corresponding objects within a preconfigured range, using monadic and dyadic relation tables to provide accurate information and adjust the search range based on accuracy, allowing for efficient use without additional indications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional referential expression processing is used, then the system can process explicit referential expressions, but it cannot utilize non-verbal information such as motions, gestures, and facial expressions

Engineering Contradiction:
Improveability to process diverse input typesVSAvoidnon-verbal information utilization
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The conversational processing system is designed to handle multiple types of inputs including verbal expressions, motions, gestures, and facial expressions through a unified processing framework. The system extracts referential expressions from diverse input modalities and processes them through common intermediate representation and object search mechanisms, enabling the system to function across multiple input types without requiring separate processing paths for each modality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If explicit indications of referential expressions are required, then processing accuracy is maintained, but real-time conversational processing becomes difficult

Engineering Contradiction:
Improvereferential expression identification accuracyVSAvoidprocessing time for real-time interaction
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system pre-configures object search ranges and prepares intermediate representation templates before actual conversational processing occurs. By having the object search range and processing framework ready in advance, the system can quickly process referential expressions as they appear in real-time conversation without requiring extensive analysis or configuration during the interaction itself, thus maintaining both accuracy and real-time performance.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the object search range is expanded to cover more possibilities, then object identification accuracy improves, but processing complexity increases

Engineering Contradiction:
Improveobject identification accuracyVSAvoidsearch space complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies different processing strategies to different regions of the search space. Rather than uniformly searching all possible objects, the system identifies and focuses computational resources on local regions of the object search range that are most relevant to the current conversational context. This allows the system to maintain high object identification accuracy while avoiding the computational complexity of exhaustively searching the entire object space.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS9966069B2Method for processing dialogue based on processing instructing expression and apparatus therefor
Publication Date: 2018.05.08 POSTECH ACADEMY INDUSTRY FOUNDATION
  • US9966069B2 patent drawing
  • US9966069B2 patent drawing
  • US9966069B2 patent drawing

AI summary

Disclosed are a method for processing a dialog based on processing instructing expression in a multi-modal environment and an apparatus therefor. The method for processing a dialog in an information processing device capable of processing digital signals includes the steps of: extracting an instructing expression from an inputted sentence; generating an intermediate instructing expression representing the modifying relations between the words constituting the extracted instructing expression; and searching the object corresponding with the intermediate instructing expression in a predetermined object search range. Thus, a terminal can be effectively and conveniently used without separately clarifying various instructing expressions representing things or objects with the terminal.