Conversational Processing Apparatus for Multi-Modal Referential Expression Handling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional conversational processing systems in multi-modal environments are limited by their inability to utilize non-verbal information such as motions, gestures, and facial expressions, requiring explicit indications for referential expressions, which hinders real-time and daily-life interactions.
Innovation Solution
A conversational processing method and apparatus that extracts referential expressions from input sentences, generates intermediate expressions to represent modifying relations among words, and searches for corresponding objects within a preconfigured range, using monadic and dyadic relation tables to provide accurate information and adjust the search range based on accuracy, allowing for efficient use without additional indications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional referential expression processing is used, then the system can process explicit referential expressions, but it cannot utilize non-verbal information such as motions, gestures, and facial expressions
Solution Approach 1:
The conversational processing system is designed to handle multiple types of inputs including verbal expressions, motions, gestures, and facial expressions through a unified processing framework. The system extracts referential expressions from diverse input modalities and processes them through common intermediate representation and object search mechanisms, enabling the system to function across multiple input types without requiring separate processing paths for each modality.
2Measurement precision
If explicit indications of referential expressions are required, then processing accuracy is maintained, but real-time conversational processing becomes difficult
Solution Approach 1:
The system pre-configures object search ranges and prepares intermediate representation templates before actual conversational processing occurs. By having the object search range and processing framework ready in advance, the system can quickly process referential expressions as they appear in real-time conversation without requiring extensive analysis or configuration during the interaction itself, thus maintaining both accuracy and real-time performance.
3Measurement precision
If the object search range is expanded to cover more possibilities, then object identification accuracy improves, but processing complexity increases
Solution Approach 1:
The system applies different processing strategies to different regions of the search space. Rather than uniformly searching all possible objects, the system identifies and focuses computational resources on local regions of the object search range that are most relevant to the current conversational context. This allows the system to maintain high object identification accuracy while avoiding the computational complexity of exhaustively searching the entire object space.
Data Source
AI summary
Disclosed are a method for processing a dialog based on processing instructing expression in a multi-modal environment and an apparatus therefor. The method for processing a dialog in an information processing device capable of processing digital signals includes the steps of: extracting an instructing expression from an inputted sentence; generating an intermediate instructing expression representing the modifying relations between the words constituting the extracted instructing expression; and searching the object corresponding with the intermediate instructing expression in a predetermined object search range. Thus, a terminal can be effectively and conveniently used without separately clarifying various instructing expressions representing things or objects with the terminal.


