Nonverbal Message Extraction With Context-Aware Text Annotation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer technologies are incapable of automatically extracting nonverbal messages from unstructured corpora, such as facial expressions and body postures, which are crucial for tasks like dialogue understanding and generation, due to the lack of public datasets and resource-consuming data collection methods.
Innovation Solution
A method and apparatus using a machine learning model to receive text, extract nonverbal messages, and annotate them, involving kinesics, internal states, and vocal types, with context analysis and minimization of a corruption objective to determine the messages, and employing encoder-decoder models for multi-span extraction and generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If machine learning models are used to extract nonverbal messages from unstructured text, then extraction capability is improved, but model complexity and training requirements increase
Solution Approach 1:
The patent introduces a pre-trained language model as an intermediary component that processes unstructured text and extracts nonverbal messages. This mediator model serves as a bridge between raw text input and structured nonverbal message output, enabling automatic extraction without requiring complex custom processing systems. The pre-trained model leverages existing linguistic knowledge to handle the extraction task efficiently.
Solution Approach 2:
The system performs preliminary actions by pre-training the language model on extensive corpora before actual extraction tasks. This pre-training phase equips the model with fundamental language understanding and pattern recognition capabilities, allowing it to extract nonverbal messages from unstructured text without requiring complex real-time processing during actual use cases.
2Ease of manufacture
If nonverbal messages are extracted from small-scale well-structured corpora, then extraction ease is improved, but applicability to real-world unstructured data deteriorates
Solution Approach 1:
The patent creates a universal extraction system using a pre-trained language model that can handle both well-structured and unstructured text corpora. The model's general language understanding capabilities allow it to adapt to various text formats and structures, making it applicable to real-world data while maintaining ease of extraction. This multi-functional approach eliminates the need for separate extraction systems for different data types.
Solution Approach 2:
The system adapts to different data structures by changing its processing parameters dynamically. The pre-trained model adjusts its attention mechanisms and prediction patterns based on the input text structure, whether well-structured or unstructured. This parameter adaptation allows the same extraction system to maintain high performance across diverse corpora without requiring manual reconfiguration.
3Speed
If existing heuristics are used to extract NMs from scripts, then extraction speed is improved, but extraction accuracy and completeness deteriorate
Solution Approach 1:
The patent replaces mechanical heuristic rules with a neural network-based language model for extraction. This substitution eliminates the rigid, rule-based approach that achieves speed but sacrifices accuracy, and replaces it with a probabilistic model that can handle ambiguity and nuance in text. The neural model processes text through multiple layers of transformation, achieving both speed and high accuracy simultaneously.
Solution Approach 2:
The extraction system incorporates feedback mechanisms where the model continuously refines its predictions based on contextual information. The pre-trained model uses attention mechanisms to focus on relevant portions of the text and adjusts its extraction decisions based on surrounding context, ensuring high accuracy while maintaining processing speed through efficient attention computation.
Data Source
AI summary
A method and apparatus comprising computer code configured to cause a processor or processors to receive a text comprising a plurality of sentences, by a machine learning model, extract a nonverbal message from one of the sentences and add an annotation to the text, the annotation indicating the nonverbal message, and output a version of the text including the annotation.


