Multimodal Tag Creation by Discarding Redundant Training Signal Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Portable communication devices face limitations in navigation, data entry, and input due to their small size, and existing input modalities like voice tags, while natural, do not fully leverage other available input modalities such as motion, vibration, and gesture sensing.
Innovation Solution
The method and apparatus enable multimodal tags by combining training signals from different input modalities like audio, tactile, motion, and gesture signals to create a unified input mechanism, associating specific functions with these combinations, and using pattern matching techniques like Hidden Markov Models or Dynamic Time Warping to recognize and execute user-defined actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If multiple input modalities (audio, tactile, motion, gesture) are combined to create multimodal tags, then the user interface becomes more intuitive and convenient, but the device complexity increases
Solution Approach 1:
The patent combines multiple input modalities (audio, tactile, motion, gesture) into unified multimodal tags by merging their respective training signals. This integration allows the system to process diverse input types through a common framework, improving user interface convenience while managing complexity through systematic combination rather than separate handling of each modality
Solution Approach 2:
The multimodal tag system serves multiple functions simultaneously: it processes voice commands, motion gestures, tactile inputs, and camera-based gestures through a single unified mechanism. This multi-functionality allows the device to handle various input types without requiring separate processing systems for each, thereby improving ease of operation while controlling device complexity
2Adaptability or versatility
If training signals from multiple modalities are processed and stored, then the system can recognize more user actions, but the information processing and storage requirements increase
Solution Approach 1:
The patent extracts essential features from training signals across multiple modalities while discarding redundant information. By focusing on extracting only the necessary characteristics needed for action recognition, the system achieves high adaptability in recognizing diverse user actions without proportionally increasing information processing and storage requirements
Solution Approach 2:
The system discards redundant information present in training signals from multiple modalities while retaining the essential features needed for recognition. This selective discarding reduces the information burden during training and operation, allowing the system to maintain high versatility in action recognition without proportionally increasing storage and processing demands
Data Source
AI summary
A method and apparatus for enabling multimodal tags in a communication device is disclosed. The method comprises receiving a first training signal and receiving a second training signal in conjunction with the first training signal. A multimodal tag is created by discarding redundant or non-discriminative information associated with each of the first and second training signals to represent a combination of the first training signal and the second training signal and a function is associated with the created multimodal tag.


