Voice-Content Pairing for Multimodal Input Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic devices face challenges in seamlessly integrating and processing multimodal interface inputs, such as voice and content inputs associated with the same event, leading to difficulties in providing convenient and comprehensive information retrieval.
Innovation Solution
The device maps and stores voice and content inputs contemporaneously, creating voice-content pairs indexed by time and keyword, allowing for simultaneous recording and playback of both inputs, enabling efficient retrieval and visualization of associated data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the electronic device separately processes recording data and content data from the same event, then the processing complexity is reduced, but the user cannot conveniently identify and access related information from both inputs simultaneously
Solution Approach 1:
The patent merges separate recording data and content data into a unified data structure by creating voice-content pairs that link audio recordings with their corresponding transcribed text and metadata. This integration allows users to access both types of information simultaneously through a single interface, resolving the contradiction between ease of information access and processing complexity.
Solution Approach 2:
The unified data structure serves multiple functions: it stores audio recordings, transcribed text, timestamps, and keyword information in a single organized format. This multi-functional approach enables the system to handle various input types (voice, text, metadata) through a common processing framework, reducing overall system complexity while improving user access convenience.
2Adaptability or versatility
If the electronic device integrates and processes multiple input modes (voice, text, multimodal) simultaneously, then the comprehensiveness of information retrieval is improved, but the system complexity increases
Solution Approach 1:
The patent segments the complex multimodal processing task into distinct components: audio recording, voice-to-text transcription, metadata extraction, and keyword identification. Each component is processed independently and then integrated through the unified data structure, allowing the system to handle multiple input modes comprehensively while managing complexity through modular organization.
Solution Approach 2:
The unified data structure acts as an intermediary layer between various input modalities and the user interface. It standardizes different input types (audio, text, metadata) into a common format with consistent fields (timestamps, content, keywords), enabling comprehensive information retrieval without requiring complex ad-hoc processing for each input type.
3Productivity
If the electronic device stores and manages separate data structures for recording data and content data, then the storage organization is simpler, but the efficiency of retrieving and correlating related information decreases
Solution Approach 1:
The patent combines separate recording data and content data into voice-content pairs within a unified data structure. This merging enables efficient retrieval of related information by allowing the system to access both audio and text content through a single data access point, significantly improving productivity while the structured organization maintains manageable complexity.
Solution Approach 2:
The system performs preliminary actions by pre-processing and organizing data into the unified structure during the recording and transcription process. Timestamps, keywords, and content are extracted and linked in advance, so that when retrieval is needed, the information is already correlated and ready for efficient access without requiring complex real-time processing.
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
An electronic device is provided. The electronic device includes a microphone, a touch screen display, a processor, and memory. The memory stores instructions, when executed, causing the at least one processor to: output a screen where a specified application is executed, the screen including a first interface configured to support recording a voice input through the at least one microphone and a second interface configured to receive content input of a user, on the touch screen display; pairing and storing in the memory each one of a plurality of recorded voice inputs to the first interface with a corresponding one of a plurality of content inputs to the second interface, thereby forming a voice-content pair, wherein the voice input and content input of each voice-content pair are contemporaneously recorded and received; receive a user input based on a keyword for specified content; select at least one recorded voice-content pair in response to the user input; convert recorded voice inputs of the at least one voice-content pair into text data; and output the converted text data on the touch screen display.