Text Voice Synchronization in Meeting Recording Terminals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for recording meeting content, such as shooting PowerPoint presentations, often result in asynchronous text and voice content, leading to missed information and lack of synchronization between visual and audio recordings.
Innovation Solution
A text and voice information processing method that recognizes text information in target pictures, extracts key information, and maps corresponding voice files to specific coordinate positions based on word frequency thresholds, allowing for synchronized display and playback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If voice recording is performed separately after PPT shooting, then audio can be recorded, but text and voice content become asynchronous and cannot be synchronized
Solution Approach 1:
The patent merges the PPT shooting function with voice recording into a single integrated operation. When the user shoots a PPT page, the system automatically records the ambient voice at that moment, creating a direct temporal and functional link between the visual content and audio content. This eliminates the need for separate recording operations and ensures automatic synchronization between text (from PPT recognition) and voice content.
Solution Approach 2:
The system performs preliminary voice recording during the PPT shooting process itself, rather than waiting for post-processing. By capturing voice data at the moment each PPT page is photographed, the system establishes the synchronization relationship in advance, allowing text and voice to be paired correctly from the beginning without requiring later alignment operations.
2Reliability
If all text content is processed with voice mapping, then complete synchronization is achieved, but power consumption and storage requirements increase
Solution Approach 1:
Instead of uniformly processing all text content, the system applies voice mapping selectively based on local characteristics. It identifies key information regions in the PPT (such as titles, bullet points, or highlighted text) and performs voice mapping only for those specific areas. This localized approach maintains synchronization accuracy for important content while reducing overall processing load and power consumption compared to processing every text element.
Solution Approach 2:
The system performs partial voice mapping rather than complete mapping of all text content. By focusing on mapping voice to key information regions rather than every text element, it achieves sufficient synchronization for the most important content while avoiding the excessive power consumption and storage requirements that would result from processing the entire text corpus.
3Reliability
If voice files are mapped to every text keyword, then complete synchronization is achieved, but storage space requirements increase
Solution Approach 1:
The system applies voice mapping to specific local regions of the PPT content rather than uniformly to all text keywords. It identifies key information regions and maps voice files only to those areas, creating a sparse but effective synchronization structure that reduces storage requirements while maintaining synchronization for the most important content.
Solution Approach 2:
The system extracts and identifies key information regions from the PPT content, separating them from the rest of the text. By taking out only the essential content areas for voice mapping, it reduces the total amount of data that needs to be stored and processed, thereby reducing storage space requirements while maintaining synchronization accuracy for the most critical information.
Data Source
AI summary
A text and voice information processing method includes: photographing a document to obtain a first picture, wherein the document comprises a first text keyword and a second text keyword; recording audio to obtain a voice file corresponding to the document; obtaining a first voice segment matching the first text keyword and a second voice segment matching the second text keyword from the voice file; displaying a second picture with a first play button and a second play button; in response to a user input, playing back the first voice segment or playing back the second voice segment.


