Multimedia Stream Search via Speech Recognition Timing Tags
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional videoconferencing and streaming systems lack an efficient method for searching archived multimedia content, relying on manual metadata association which is labor-intensive and unreliable.
Innovation Solution
A system and method that converts multimedia streams from conventional conference formats into a searchable format by using a speech recognition engine to generate models of sound fragments, comparing them with reference models, and associatively storing timing information with recognized words, enabling keyword search and playback of specific conference segments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual metadata association is used to enable search in archived stream data, then search capability is provided, but a lot of manual work is required and reliability is low
Solution Approach 1:
The system automatically generates metadata by analyzing audio content through speech recognition and scene detection algorithms. The conversion engine autonomously extracts speaker information, speech content, and scene descriptions without requiring manual intervention, making the system self-serve the metadata generation task
Solution Approach 2:
Manual metadata creation is replaced by automated computational processes including speech-to-text conversion, scene detection algorithms, and automatic tagging systems. These mechanical/computational systems perform the metadata generation task that previously required human operators
2Ease of operation
If manual metadata association is used to enable search in archived stream data, then search capability is provided, but accuracy of metadata correspondence is not guaranteed
Solution Approach 1:
The system uses timing information and synchronization mechanisms to ensure metadata accurately corresponds to the actual stream content. Speech recognition results are time-stamped and linked to specific audio segments, providing feedback verification that the metadata correctly represents the archived content
Solution Approach 2:
Manual metadata creation prone to human error is replaced by automated speech recognition and scene detection systems that systematically analyze the actual content. This substitution ensures metadata accuracy through consistent, repeatable computational processes rather than variable human performance
3Ease of operation
If conventional conference format coded data stream is converted to multimedia streaming format with timing information, then searchability is enabled, but system complexity increases
Solution Approach 1:
The conversion engine performs format conversion and metadata generation in advance during the archiving process. By completing the conversion to multimedia streaming format with embedded timing information upfront, the system eliminates the need for complex real-time processing during search operations
Solution Approach 2:
A dedicated conversion engine acts as an intermediary component between the conventional conference system and the archiving system. This specialized module handles the complex format conversion and metadata extraction tasks, isolating the complexity from the main search and playback systems
Data Source
AI summary
The present invention provides a system and a method making an archived conference or presentation searchable after being stored in the archive server. According to the invention, one or more media streams coded according to H.323 or SIP are transmitted to a conversion engine for converting multimedia content into a standard streaming format, which may be a cluster of files, each representing a certain medium (audio, video, data) and/or a structure file that synchronizes and associates the different media together. When the conversion is carried out, the structure file is copied and forwarded to a post-processing server. The post-processing server includes i.a. a speech recognition engine generating a text file of alphanumeric characters representing all recognized words in the audio file. The text file is then entered into the cluster of files associating each identified word to a timing tag in the structure file. After this post-processing, finding key words and associated points of time in the media stream could easily be executed by a conventional search engine.


