Voice Search via Closed Caption Matching for Streamed Media
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for selecting and accessing specific video clips in streamed media are time-consuming and cumbersome, requiring users to navigate through multiple sources and electronic program guides.
Innovation Solution
A content-focused television receiver system that utilizes cloud-based voice searching to convert voice requests to text, matching the text with closed caption or subtitle text in streamed media to identify and play specific video clips.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If conventional channel surfing and electronic program guide navigation are used to locate video clips, then users can access available programming content, but the process becomes time-consuming and cumbersome
Solution Approach 1:
The patent replaces the mechanical navigation system (remote control, channel surfing, EPG browsing) with a voice-based search system. Users speak natural language queries about video content, and the system processes these语音 requests through speech-to-text conversion and natural language processing to directly locate and play the desired video clips, eliminating the need for manual navigation through multiple sources and guides.
2Ease of operation
If users manually navigate through multiple sources and electronic program guides to find specific content, then they can access programming from various sources, but the operation becomes complex and frustrating
Solution Approach 1:
The patent introduces a voice processing intermediary system that acts as a mediator between the user and the complex multi-source content delivery system. The speech-to-text converter and natural language processing components serve as intermediaries that translate user intent into precise content retrieval operations, simplifying the interaction model while maintaining access to diverse programming sources.
3Productivity
If cloud-based voice searching with speech-to-text conversion is implemented, then users can quickly locate video clips by voice, but the system complexity increases
Solution Approach 1:
The patent introduces a voice processing intermediary system that acts as a mediator between the user and the complex multi-source content delivery system. The speech-to-text converter and natural language processing components serve as intermediaries that translate user intent into precise content retrieval operations, simplifying the interaction model while maintaining access to diverse programming sources.
Solution Approach 2:
The system performs preliminary speech-to-text conversion and natural language processing of voice requests before initiating the video clip search. By pre-processing the voice input into text format and extracting key search terms in advance, the system prepares the query for efficient matching against video metadata and closed caption text, thereby accelerating the overall content retrieval process.
Data Source
AI summary
Methods, systems, and apparatuses are described to implement voice search in media content for requesting media content of a video clip of a scene contained in the media content streamed to the client device; for capturing the voice request for the media content of the video clip to display at the client device wherein the streamed media content is a selected video streamed from a video source; for applying a NLP solution to convert the voice request to text for matching to a set of one or more words contained in at least close caption text of the selected video; for associating matched words to close caption text with a start index and an end index of the video clip contained in the selected video; and for streaming the video clip to the client device based on the start index and the end index associated with matched closed caption text.


