Voice Search Query Building With Conjunction-Triggered Entity Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face challenges in efficiently searching for content items featuring specific entities when they can only recall partial details, such as actors or locations, leading to time-consuming manual searches without effective recall mechanisms.
Innovation Solution
A system that combines voice input with user selection of on-screen and real-world entities, using cameras and natural language processing to construct search queries based on gestures and pronouns, allowing for the identification of on-screen entities and real-world landmarks, and utilizing conjunctions to manage multiple entity selections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If users manually search for content items by entering search queries and reviewing results, then they can locate content featuring specific entities, but the process becomes time-consuming and inefficient
Solution Approach 1:
The system enables users to perform searches through natural voice commands and gestures without requiring manual typing or navigation. The voice-activated search interface automatically processes queries and returns results, allowing the system to serve itself by interpreting user intentions directly from speech and motion inputs.
Solution Approach 2:
The patent replaces manual mechanical interactions (typing, clicking, navigating menus) with acoustic and gestural inputs. The voice-activated search mechanism substitutes the mechanical keyboard and mouse operations with speech recognition and gesture recognition systems, significantly reducing the physical effort and time required for searching.
2Loss of information
If users rely on memory to recall details about content items, then they can identify what they are looking for, but they often cannot recall sufficient details for effective searching
Solution Approach 1:
The system introduces an intermediary layer between the user's partial memories and the search database. By capturing gestures and processing voice commands, the system acts as a mediator that translates incomplete user recollections into effective search queries, bridging the gap between what users can remember and what can be searched.
Solution Approach 2:
The patent changes the parameters of information input from precise textual data to flexible acoustic and gestural signals. This allows users to search using approximate memories and natural speech patterns rather than requiring exact entity names or detailed search criteria, making the search process more tolerant of imperfect recall.
3Measurement precision
If the search interface requires detailed search parameters, then search results can be precise, but users struggle to provide accurate information when recalling content
Solution Approach 1:
The search interface segments the search process into multiple stages: voice command interpretation, gesture recognition, entity identification, and query construction. This segmentation allows each component to handle specific tasks independently, processing user inputs in manageable steps rather than requiring all detailed parameters to be provided simultaneously.
Solution Approach 2:
The system performs preliminary actions by automatically interpreting voice commands and gestures before constructing the final search query. Entity recognition and query preparation occur in advance, transforming raw user inputs into structured search parameters, so that by the time the actual search executes, all necessary precise parameters are already prepared.
Data Source
AI summary
A search is performed based on a voice input combined with user selection of entities displayed on a display screen as well as real-world entities. A voice input is received from the user by a media device, as well as a selection of a first entity being displayed on the media device. A conjunction spoken in the voice input triggers the media device to wait for selection of a second entity before performing the search. After receiving selection of the second entity, a search query is constructed based on the voice input, the first entity, and the second entity. The search query is transmitted to a database and, in response, the media device receives at least one identifier of a least one content item. The at least one identifier is then generated for display to the user.


