Structured Content Search Engine for Web Page Constituents
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current search technologies fail to accurately retrieve relevant content from web pages, leading to wasted bandwidth and screen real estate, and are ineffective in handling complex search expressions and structural proximity relationships, resulting in false hits and sub-optimal search results on mobile devices.
Innovation Solution
A structured content search engine that analyzes content structures, using tree and graph structures, layout information, and category data to identify and return relevant document constituents, rather than entire documents, and supports advanced search expressions with structural proximity operators to improve search accuracy and relevance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If standard web search engines return entire web pages, then search coverage is comprehensive, but bandwidth is wasted and screen real estate is wasted
Solution Approach 1:
The patent segments web pages into hierarchical constituents (document, section, paragraph, sentence, word) and enables searching at different granularities. The search engine can return only the specific constituent containing the search term rather than the entire page, reducing bandwidth consumption while maintaining search effectiveness.
Solution Approach 2:
The patent extracts and returns only the relevant constituents (such as specific paragraphs or sentences containing search terms) rather than delivering entire web pages. This extraction approach eliminates unnecessary content transmission, directly addressing bandwidth waste while preserving the essential search results.
2Reliability
If standard web search engines return entire web pages, then search coverage is comprehensive, but screen real estate is wasted
Solution Approach 1:
The patent enables hierarchical navigation through document constituents, allowing users to drill down from document level to section, paragraph, and sentence level. This segmentation lets users view only the relevant portions of web pages on their screens, reducing the amount of content displayed while maintaining comprehensive search coverage.
Solution Approach 2:
The search engine extracts and displays only the specific constituents containing search terms rather than rendering entire web pages. This extraction reduces screen real estate consumption by eliminating unrelated content from the display while preserving all search results.
3Ease of manufacture
If sub-document search engines use string-based algorithms, then implementation is simple, but structural proximity relationships are not exploited
Solution Approach 1:
The patent adds a structural dimension to search by incorporating tree-based algorithms that analyze the hierarchical structure of web pages. Instead of treating documents as flat strings of text, the system traverses the document tree to identify constituents based on both text content and structural position, improving search accuracy without abandoning simple string matching.
Solution Approach 2:
The patent combines string-based algorithms with tree-based algorithms to create a composite search approach. The string-based component handles simple text matching while the tree-based component analyzes structural relationships, creating a more powerful search system that leverages the strengths of both approaches.
4Measurement precision
If search engines ignore constituent boundaries, then proximity relationships are captured, but false hits increase
Solution Approach 1:
The patent segments documents into hierarchical constituents and uses these segments as the basis for proximity calculations. By respecting constituent boundaries (such as paragraphs or sections), the system can determine whether search terms are truly proximate within the same logical unit, reducing false hits while maintaining accurate proximity detection.
Data Source
AI summary
Embodiments of methods and apparatuses for searching contents, including structured search for atomic search expressions, including proximately associated atomic search expressions, are described herein. Embodiments may use tree structures (or more generally, graph structures), layout structures, and/or other information to capture within search results relevant content, include sub-document constituents, to reduce the incidence of false positives within search results, and/or to improve the accuracy of rankings within search results. Embodiments may use distance and/or scoring functions to generate scores for the structures to indicate relevance, including usage of local geometry, and linear iteration over portions of the content at a level to capture potential of a portion to influence other portions of the level, and influence received by a portion from the other portions of the level. Other embodiments may be described and claimed.


