Traffic-Based Content Acquisition and Indexing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional web crawlers struggle to keep information up-to-date as they infrequently revisit sites, leading to outdated search results and increased network load when visiting sites more frequently.
Innovation Solution
A method and apparatus for traffic-based content acquisition and indexing that scans packets as content is transferred across the network, building a keyterm index in real-time to update information resources without relying on hyperlinks, allowing for timely and efficient identification of changes in web page content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a web crawler visits sites more frequently to keep information fresh, then the freshness of search results is improved, but the network load and server load increase
Solution Approach 1:
The patent applies preliminary action by scanning and indexing content elements as they are being transferred across the network, before the content is fully delivered to the destination. This allows the search engine to have updated information ready in advance without requiring frequent revisits to the source site, thus maintaining result freshness while reducing overall network load.
Solution Approach 2:
The patent introduces an intermediary scanning mechanism that intercepts content during its transfer across the network. This intermediary process extracts and indexes content elements on-the-fly, acting as a mediator between the content source and the search engine, thereby eliminating the need for direct frequent connections between crawlers and sites.
2Loss of energy
If a web crawler infrequently visits sites to reduce network load, then the network load is reduced, but the search results become outdated
Solution Approach 1:
The patent implements continuity of useful action by continuously scanning content elements as they traverse the network, rather than periodically revisiting sites. This continuous monitoring ensures that content changes are captured in real-time, maintaining up-to-date search results while avoiding the overhead of frequent complete site revisits.
Solution Approach 2:
The intermediary scanning process continuously intercepts content during transfer, providing ongoing updates to the index without requiring direct frequent communication between crawlers and content sources, thus maintaining reliability while minimizing network load.
3Ease of operation
If a web crawler relies on hyperlinks to discover content, then the crawling process is simplified, but content that is not published via hyperlinks remains undetected
Solution Approach 1:
The patent applies universality by implementing a scanning mechanism that can detect and index content elements regardless of how they are published or delivered. This multi-functional approach allows the system to capture content through network traffic scanning rather than relying solely on hyperlink following, ensuring both simplicity and completeness.
Solution Approach 2:
The intermediary scanning process acts as a universal detector that intercepts all content transfers across the network, regardless of their publication method. This mediator approach ensures that both hyperlink-published and non-hyperlink-published content is detected and indexed, eliminating information loss while maintaining operational simplicity.
Data Source
AI summary
A method and apparatus for processing packets in a network are disclosed. For example, the method scans one or more packets representing a content that is being transferred via the network, where the scanning acquires one or more content elements. The method then builds a keyterm index from the one or more content elements, and stores the keyterm index in a repository. A query handler then responds to queries in accordance with the keyterm index.


