Web Page Categorization Using Search Engine Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current web page categorization methods for parental control and malicious content prevention are inefficient due to the rapid creation of new web pages, leading to slow and resource-intensive content scanning solutions that cannot keep up with the pace of new content distribution.
Innovation Solution
A method and apparatus that detect and analyze web content elements from search engine results to categorize web pages, using a processor and memory with computer program code to identify and categorize content based on analysis, enabling efficient categorization without sacrificing speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If content scanning is performed on every web page by analyzing page content itself, then categorization accuracy is improved, but processing speed deteriorates and system resources are consumed
Solution Approach 1:
The patent applies preliminary action by using search engine results to pre-categorize web pages before actual content scanning occurs. The system leverages the search engine's existing analysis of web content (titles, descriptions, keywords) to assign preliminary categories, so that when a user requests a web page, the categorization is already done or can be quickly verified against the pre-analyzed search data, avoiding the need to scan every page from scratch
Solution Approach 2:
The patent introduces an intermediary mechanism - a local database or cache that stores categorization information from search engine results. This intermediary layer sits between the search engine and the content scanner, allowing the system to retrieve pre-categorized information quickly without triggering full content analysis, thus maintaining accuracy while improving speed
2Measurement precision
If content scanning is performed on every web page by analyzing page content itself, then categorization accuracy is improved, but system resource consumption increases
Solution Approach 1:
The search engine performs the energy-intensive content analysis in advance during its normal web crawling and indexing operations. By the time a user requests a web page, the categorization work has already been done by the search engine, eliminating the need for duplicate resource-intensive scanning at the client or gateway level
Solution Approach 2:
The patent uses copying by storing categorization metadata from search engine results in a local cache or database. Instead of re-analyzing content, the system copies and reuses the categorization information already extracted by the search engine, significantly reducing computational resources required at the point of use
3Productivity
If URL comparison to database of unwanted/indexed URLs is performed, then processing speed is improved, but categorization capability deteriorates due to inability to keep up with new web pages
Solution Approach 1:
The system performs preliminary categorization by leveraging search engine results that are continuously updated as the search engine crawls new web pages. This preliminary categorization happens in advance and is stored locally, enabling fast retrieval while automatically adapting to new content as the search engine indexes it, without requiring manual database updates
4Reliability
If real-time categorization is implemented before user selection, then parental control effectiveness is improved, but processing overhead increases
Solution Approach 1:
The patent introduces an intermediary cache or local database that stores categorization information from search engine results. This intermediary layer provides real-time categorization by quickly retrieving pre-analyzed data without requiring complex real-time analysis, thus maintaining parental control effectiveness while minimizing processing overhead
Solution Approach 2:
The system copies categorization metadata from search engine results to a local storage mechanism, enabling instant retrieval and real-time blocking decisions without performing complex analysis at the moment of user interaction, thereby reducing processing overhead while maintaining control effectiveness
Data Source
AI summary
In accordance with an example embodiment of the present invention, there is provided an apparatus, including at least one processor; and at least one memory including computer program code the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following: detecting a listing of web content elements provided by a web search engine, the web content elements relating to web pages retrieved by the web search engine; analyzing one or more web content elements of the detected listing; and categorizing the content of one or more web pages on the basis of the analysis.


