Network Page Generation for Search Engine Indexing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Network site owners face challenges in protecting their valuable data from being indexed by search engines while still generating attractive network pages that can drive search traffic to their sites.

Innovation Solution

A system comprising a network page generation application and data extraction application processes a corpus of data to extract relevant information, generating network pages optimized for search engines while protecting sensitive data from indexing, using techniques like latent Dirichlet allocation and tf-idf analysis, and employing 'nofollow' attributes and captchas to control access.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the entire corpus of data is made accessible to search engines, then search traffic to the network site increases, but the network site owner loses control over their valuable data

Engineering Contradiction:
Improvesearch trafficVSAvoiddata loss of control
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent divides the corpus of data into two distinct segments: (1) a processed representation that is extracted and made accessible to search engines for indexing, and (2) the original protected corpus that remains inaccessible to search engines. This segmentation allows search traffic to be generated from the processed segment while the protected segment retains its security and control

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary processing step where a application extracts and processes data from the protected corpus to create a search-engine-friendly representation. This intermediary layer acts as a mediator that translates the protected data into a form that search engines can index without granting them direct access to the original corpus, thus maintaining control while enabling search visibility

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If search engines crawl and index all network pages, then search engine visibility improves, but sensitive data may be indexed without authorization

Engineering Contradiction:
Improvesearch engine visibilityVSAvoiddata protection
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies different quality characteristics to different parts of the data structure: the processed representation is optimized for search engine indexing with appropriate metadata and formatting, while the original corpus maintains its protected status with restricted access controls. This local differentiation in quality and accessibility allows each part to serve its specific function without compromising the other

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent creates a processed copy or representation of the data from the protected corpus. This copy contains the essential information needed for search engine indexing but is fundamentally different from and does not provide direct access to the original protected data. The copy serves the search visibility function while the original remains secure

Inventive Principle:
Principle #26Copying

3Object-affected harmful factors

If the network site owner uses robot exclusion standards and nofollow attributes to protect data, then data control improves, but search traffic from protected pages decreases

Engineering Contradiction:
Improvedata controlVSAvoidsearch traffic
Core Design Contradiction:
Object-affected harmful factorsVSProductivity

Solution Approach 1:

The patent extracts the essential information and value from the protected corpus and places it into a separate processed representation that is specifically designed for search engine access. This extracted version is published and linked in a way that attracts search traffic, while the original protected pages can maintain their exclusion directives without losing search visibility, as the value has been transferred to the extracted representation

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS8438149B1Generating network pages for search engines
Publication Date: 2013.05.07 AMAZON TECH INC
  • US8438149B1 patent drawing
  • US8438149B1 patent drawing
  • US8438149B1 patent drawing

AI summary

Disclosed are various embodiments generating network pages for search engines from data protected from search engines. A portion of data from a corpus of data that is protected from indexing by a search engine is extracted. A first network page is generated based at least in part on the portion of data. The first network page is configured for the search engine to index the portion of data. The first network page omits a context for the portion of data from the corpus of data. The first network page includes one or more links to a second network page that is protected from indexing by the search engine. The second network page provides access to the corpus of data.