Document Ranking via Content Originality Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Search engines struggle to distinguish between original and redundant content, leading to obscured unique results due to the proliferation of nearly identical documents in search results.
Innovation Solution
A method to rank documents based on the originality of their content by identifying and scoring unique content pieces, where the score is determined by whether the content occurs first in the corpus, and using this score to enhance or reduce the rank of associated authors and documents, thereby prioritizing original content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If search engines include all documents in search results, then comprehensive coverage is achieved, but redundant and duplicate content obscures unique results
Solution Approach 1:
The patent segments the document collection into original documents and duplicate documents by analyzing content similarity. Search results are segmented to include only original documents, excluding duplicates. This segmentation resolves the contradiction by maintaining comprehensive coverage of unique content while eliminating redundant entries that obscure valuable information.
Solution Approach 2:
The patent extracts and identifies duplicate content from the document collection using similarity analysis. By taking out duplicate documents from the search results and keeping only original documents, the system maintains comprehensive coverage while eliminating redundancy, thus resolving the contradiction between quantity and information visibility.
2Measurement precision
If search engines rank documents by explicit references (hyperlinks), then document connectivity is measured, but content originality is not captured
Solution Approach 1:
The patent introduces a feedback mechanism where content similarity analysis results are fed back into the ranking process. By analyzing whether documents contain duplicate content and adjusting rankings accordingly, the system improves measurement precision by incorporating content originality information, thus resolving the contradiction between connectivity-based ranking and originality detection.
Solution Approach 2:
The patent introduces content similarity analysis as an intermediary mechanism between document retrieval and ranking. This intermediary layer analyzes content originality and provides adjusted rankings that reflect both connectivity and uniqueness, resolving the contradiction by mediating between the two competing ranking criteria.
3Quantity of substance
If multiple copies of the same content appear in search results, then all instances are captured, but information value is reduced
Solution Approach 1:
The patent segments documents into original and duplicate categories based on content similarity analysis. By segmenting search results to include only original documents and excluding duplicates, the system reduces the number of results while significantly improving information value and reliability, thus resolving the contradiction between quantity and quality.
Solution Approach 2:
The patent conceptually 'colors' documents as either original or duplicate based on content analysis. This classification system allows the search engine to differentiate between valuable original content and redundant copies, selecting only original documents for display. This resolves the contradiction by filtering out low-value duplicates while maintaining high-value original content.
Data Source
AI summary
Methods, systems, and apparatus, including computer program products for identifying original content. In one aspect a method is described that includes identifying a first document in a collection of documents. The first document contains a content piece and the content piece does not occur in any earlier document in the collection. The first document is associated with a first author and the first author associated with a first rank. The first rank of the first author is determined using a score of the content piece. The score is a figure of merit of the content piece.


