Web Content Deduplication via Fuzzy Indexing and API Integration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing comparison websites rely on centralized databases that require manual updates and massive crawls to track dynamic web content, leading to inefficiencies and time-consuming data maintenance due to constant changes in web content across multiple websites.

Innovation Solution

A method and system for deduplication of web content that converts search results from multiple websites into fuzzy indexes for real-time comparison and matching, eliminating the need for centralized databases by auto-matching and simplifying data, thereby reducing processing power and time needed for comparisons.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a centralized database is used to store web content from multiple websites, then web content can be collected and stored, but the database requires manual updates and massive crawls to track dynamic changes, leading to significant time consumption and inefficiency

Engineering Contradiction:
Improveweb content storage capacityVSAvoidtime for manual updates and crawls
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system enables self-service by allowing websites to automatically provide their own content through integrated APIs. Each website acts as its own data source, automatically updating content without requiring manual crawling or intervention from the comparison platform. This resolves the time consumption issue while maintaining comprehensive content storage.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Websites pre-process and structure their data according to standardized schemas before making it available through APIs. This preliminary action ensures that when data is retrieved by the comparison platform, it is already in the correct format for immediate use, eliminating the need for time-consuming manual crawling and processing.

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If manual pointers are generated to track web content from multiple sites, then content can be organized for comparison, but significant time and delay are required for manual updates when hotels open or close

Engineering Contradiction:
Improvecontent organization capabilityVSAvoidtime for manual pointer updates
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system eliminates manual pointer management by implementing automated discovery mechanisms. The comparison platform automatically discovers new websites and content sources through structured data integration, continuously updating its network of data sources without human intervention. This maintains ease of operation while removing the time burden of manual updates.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements continuous feedback loops where the comparison platform monitors for changes in website availability and content structure. When new hotels open or websites change their data formats, the system automatically detects these changes through API responses and adjusts its pointers accordingly, eliminating manual update requirements.

Inventive Principle:
Principle #23Feedback

3Reliability

If a centralized database performs massive daily crawls to keep track of data changes, then the most recent changes can be maintained, but the process is significant and time-consuming

Engineering Contradiction:
Improvedata currencyVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

Websites automatically serve their own updated content through integrated APIs, eliminating the need for the comparison platform to perform crawling operations. Each website pushes or makes available its latest data through standardized interfaces, ensuring data currency while dramatically improving processing efficiency by replacing resource-intensive crawls with efficient API calls.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system replaces the mechanical crawling process with electronic API-based data retrieval. Instead of using web crawlers that systematically scan and parse HTML content (a resource-intensive mechanical process), the system uses structured API calls that directly access standardized data formats, significantly reducing processing time and computational resources while maintaining data reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If duplicate search results from multiple websites are included in combined lists, then comprehensive results are provided, but redundant information increases processing load and user confusion

Engineering Contradiction:
Improvemulti-source content aggregationVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Each website automatically provides standardized data through its API, including unique identifiers and structured attributes. This self-service approach enables the comparison platform to efficiently match and deduplicate results across sources by comparing standardized fields, maintaining comprehensive multi-source aggregation while reducing processing complexity through automated deduplication algorithms.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system transforms unstructured web content into standardized parameters and attributes through API responses. By converting diverse website data formats into uniform parameter structures (price, availability, hotel name, location, etc.), the system enables efficient automated deduplication and comparison, reducing processing complexity while maintaining the ability to aggregate from multiple diverse sources.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10977321B2System and method for web content matching
Publication Date: 2021.04.13 DECKARD TECHNOLOGIES INC
  • US10977321B2 patent drawing
  • US10977321B2 patent drawing
  • US10977321B2 patent drawing

AI summary

Provided are a system and method for performing deduplication of web content. In one example, the method includes converting search results of a first website into a first fuzzy index and converting search results of a second website into a second fuzzy index, determining a search result of the first website corresponds to a same item as a search result of the second website based on a comparison of the first fuzzy index and the second fuzzy index, and displaying a comparison of web content associated with the item from the first search result and web content associated with item from the second search result. The deduplication of content according to various embodiments may be performed on the fly without storing web content in a centralized database.