Content Extraction Engine Noise Removal

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing content extraction systems face challenges such as high processing requirements, false positives, format variations, and high manpower needs when extracting product data from online content pages, particularly due to noise content and varying data elements.

Innovation Solution

A content extraction engine that analyzes HTML content pages, removes noise content using synonym lists, and automatically identifies and extracts target product data, reducing processing requirements and false positives while adjusting to different formats.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If content extraction systems process all content pages including noise content, then extraction coverage is complete, but processing requirements and network bandwidth increase significantly

Engineering Contradiction:
Improveextraction efficiencyVSAvoidprocessing requirements
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The system extracts and removes noise content (such as advertisements, navigation elements, and other non-product-related elements) from web pages before performing data extraction. This is achieved by identifying and eliminating unnecessary content elements, thereby reducing the processing load while maintaining focus on relevant product data extraction.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The content extraction process is divided into distinct stages: first identifying and removing noise content, then extracting target product data from the cleaned content. This segmentation allows the system to process only relevant content, improving efficiency while reducing overall processing requirements.

Inventive Principle:
Principle #1Segmentation

2Productivity

If content extraction systems process all content pages including noise content, then extraction coverage is complete, but network bandwidth requirements increase

Engineering Contradiction:
Improveextraction efficiencyVSAvoidnetwork bandwidth requirements
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The system extracts and removes noise content (such as advertisements, navigation elements, and other non-product-related elements) from web pages before performing data extraction. This is achieved by identifying and eliminating unnecessary content elements, thereby reducing the processing load while maintaining focus on relevant product data extraction.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If content extraction systems manually analyze content pages, then extraction accuracy is high, but manpower requirements and processing time increase

Engineering Contradiction:
Improveextraction accuracyVSAvoidmanpower involvement
Core Design Contradiction:
Measurement precisionVSExtent of automation

Solution Approach 1:

The system performs automated content extraction by analyzing HTML content structures, identifying product data elements, and extracting relevant information without requiring manual intervention. The automated process uses parsing algorithms and data extraction techniques to achieve high accuracy while eliminating the need for manual analysis.

Inventive Principle:
Principle #25Self-service

4Adaptability or versatility

If content extraction systems use fixed extraction methods, then processing is simple, but adaptability to varying content formats is poor

Engineering Contradiction:
Improveformat adaptabilityVSAvoidextraction system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The content extraction system dynamically adapts to different content formats by analyzing the structure of HTML content pages and adjusting extraction methods accordingly. The system identifies different content patterns and formats, then applies appropriate extraction techniques for each type, enabling versatile processing without requiring overly complex predefined rules for every possible format.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11556232B2Content extraction system
Publication Date: 2023.01.17 EBAY INC
  • US11556232B2 patent drawing
  • US11556232B2 patent drawing
  • US11556232B2 patent drawing

AI summary

A system includes a content extraction engine comprising at least one processor and configured to receive a content page for a target product including product data for the target product and noise content unrelated to the target product, identify noise content pertaining to data unrelated to the target product, remove noise content from the content page, thereby generating a remainder content page containing target product data usable to enable product comparison between multiple sources.