Distributed Data Extraction System Morphological Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data extraction systems, like the one described in Patent Document 1, face challenges in handling large volumes of web page data due to the burden of performing morphological analysis, extraction, and accumulation on a single apparatus, which is not feasible for websites with complex content like sounds and images, and requires significant data capacity.

Innovation Solution

A distributed data extraction system comprising multiple terminals and a server, where the terminal performs morphological analysis and extraction of prescribed data, and the server verifies and accumulates the data, sharing new phrases, images, and sounds among terminals, reducing the burden on each apparatus by distributing processes and compressing data for efficient processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single apparatus performs all data extraction processes (morphological analysis, extraction, accumulation), then the system structure is simple, but the burden on the single apparatus becomes excessive and data capacity requirements become unrealistic

Engineering Contradiction:
Improvesystem structureVSAvoiddata extraction capacity
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent divides the data extraction system into multiple terminals and a server. Terminals perform morphological analysis and extract phrases from web pages, while the server accumulates and verifies extracted data. This segmentation distributes the computational burden across multiple devices, enabling the system to handle large volumes of web data that would be impossible for a single apparatus to process.

Inventive Principle:
Principle #1Segmentation

2Stability of the object's composition

If a single apparatus accumulates all extracted data, then data consistency is maintained, but the data capacity requirement becomes huge and unrealistic

Engineering Contradiction:
Improvedata consistencyVSAvoiddata capacity
Core Design Contradiction:
Stability of the object's compositionVSQuantity of substance

Solution Approach 1:

The patent combines multiple terminals and a server into a distributed system where each terminal maintains local data accumulation capabilities. The server consolidates data from multiple terminals and performs verification to ensure consistency. This merging approach allows the system to handle huge quantities of data without requiring any single apparatus to have enormous data capacity.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If multiple terminals extract data independently, then the extraction capacity increases, but duplicate data accumulation occurs across terminals

Engineering Contradiction:
Improveextraction capacityVSAvoiddata redundancy
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent implements a feedback mechanism where terminals send extracted data to the server, which verifies whether the data already exists in its accumulation database. The server provides feedback to terminals about duplicate data, allowing terminals to avoid redundant extraction and storage. This feedback loop maintains data consistency while preserving the high extraction capacity of multiple terminals.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8321198B2Data extraction system, terminal, server, programs, and media for extracting data via a morphological analysis
Publication Date: 2012.11.27 SQUARE ENIX HLDG CO LTD
  • US8321198B2 patent drawing
  • US8321198B2 patent drawing
  • US8321198B2 patent drawing

AI summary

This invention provides a terminal searching for web pages on the web and extracting the prescribed data from the web pages and a server verifying and accumulating the extracted data. The prescribed data can be extracted from the web pages on the web in a manner that the process relating to the data extraction is distributed between the terminal and the server. Therefore, necessary processes up to the data extraction are distributed, and the burden placed on each apparatus can be lessened. Further, new data not formerly found in the web pages can be found out and extracted from the web pages that has been updated or newly made.