Distributed Data Extraction System Morphological Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data extraction systems, like the one described in Patent Document 1, face challenges in handling large volumes of web page data due to the burden of performing morphological analysis, extraction, and accumulation on a single apparatus, which is not feasible for websites with complex content like sounds and images, and requires significant data capacity.
Innovation Solution
A distributed data extraction system comprising multiple terminals and a server, where the terminal performs morphological analysis and extraction of prescribed data, and the server verifies and accumulates the data, sharing new phrases, images, and sounds among terminals, reducing the burden on each apparatus by distributing processes and compressing data for efficient processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single apparatus performs all data extraction processes (morphological analysis, extraction, accumulation), then the system structure is simple, but the burden on the single apparatus becomes excessive and data capacity requirements become unrealistic
Solution Approach 1:
The patent divides the data extraction system into multiple terminals and a server. Terminals perform morphological analysis and extract phrases from web pages, while the server accumulates and verifies extracted data. This segmentation distributes the computational burden across multiple devices, enabling the system to handle large volumes of web data that would be impossible for a single apparatus to process.
2Stability of the object's composition
If a single apparatus accumulates all extracted data, then data consistency is maintained, but the data capacity requirement becomes huge and unrealistic
Solution Approach 1:
The patent combines multiple terminals and a server into a distributed system where each terminal maintains local data accumulation capabilities. The server consolidates data from multiple terminals and performs verification to ensure consistency. This merging approach allows the system to handle huge quantities of data without requiring any single apparatus to have enormous data capacity.
3Productivity
If multiple terminals extract data independently, then the extraction capacity increases, but duplicate data accumulation occurs across terminals
Solution Approach 1:
The patent implements a feedback mechanism where terminals send extracted data to the server, which verifies whether the data already exists in its accumulation database. The server provides feedback to terminals about duplicate data, allowing terminals to avoid redundant extraction and storage. This feedback loop maintains data consistency while preserving the high extraction capacity of multiple terminals.
Data Source
AI summary
This invention provides a terminal searching for web pages on the web and extracting the prescribed data from the web pages and a server verifying and accumulating the extracted data. The prescribed data can be extracted from the web pages on the web in a manner that the process relating to the data extraction is distributed between the terminal and the server. Therefore, necessary processes up to the data extraction are distributed, and the burden placed on each apparatus can be lessened. Further, new data not formerly found in the web pages can be found out and extracted from the web pages that has been updated or newly made.


