Universal Internet Data Mining via Segmentation and Local Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current internet data mining methods face significant challenges in efficiently extracting valuable information from massive amounts of structured, semi-structured, and non-structured data due to varying value density and volume, lacking a systematic and practical solution for universal application.
Innovation Solution
A universal internet information data mining method utilizing a human-machine interaction template to acquire and process topic, pragmatic, and common keywords, enabling data mining operations such as search, statistics, analysis, and modeling across structured, semi-structured, and non-structured data types, leveraging the 'Double Ten Law' for effective information extraction and publishing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data mining is performed on massive internet information including structured, semi-structured and non-structured data, then the quantity of mined information increases, but the value density decreases due to the large volume of low-value non-structured data
Solution Approach 1:
The patent segments internet information into three distinct categories: structured data, semi-structured data, and non-structured data. Each category is processed through dedicated processing paths with appropriate techniques, allowing the system to handle large quantities of diverse data while maintaining value density by applying targeted extraction methods to each segment rather than treating all data uniformly
Solution Approach 2:
The patent applies different processing qualities and methods to different data segments based on their characteristics. Structured data receives database querying operations, semi-structured data receives parsing and extraction operations, and non-structured data receives text mining operations. This local quality approach ensures that each data type is processed with the most appropriate technique, maximizing value extraction while managing the overall quantity efficiently
2Adaptability or versatility
If comprehensive data mining is performed on all types of internet information, then the coverage of mined information increases, but the system complexity increases due to handling multiple data types and processing requirements
Solution Approach 1:
The patent creates a universal data mining system that handles multiple data types (structured, semi-structured, and non-structured) through a unified architecture. The system uses a common processing framework that automatically routes different data types to appropriate processing pipelines, providing multi-functionality without requiring separate systems for each data type, thus managing complexity while maintaining broad coverage
Solution Approach 2:
The patent introduces intermediary components including a data classification module that categorizes incoming data, a routing module that directs data to appropriate processing paths, and a result integration module that consolidates outputs. These intermediaries simplify the overall system complexity by managing the flow between diverse data sources and processing techniques through standardized interfaces
3Loss of information
If data mining focuses on structured data with high value density, then the quality of mining results improves, but the quantity of processed information decreases due to limited data volume
Solution Approach 1:
The patent merges the processing of structured, semi-structured, and non-structured data into a unified data mining operation. By combining these different data types, the system processes a much larger quantity of information than would be available from structured data alone, while maintaining result quality through the use of appropriate extraction techniques for each data type and integrating results from all sources
Data Source
AI summary
By means of providing directly a data mining requiring user with a universal internet information data mining requirement description human-machine interaction template, the present invention provides big internet data with a set of both open and strictly-defined constraints for concept collection, data structures, and data mining operations, thus satisfying three factors for establishing a data mining model, providing an important condition for increasing the value density of an internet mining service, and allowing for implementation of universal and parallel mining of structured data, semi-structured data, and non-structured data of the internet.


