Universal Internet Data Mining via Segmentation and Local Quality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current internet data mining methods face significant challenges in efficiently extracting valuable information from massive amounts of structured, semi-structured, and non-structured data due to varying value density and volume, lacking a systematic and practical solution for universal application.

Innovation Solution

A universal internet information data mining method utilizing a human-machine interaction template to acquire and process topic, pragmatic, and common keywords, enabling data mining operations such as search, statistics, analysis, and modeling across structured, semi-structured, and non-structured data types, leveraging the 'Double Ten Law' for effective information extraction and publishing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data mining is performed on massive internet information including structured, semi-structured and non-structured data, then the quantity of mined information increases, but the value density decreases due to the large volume of low-value non-structured data

Engineering Contradiction:
Improvequantity of mined informationVSAvoidvalue density
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments internet information into three distinct categories: structured data, semi-structured data, and non-structured data. Each category is processed through dedicated processing paths with appropriate techniques, allowing the system to handle large quantities of diverse data while maintaining value density by applying targeted extraction methods to each segment rather than treating all data uniformly

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different processing qualities and methods to different data segments based on their characteristics. Structured data receives database querying operations, semi-structured data receives parsing and extraction operations, and non-structured data receives text mining operations. This local quality approach ensures that each data type is processed with the most appropriate technique, maximizing value extraction while managing the overall quantity efficiently

Inventive Principle:
Principle #3Local quality

2Adaptability or versatility

If comprehensive data mining is performed on all types of internet information, then the coverage of mined information increases, but the system complexity increases due to handling multiple data types and processing requirements

Engineering Contradiction:
Improvecoverage of mined informationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal data mining system that handles multiple data types (structured, semi-structured, and non-structured) through a unified architecture. The system uses a common processing framework that automatically routes different data types to appropriate processing pipelines, providing multi-functionality without requiring separate systems for each data type, thus managing complexity while maintaining broad coverage

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces intermediary components including a data classification module that categorizes incoming data, a routing module that directs data to appropriate processing paths, and a result integration module that consolidates outputs. These intermediaries simplify the overall system complexity by managing the flow between diverse data sources and processing techniques through standardized interfaces

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If data mining focuses on structured data with high value density, then the quality of mining results improves, but the quantity of processed information decreases due to limited data volume

Engineering Contradiction:
Improvequality of mining resultsVSAvoidquantity of processed information
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent merges the processing of structured, semi-structured, and non-structured data into a unified data mining operation. By combining these different data types, the system processes a much larger quantity of information than would be available from structured data alone, while maintaining result quality through the use of appropriate extraction techniques for each data type and integrating results from all sources

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10108717B2Universal internet information data mining method
Publication Date: 2018.10.23 CHONGQING SIZAI INFORMATION TECH CO LTD
  • US10108717B2 patent drawing
  • US10108717B2 patent drawing
  • US10108717B2 patent drawing

AI summary

By means of providing directly a data mining requiring user with a universal internet information data mining requirement description human-machine interaction template, the present invention provides big internet data with a set of both open and strictly-defined constraints for concept collection, data structures, and data mining operations, thus satisfying three factors for establishing a data mining model, providing an important condition for increasing the value density of an internet mining service, and allowing for implementation of universal and parallel mining of structured data, semi-structured data, and non-structured data of the internet.