Multi-stage Risk Filtering for Big Data Collection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data collection systems for big data face challenges in efficiently filtering out malicious data and managing computational resources, leading to security risks and excessive resource consumption due to excessive keywords or rules.
Innovation Solution
A data collection system utilizing a series of risk filtering modules (first-order, second-order, and third-order) combined with data classification, normalization, and clustering analysis to selectively extract and filter raw data from various sources, including text, images, and executable scripts, while addressing cyber security and system security issues.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional filtering methods with keywords or rules are used to extract desired results from big data, then the system can identify and collect specific data, but the system consumes excessive computational resources and collects malicious data or data beyond usable extents
Solution Approach 1:
The patent divides the data filtering process into multiple sequential stages: initial keyword-based filtering, followed by risk level assessment (first-order, second-order, third-order risks), and final data extraction. This segmentation allows the system to process data in manageable chunks, reducing overall computational burden while maintaining filtering accuracy through progressive refinement.
Solution Approach 2:
The system performs preliminary risk assessment and filtering before detailed data extraction. By pre-identifying and removing high-risk data (malicious content, security threats, undesirable data) in early stages, the system reduces the volume of data requiring intensive processing later, thereby lowering computational resource consumption while preserving filtering precision.
2Measurement precision
If traditional filtering methods with keywords or rules are used, then the system can collect specific data, but the system collects malicious data or data out of usable extents leading to information security risks
Solution Approach 1:
The patent implements preliminary anti-action by proactively identifying and removing malicious data, security threats, and undesirable content before they can cause harm. The system uses multiple risk filtering modules to detect and eliminate potential security risks in advance, preventing harmful data from entering the usable data pool and thus protecting against information security risks.
Solution Approach 2:
The system converts the potential harm of malicious data into benefit by using it as training material for improving risk detection algorithms. By analyzing patterns in malicious data captured during filtering, the system enhances its ability to identify and neutralize security threats, turning harmful inputs into opportunities for strengthening security defenses.
3Measurement precision
If excessive rules or keywords are used in filtering methods, then the system can attempt to filter data more thoroughly, but the filtering results suffer from mutual interference
Solution Approach 1:
The patent segments the filtering rules into organized categories (risk levels, data types, source credentials) rather than using a single complex rule set. This segmentation reduces mutual interference by structuring rules hierarchically, where each category operates semi-independently, simplifying the overall system while maintaining comprehensive filtering capability.
4Loss of information
If the system processes all received raw data without filtering, then no data is lost, but the system consumes excessive computational resources and security risks increase
Solution Approach 1:
The patent applies partial action by selectively processing only the most critical aspects of raw data through multiple filtering stages. Rather than exhaustively analyzing every byte of data, the system performs targeted risk assessments and filtering on key parameters, achieving adequate data retention while significantly reducing computational resource consumption through selective processing.
Data Source
AI summary
Data collection system for effectively processing big data is provided. The data collection system includes multiple risk filtering modules up to third order or higher and a specific data extractor, wherein the multiple risk filtering modules and the specific data extractor are connected in series. The data collection system is capable of filtering received raw data through the multiple risk filtering modules so as to remove data with cyber security risks or system security issues, and keeping required data by the specific data extractor. In addition, the system can assist the user automatically to carefully select raw data through a combination of means of data classification, data normalization, and data clustering analysis. Thereby the system effectively enhances usability and security of data collection.


