Autonomous Mining Engine for Industrial Big Data Model Sets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data mining technologies face challenges in automating the process, integrating prior knowledge, and optimizing results, particularly in industrial applications, where supervised methods rely on single models and unsupervised methods are 'blind' and lack integration of knowledge.
Innovation Solution
An autonomous mining method for industrial big data based on model sets is introduced, which builds a mining engine using domain knowledge and structural characteristics of multi-source heterogeneous data, performs fault-tolerant estimation, and integrates knowledge for automatic mining and optimization through a fault-tolerant mining engine and VV&A testing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If supervised data mining methods are used to extract knowledge from datasets, then the mining process can achieve certain accuracy and reliability, but the process requires manual intervention to select model forms and data files, increasing operation complexity and reducing automation
Solution Approach 1:
The system enables self-service by automatically selecting appropriate model forms and data files through the mining engine, which autonomously performs model selection and parameter optimization without requiring manual intervention. The engine evaluates multiple models and automatically chooses the optimal configuration based on the dataset characteristics.
Solution Approach 2:
The mining engine is designed with multi-functionality to handle various data types and model forms universally. It can automatically adapt to different supervised data mining tasks (classification, evaluation, prediction) and select appropriate models from multiple candidates, making the system versatile across different industrial applications.
2Reliability
If multiple models are used to extract knowledge from datasets, then the quality and comprehensiveness of mining results improve, but the process becomes more complex and difficult to automate compared to single-model approaches
Solution Approach 1:
The system segments the complex task of multi-model selection into manageable components: the mining engine divides the evaluation process into individual model assessments, automatically tests each model's performance on the dataset, and selectively combines results from the most suitable models. This segmentation makes the multi-model process automated and tractable.
Solution Approach 2:
The model selection process is dynamic rather than static. The mining engine automatically adjusts which models to use based on the specific dataset characteristics and mining objectives. The system dynamically evaluates model performance and adapts the model combination strategy, making the complexity manageable through automated adaptation.
3Adaptability or versatility
If unsupervised data mining methods are used to find relationships in all attributes, then the system can explore data without prior knowledge, but the mining process becomes 'blind' and cannot guarantee result quality or integrate prior knowledge
Solution Approach 1:
The mining engine acts as an intermediary between unsupervised exploration and supervised knowledge integration. It first performs unsupervised pattern discovery to identify relationships in all attributes, then uses the engine to evaluate these patterns against domain knowledge and prior information, ensuring result quality while maintaining exploratory capability.
4Measurement precision
If manual model selection and parameter optimization are performed in data mining, then the mining process can achieve accurate results, but the workload and procedures become complicated, reducing productivity for large-scale data mining
Solution Approach 1:
The system replaces manual mechanical operations (manual model selection, parameter tuning, and evaluation) with an automated mining engine that uses algorithmic methods to perform the same functions. The engine automatically selects models and optimizes parameters through computational procedures, dramatically improving productivity while maintaining accuracy.
Data Source
AI summary
Disclosed is an autonomous mining method of industrial big data based on model sets, which comprises the following steps: S1, building model sets and a mining engine based on domain knowledge and structural characteristics of multi-source heterogeneous data; S2, carrying out data sampling on the multi-source heterogeneous data, and counting the fault-tolerant estimation of random error variance; S3, mining data sets by using the mining engine, and determining the optimal fault-tolerant model of each sampled data sequence and the optimal fault-tolerant estimation of model parameters; S4, performing goodness-of-fit statistics calculation and VV&A test by using the optimal fault-tolerant model; S5, acquiring data model representation and connotation knowledge based on model clustering. The method can realize the automation of the mining process of big data, the integration of associated knowledge, the expansion of model sets, the integration of mining and modeling and the optimization of mining results.


