Geoscientific Data Analysis Using Order Distribution Grouping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing geoscientific data estimation methods using machine learning often ignore regions with high-value data, leading to low estimation accuracy due to data preprocessing techniques that recognize these regions as singular points or ignore them, thus neglecting important information.
Innovation Solution
A data analysis apparatus and method that aligns geoscientific data by size, groups it based on order distribution, and generates classification and regression models for each group using satellite data as explanatory variables, ensuring that high-value regions are included in the learning model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If data preprocessing techniques are used to remove or ignore high-value regions as singular points, then data processing simplicity is improved, but estimation accuracy deteriorates due to loss of important information
Solution Approach 1:
The patent extracts high-value regions from the training data and processes them separately through a dedicated extraction unit. Instead of removing these regions as conventional preprocessing does, the system isolates them to apply appropriate handling that preserves their important information while maintaining overall system simplicity.
Solution Approach 2:
Rather than removing high-value regions as conventional methods do, the patent inverts this approach by specifically extracting and preserving them. The inversion transforms the problem from filtering out anomalies to actively retaining and utilizing valuable extreme data points for improved estimation accuracy.
2Productivity
If conventional machine learning models are used that ignore high-value regions, then model training speed is improved, but information completeness deteriorates
Solution Approach 1:
The patent segments the training data into normal regions and high-value regions, processing them through different pathways. The extraction unit isolates high-value regions while the main training process handles常规数据, allowing both speed and completeness to be optimized through divided processing strategies.
Solution Approach 2:
The extraction unit acts as an intermediary component that bridges the gap between fast conventional training and complete information utilization. It prepares high-value region data in a way that allows efficient integration into the overall model training without compromising speed or completeness.
3Measurement precision
If all data including high-value regions are included in the learning model, then estimation accuracy is improved, but data processing complexity increases
Solution Approach 1:
The patent divides the data processing pipeline into distinct segments: an extraction unit for high-value regions, a main training unit for conventional data, and an integration mechanism. This segmentation manages complexity by organizing complex tasks into modular, manageable components with clear responsibilities.
Solution Approach 2:
The extraction unit performs preliminary processing of high-value regions before they are integrated into the main model training. This preliminary action prepares the data in advance, reducing the computational burden during main training and overall system complexity while ensuring accurate utilization of all data.
Data Source
AI summary
A data analysis apparatus 10 includes; an align unit 11 that acquires a pair data of a first data indicating a characteristic of a specific region and a second data corresponding to the first data and indicating another characteristic of the specific region, and aligns the first data in order of their sizes, a classification model generation unit that groups the pair data based on a characteristic of an order distribution of the first data after alignment, classifies the pair data, and generates a classification model for classifying the pair data using the classification result, a regression model generation unit that performs machine learning for each group, using the first data constituting the pair data and the second data constituting the same pair data, and generates a regression model indicating a relation with the first data and the second data.


