Adaptive Sampling via Optimal Experimental Designs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face inefficiencies in extracting useful information from large data repositories due to the rapid growth of data volumes, where the ability to process and analyze all data becomes inefficient and often fails to keep pace.
Innovation Solution
An adaptive sampling process using systematic sampling procedures derived from optimal experimental designs to target specific observations of interest within large data sets, allowing for efficient information extraction by selecting a smaller, more diagnostic sample matrix that maximizes information value for analytic tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If all data within a large data repository is processed and analyzed, then complete information extraction is achieved, but computational time and resources become excessive and inefficient
Solution Approach 1:
The patent extracts and processes only a selected subset of data from the large data repository rather than processing all data. The system identifies and extracts specific observations that are most diagnostic for the analytic task, thereby achieving sufficient information extraction without the excessive computational burden of processing the entire dataset.
Solution Approach 2:
The patent applies partial action by processing a carefully selected portion of the data that provides maximum information value. The adaptive sampling process determines the optimal sample size and composition needed to achieve the analytic objectives, avoiding both insufficient sampling and excessive processing of unnecessary data.
2Quantity of substance
If data volumes grow rapidly, then data repository capacity increases, but the ability to process and analyze all data becomes inefficient and fails to keep pace
Solution Approach 1:
The system extracts a representative and diagnostic subset of observations from the growing data repository. By focusing computational resources on a carefully selected sample rather than attempting to process all available data, the system maintains processing efficiency even as data volumes increase rapidly.
Solution Approach 2:
The adaptive sampling process dynamically adjusts sampling parameters based on the analytic task requirements and data characteristics. This allows the system to optimize the balance between sample size and information content, maintaining processing capability as data volumes grow by changing the sampling strategy rather than attempting to process all data.
3Loss of time
If systematic sampling procedures from optimal experimental designs are used, then computational effort is reduced, but the sampling strategy complexity increases
Solution Approach 1:
The system performs preliminary actions by pre-determining the optimal sampling strategy based on the analytic task requirements and data characteristics. The adaptive sampling process establishes the sampling design beforehand, identifying which observations are most likely to provide diagnostic information, thereby reducing computational effort during the actual analysis phase.
Data Source
AI summary
A system, method, and computer-readable medium for extracting the samples from big data to extract most information about the relationships of interest between dimensions and variables in the data repository. More specifically, extracting information from large data repositories follows an adaptive process that uses systematic sampling procedures derived from optimal experimental designs to target from a large data set specific observations with information value of interest for the analytic task under consideration. The application of adaptive optimal design to guide exploration of large data repositories provides advantages over known big data technologies.


