Virtual Classes for Distributed Statistical Test Simulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for performing statistical tests require large data sets, often in the order of terabytes, which can be computationally intensive and time-consuming, especially for moderately complex tests, making it challenging to determine statistical significance efficiently.

Innovation Solution

A distributed processing system is used to simulate statistical tests, generating simulated data and distributing it across multiple nodes to reduce computational resources and time, employing a simulation control engine to coordinate task execution and a virtual software class to manage operations across the distributed system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional statistical testing methods are used, then accurate statistical significance can be determined, but massive data sets (terabytes) and extensive computational resources are required

Engineering Contradiction:
Improvestatistical significance accuracyVSAvoiddata set size
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent creates virtual copies of data sets through simulated data generation. Instead of requiring massive actual data sets, the system generates synthetic data that replicates the statistical properties and characteristics needed for accurate testing. This copying approach maintains measurement precision while dramatically reducing the quantity of physical data required.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system performs preliminary actions by pre-generating and storing simulated data sets before actual statistical testing is needed. The simulated data is created in advance with known statistical properties, allowing rapid execution of statistical tests without requiring collection and processing of massive real-world data sets at the time of testing.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If traditional statistical testing methods are used, then accurate results are obtained, but the process is time-consuming and computationally intensive

Engineering Contradiction:
Improvetest result accuracyVSAvoidtesting duration
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary statistical computations and generates simulated data sets in advance, before actual testing requirements arise. This pre-computation stores statistical properties and distributions that can be rapidly queried and applied during actual testing, dramatically reducing the time required while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the statistical testing process into distinct phases: simulated data generation, statistical property extraction, and actual hypothesis testing. By pre-computing and storing statistical properties from simulated data, the system separates the computationally intensive portions from the actual testing phase, reducing overall testing time while preserving result accuracy.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If complex statistical tests are performed, then comprehensive analysis is achieved, but computational resources and time requirements increase significantly

Engineering Contradiction:
Improvetest complexity capabilityVSAvoidtesting efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system creates virtual representations of complex data sets through simulated data generation. Instead of requiring actual massive data sets for complex tests, the system generates synthetic data that replicates the necessary statistical structures, correlations, and distributions. This allows comprehensive complex analysis while maintaining high productivity by avoiding the computational burden of processing actual terabyte-scale data sets.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11106486B2Techniques to manage virtual classes for statistical tests
Publication Date: 2021.08.31 SAS INSTITUTE INC
  • US11106486B2 patent drawing
  • US11106486B2 patent drawing
  • US11106486B2 patent drawing

AI summary

Techniques to manage virtual classes for statistical tests are described. An apparatus may comprise a simulated data component to generate simulated data for a statistical test, statistics of the statistical test based on parameter vectors to follow a probability distribution, a statistic simulator component to simulate statistics for the parameter vectors from the simulated data with a distributed computing system comprising multiple nodes each having one or more processors capable of executing multiple threads, the simulation to occur by distribution of portions of the simulated data across the multiple nodes of the distributed computing system, and a distributed control engine to control task execution on the distributed portions of the simulated data on each node of the distributed computing system with a virtual software class arranged to coordinate task and sub-task operations across the nodes of the distributed computing system. Other embodiments are described and claimed.