Automated Data Analytics Lifecycle Automation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data analytics solutions face limitations in handling large data sets, including cost estimation and inflexibility, leading to costly and inadequate computing systems that struggle to manage growing data sets effectively.
Innovation Solution
An automated data analytics lifecycle method that defines an initial data analytic plan, conditions data, selects models, executes them, communicates results, and provisions computing resources to create a refined plan, allowing for flexible and cost-effective data analytics solutions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional data analytics solutions are used to handle large data sets, then data analysis can be performed, but the cost estimation becomes inadequate and the solution becomes inflexible
Solution Approach 1:
The system performs preliminary actions by automatically provisioning computing resources before data analytics execution and by pre-defining data analytics plans with cost estimates. This allows cost estimation to be performed in advance rather than retrospectively, making the solution both flexible and cost-certain.
Solution Approach 2:
The system implements feedback mechanisms where execution information from data analytics runs is fed back to automatically adjust and refine future resource provisioning decisions. This continuous feedback loop enables the system to learn from actual usage patterns and optimize both cost estimation and resource allocation over time.
2Productivity
If computing resources are provisioned to handle growing data sets, then data analytics capability is improved, but the system becomes costly and inflexible
Solution Approach 1:
The system dynamically provisions computing resources based on actual data analytics execution needs rather than static pre-allocation. The resource provisioning automatically adjusts up or down based on real-time demands, enabling the system to maintain high productivity while remaining flexible and cost-effective.
Solution Approach 2:
The system changes provisioning parameters dynamically based on execution information and performance metrics. By adjusting resource allocation parameters in response to actual usage patterns, the system optimizes both productivity and flexibility without being locked into fixed configurations.
3Ease of manufacture
If data analytics solutions are defined with fixed parameters, then implementation is simplified, but the solution cannot adapt to changing requirements
Solution Approach 1:
Data analytics plans are defined in advance with preliminary parameter specifications, providing implementation ease. However, the system maintains flexibility by allowing parameter adjustments based on execution feedback, combining the benefits of pre-planning with adaptive capability.
Solution Approach 2:
The system performs self-service by automatically adjusting data analytics plan parameters based on execution information without requiring manual reconfiguration. This enables the solution to adapt to changing requirements while maintaining implementation simplicity through automated self-optimization.
Data Source
AI summary
An initial data analytic plan for analyzing a given data set associated with a given data problem is defined. At least a portion of original data in the given data set is conditioned to generate conditioned data. At least one model is selected to analyze at least one of the original data and the conditioned data. The at least one selected model is executed on at least one of a portion of the original data and a portion of the conditioned data. Results of the model execution are communicated to at least one entity, the results comprising a refined data analytic plan for analyzing the given data set. One or more computing resources are provisioned to implement the refined data analytic plan. The defining, conditioning, selecting, executing, communicating and provisioning steps are performed on one or more processing elements associated with a computing system and automate a data analytics lifecycle.


