Cloud EDA Job Scheduling Using ML Resource Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for managing electronic design automation (EDA) on cloud face challenges such as identifying the best cloud service provider, optimizing resource utilization, determining suitable jobs for cloud bursting, and predicting costs, leading to inefficiencies like costly burst patterns and server farm tilt due to static priority-based policies.
Innovation Solution
A method and system using machine learning models to predict optimal resource configurations and configuration circuits on cloud, evaluating different cloud service providers for lowest cost, calculating job completion and burst times, and determining wait times to deploy EDA jobs efficiently on-premises or on cloud, thereby overcoming inefficiencies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If cloud bursting is used to handle volume of design simulation jobs, then job throughput is improved, but cost increases and resource utilization becomes inefficient
Solution Approach 1:
The system dynamically adjusts resource allocation between on-premises server farm and cloud infrastructure based on real-time job queue conditions, utilization metrics, and cost parameters. The hybrid architecture enables flexible scaling where jobs are processed on-premises when resources are available and burst to cloud when needed, optimizing both throughput and cost efficiency
Solution Approach 2:
The system changes operational parameters by monitoring utilization thresholds, cost parameters, and job characteristics to determine optimal deployment locations. By adjusting these parameters dynamically, the system achieves cost-effective resource utilization while maintaining high job throughput through intelligent cloud bursting decisions
2Ease of operation
If static priority-based policies are used for job scheduling, then implementation is simple, but resource utilization becomes inefficient and server farm tilt occurs
Solution Approach 1:
The scheduling system transitions from static priority-based policies to dynamic multi-parameter decision-making that considers job characteristics, resource availability, cost parameters, and utilization metrics in real-time, optimizing resource allocation without sacrificing operational simplicity
Solution Approach 2:
The system implements feedback mechanisms by continuously monitoring job completion rates, resource utilization, and cost parameters, then using this information to adjust scheduling decisions and prevent server farm tilt, improving overall resource utilization efficiency
3Productivity
If cloud resources are used to meet design simulation demands, then processing capacity is improved, but identifying best fit cloud service provider becomes challenging
Solution Approach 1:
The system creates a universal interface layer that abstracts the complexity of multiple cloud service providers, enabling the hybrid architecture to work with different CSPs through standardized protocols and metrics, simplifying provider selection while maintaining access to diverse processing capacities
4Productivity
If server farm size is increased to handle job volume, then processing capacity is improved, but adaptability to varying job requirements decreases
Solution Approach 1:
The system segments processing capacity into on-premises server farm and cloud infrastructure components, allowing independent scaling and adaptation of each segment to match specific job requirements, thereby maintaining both high processing capacity and flexibility
Data Source
AI summary
Existing techniques of managing Electronic Design Automation (EDA) on cloud are based on pre-defined policies which result in costly burst patterns and server farm tilt. Embodiments of present disclosure overcomes these drawbacks by a method and system for managing EDA on cloud which employ machine learning to predict optimal resource configurations for deploying EDA jobs and configuration circuit on cloud that holds resources required by the optimal resource configuration. Further, different Cloud Service Providers (CSPs) are evaluated to determine the least cost CSP which has the desired configuration circuit. Completion time of jobs, time required to burst the jobs on cloud, and a pre-defined desired cycle time are calculated as a sum to determine the corresponding wait time for each job. The jobs are retained in the queue for corresponding wait time before deploying them on the cloud. The jobs are deployed on the on-prem infrastructure if resources are freed up before the wait time.


