Multi-Cloud Resource Orchestration for Big Data Jobs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The complexity of varying capabilities, configurations, and costs across different public cloud infrastructures makes it difficult to determine and assemble computation resources efficiently for data processing jobs, especially when accessing multiple cloud platforms, leading to inefficiencies and increased costs.
Innovation Solution
A server-based system that manages the selection, initiation, and termination of computation resources from multiple cloud providers based on job parameters, utilizing a matching process between computing applications, job schedulers, and utilization rates to configure and scale resources effectively.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple public cloud infrastructures are accessed for data processing jobs, then resource availability and flexibility improve, but system complexity and management difficulty increase
Solution Approach 1:
The patent introduces a cloud platform as an intermediary layer between users and multiple public cloud infrastructures. This platform abstracts the complexity of managing diverse cloud resources by providing unified resource orchestration, automated provisioning, and standardized interfaces, thereby maintaining high resource availability while reducing system complexity for users
Solution Approach 2:
The cloud platform implements universal resource management capabilities that work across multiple different public cloud infrastructures. By creating a multi-functional platform that can handle various cloud providers through common mechanisms, the system achieves versatility in resource access while maintaining consistent management approaches, thus reducing overall complexity
2Adaptability or versatility
If multiple public cloud infrastructures are accessed for data processing jobs, then resource flexibility improves, but management efficiency deteriorates
Solution Approach 1:
The cloud platform implements automated self-service mechanisms including automatic resource provisioning, dynamic allocation, and self-management of compute clusters across multiple cloud infrastructures. The system autonomously handles resource orchestration, scaling, and configuration without requiring manual intervention, thereby maintaining resource flexibility while significantly improving management efficiency
Solution Approach 2:
The platform performs preliminary actions by pre-configuring resource pools, pre-establishing management policies, and pre-orchestrating resource allocation strategies across multiple cloud infrastructures. This advance preparation enables flexible resource access while streamlining management operations, as the heavy lifting of resource coordination is done beforehand rather than in real-time
3Adaptability or versatility
If multiple public cloud infrastructures are accessed for data processing jobs, then compute resource options improve, but operational costs increase
Solution Approach 1:
The cloud platform dynamically changes operational parameters such as resource allocation ratios, cloud provider selection weights, and pricing models based on real-time conditions including cost fluctuations, resource utilization patterns, and job requirements. By continuously optimizing these parameters, the system maintains diverse compute resource options while minimizing operational costs through intelligent resource orchestration and automated cost management
Data Source
AI summary
Data processing approaches are disclosed that include receiving a configuration indicating a plurality of parameters for performing a data processing job, identifying available compute resources from a plurality of public cloud infrastructures, where each public cloud infrastructure of the plurality of public cloud infrastructures supports one or more computing applications, one or more job schedulers, and one or more utilization rates, selecting one or more compute clusters from one or more of the plurality of public cloud infrastructures based on a matching process between the parameters for performing the data processing job and a combination of the one or more computing applications, the one or more job schedulers, and the one or more utilization rates, and initiating the one or more compute clusters for processing the data processing job based on the selecting.


