Coordinated ML Model Distribution for Renewable-Aware Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The high energy consumption and carbon footprint associated with training artificial intelligence and machine learning models, particularly due to the use of GPUs and TPUs, pose challenges for enterprises aiming to reduce their carbon footprint while maintaining efficient training processes across heterogeneous computing resource groups with varying energy sources and hardware capabilities.
Innovation Solution
A coordinated distribution and migration strategy for machine learning models using model profiling, sustainability services, and schedulers to optimize training across computing resource groups based on renewable energy availability, hardware capabilities, and checkpointing characteristics, minimizing carbon footprint and training time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine learning models are trained using GPUs and TPUs to improve training efficiency and speed, then productivity and training speed are improved, but energy consumption and carbon footprint increase significantly
Solution Approach 1:
The patent implements dynamic model migration between computing resource groups based on real-time power supply conditions. The system continuously monitors power sources and dynamically relocates ML models from high-consumption GPU/TPU resources to alternative computing resources when renewable energy availability changes, resolving the contradiction between maintaining high training productivity and reducing energy consumption.
Solution Approach 2:
The system changes the operational parameters of ML model training by adjusting which computing resource groups execute training based on power supply characteristics. By monitoring power source parameters (renewable vs. non-renewable) and dynamically reassigning models to different resource groups, the system optimizes the balance between training efficiency and energy sustainability.
2Object-generated harmful factors
If machine learning models are migrated between computing resource groups to utilize renewable energy sources, then carbon footprint is reduced, but training time may increase due to migration overhead
Solution Approach 1:
The system performs preliminary actions by pre-monitoring power supply conditions and proactively migrating ML models to computing resource groups with favorable renewable energy availability before carbon-intensive training begins. This advance preparation reduces the need for frequent migrations during training, thereby minimizing migration overhead and total training time while maintaining low carbon footprint.
Solution Approach 2:
The system implements feedback mechanisms by continuously monitoring power supply conditions, training progress, and model performance. This feedback loop enables intelligent decision-making about when to migrate models, ensuring that migrations occur only when beneficial for carbon reduction and that training time losses are minimized through optimized migration timing and selection of appropriate target resource groups.
3Adaptability or versatility
If computing resource groups with heterogeneous hardware capabilities are used to improve adaptability and resource utilization, then ease of operation and resource efficiency are improved, but system complexity increases
Solution Approach 1:
The patent introduces an intermediary scheduling system that manages the complexity of heterogeneous computing resource groups. This intermediary layer abstracts the heterogeneity by providing unified monitoring of power supply conditions, standardized model migration protocols, and centralized coordination of training tasks across diverse GPU, TPU, and alternative computing resources, thereby maintaining high adaptability while reducing operational complexity.
Data Source
AI summary
Methods are provided for coordinated distribution of machine learning models for sustainable training. The methods involve obtaining attributes of each machine learning model. The attributes include a training constraint and a computational requirement. The methods further involve obtaining power supply information about at least two computing resource groups. The power supply information relates to one or more power sources that supply power to the at least two computing resource groups. The methods further involve generating a deployment plan for training the machine learning models across the at least two computing resource groups based on the power supply information and the attributes. The deployment plan is configured to increase a use of the power from one or more renewable energy sources. The methods further involve distributing the machine learning models to the at least two computing resource groups for sustainable training of the machine learning models based on the deployment plan.


