LLM Hardware Ranking for Pre-Execution Energy Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing tools for monitoring energy and carbon emissions of Large Language Models (LLMs) provide only high-level consumption views, are intrusive, lack compatibility, and generate emission values post-execution, limiting preemptive decision-making, and face challenges in power and processing resource consumption.
Innovation Solution
A method to estimate hardware requirements, processing time, and energy consumption for each LLM/hardware combination, enabling efficient ranking and selection of optimal configurations to minimize energy consumption and enhance operational efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing monitoring tools are used to track energy consumption, then carbon emission monitoring is provided, but the tools are intrusive and provide only high-level consumption views
Solution Approach 1:
The patent introduces an intermediary estimation system that calculates energy consumption and carbon emissions without requiring direct integration or modification of existing LLM infrastructure. This mediator layer provides precise monitoring by estimating consumption based on publicly available model specifications and hardware characteristics, avoiding the intrusiveness of existing tools while maintaining measurement accuracy.
Solution Approach 2:
The system performs preliminary estimation of energy consumption and carbon emissions before LLM execution by analyzing model architecture, parameter count, and hardware requirements. This allows carbon emission monitoring to be established in advance without requiring intrusive post-execution measurement tools, providing both precision and ease of operation.
2Measurement precision
If existing monitoring tools generate emission values after LLM execution, then carbon emissions are measured, but preemptive decision-making is restricted
Solution Approach 1:
The patent calculates energy consumption and carbon emission estimates before LLM execution by analyzing model specifications and hardware characteristics. This preliminary action enables organizations to make informed decisions about which LLMs to deploy and when, avoiding carbon-intensive models for non-critical tasks, thus eliminating the time loss associated with post-execution measurement while maintaining measurement precision.
Solution Approach 2:
The system segments the carbon emission measurement process into distinct phases: pre-execution estimation based on model specifications, during-execution tracking, and post-execution verification. This segmentation allows preemptive decision-making to occur at the estimation phase while maintaining accurate measurement through subsequent verification phases.
3Productivity
If comprehensive LLM/hardware combinations are evaluated, then optimal selection is achieved, but selection process complexity increases
Solution Approach 1:
The patent creates simplified copies or proxies for evaluating LLM/hardware combinations by using publicly available model specifications, benchmark performance data, and standardized hardware characteristics. Instead of actually deploying and testing each combination, the system uses these proxies to estimate energy consumption and performance, dramatically reducing selection process complexity while maintaining the ability to identify optimal configurations.
Solution Approach 2:
The system changes the evaluation parameters from requiring actual deployment and measurement to using estimable parameters such as model parameter count, architecture type, and hardware TDP ratings. This parameter transformation enables comprehensive evaluation of multiple LLM/hardware combinations without the complexity of actual deployment, improving selection efficiency while reducing process complexity.
Data Source
AI summary
Methods, systems, and computer-readable media for ranking large language models (LLMs). Input including list of LLMs, list of hardware and artificial intelligence (AI) prompt are provided by the user for ranking the LLMs. Based on the input, first estimating minimum number of hardware units needed to process AI prompt on each LLM/hardware combination and second estimating time to process the AI prompt using each LLM/hardware combination. Based on minimum number of hardware units and time to process AI prompt, third estimating amount of energy consumed by each LLM/hardware combination. Based on energy consumed, ranking LLM/hardware combinations for AI prompt. Based on ranking, selecting LLM and hardware, submitting AI prompt to LLM on hardware, and receiving response to submitted AI prompt from LLM.


