Calculation power evaluation and resource allocation method for large language model reasoning
By identifying virtual operators and mathematical optimization algorithms, precise resource allocation of large language models in heterogeneous GPU environments is achieved, solving the problems of low resource utilization and insufficient deployment efficiency in existing technologies, and improving the running efficiency and adaptability of LLM.
Patent Information
- Application Number
- CN202511076738.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies lack the ability to dynamically evaluate and recommend optimal hardware combinations during the pre-launch phase of large language models (LLMs), resulting in low resource utilization, insufficient deployment efficiency, and difficulty in coping with the complexity of heterogeneous computing environments and rapid iterations of model architecture changes.
By identifying virtual operators, a computing power utilization judgment table is generated, and the hardware capabilities of heterogeneous GPU pools are collected in real time. Mathematical optimization algorithms are used for precise matching and recommendation, and finally the best-fit hardware combination is selected, so that LLM can accurately evaluate and optimize resource configuration before startup.
It improves resource utilization and throughput, reduces operating costs and energy consumption, enhances adaptability to new and future LLMs, simplifies the deployment process, and improves user experience and efficiency.
Smart Images

Figure CN120950357A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computational power evaluation for large language models, and more particularly to a method for computational power evaluation and resource allocation for large language model inference. Background Technology
[0002] The rapid development of Large Language Models (LLMs) and their widespread application across multiple fields have presented unprecedented challenges to underlying computing infrastructure. LLM deployment and inference tasks typically require large and high-performance graphics processing unit (GPU) clusters. However, existing technologies for LLM computational power assessment and resource allocation primarily focus on runtime optimization rather than accurate evaluation and recommendation during the pre-launch phase. This leads to low resource utilization and insufficient deployment efficiency.
[0003] Overview of existing computational power evaluation and resource allocation methods for large language model inference
[0004] Existing methods for evaluating LLM computing power and allocating resources mainly include the following categories:
[0005] Runtime Dynamic Scheduling and Load Balancing: Many existing systems focus on resource scheduling and optimization during the actual operation of an LLM. For example, schedulers for deep learning workloads in GPU data centers aim to minimize GPU fragmentation, reduce power consumption, and improve resource utilization.1 These schedulers typically handle online tasks and lack anticipation of future task arrivals. They adjust resource allocation through dynamic load balancing3 or real-time performance modeling4. For instance, the Bullet system dynamically configures GPU resources through real-time performance modeling to address the mismatch between computationally intensive prefill and memory-intensive decode phases, enabling concurrent execution and dynamic resource allocation.4 While these methods exhibit some optimization capabilities at runtime, they are inherently reactive, meaning adjustments are made only after problems occur. This can lead to the LLM running at a suboptimal configuration during initial startup and incurring additional adjustment overhead.
[0006] Distributed Inference and Model Parallelization: To address the challenge of limited memory capacity on a single GPU, existing solutions utilize techniques such as tensor parallelism, pipeline parallelism, quantization, and model offloading to distribute LLM inference tasks across multiple heterogeneous devices.5 For example, DigiMorphLab's patent utilizes mixed-integer linear programming (MILP) to optimize the allocation of distributed inference tasks on mobile devices, comprehensively considering model floating-point operations (FLOPs), memory constraints, device processing power, and network conditions.5 The Helix system models LLM inference computation on heterogeneous GPU clusters as a maximum flow problem and uses MILP to optimize model placement and request scheduling to aggregate globally distributed GPU resources.7 These methods address the memory bottleneck of large-scale model deployment, but their optimization goals primarily focus on the efficiency of distributed execution, rather than recommending the most suitable hardware combination for a specific LLM before startup.
[0007] Hardware Adaptation and Compiler Optimization: Existing deep learning frameworks typically include hardware adaptation modules responsible for converting the model's intermediate representation (IR) into executable instruction code for specific hardware (such as GPUs, CPUs, FPGAs, and ASICs). For example, Chinese patent CN114186678A proposes a deep learning hardware adaptation method that decouples the deep learning inference framework from heterogeneous hardware by defining subgraph operators and a unified context interface, thereby facilitating compilation and deployment. NNAdapter also provides a similar generic interface, acting as a bridge between the deep learning inference framework and the hardware. These methods focus on static adaptation at compile time or deployment time; once the model is compiled to specific hardware, its runtime environment is fixed, and dynamic hardware recommendations based on real-time computing power reserves cannot be made before startup.
[0008] Performance Modeling and Analysis Tools: Various tools exist in the industry for analyzing LLM training and inference performance. For example, NVIDIA Nsight Systems 11 and NVIDIA GenAI-Perf 12 can be used to identify performance bottlenecks and benchmark LLM inference performance metrics such as first-to-term time (TTFT), inter-to-term latency (ITL), and terms per second (TPS) 12. Lumos is a trace-based performance modeling tool that accurately predicts LLM behavior by building detailed execution graphs and allows simulation of performance under different configurations 16. NeuSight uses machine learning methods to predict the performance of unseen models and kernels on new GPUs 17. These tools primarily focus on analysis and optimization, providing performance data and predictive capabilities, but they do not directly provide a decision-making or recommendation mechanism for "computing the most suitable combination of hardware resources" from a heterogeneous GPU pool.
[0009] Challenges of Deep Learning Task Scheduling and Hardware Adaptation in Heterogeneous Computing Environments
[0010] Despite some progress in existing technologies, managing and optimizing LLM computing power in heterogeneous computing environments still faces multiple challenges:
[0011] Low resource utilization: Despite surging demand for GPUs, actual utilization of GPUs in data centers is generally low. The report shows that the average and median GPU utilization rates are only 52% and 10%, respectively.2 This is primarily due to the inherent resource mismatch between the computationally intensive prefill and memory-intensive decode phases in LLM inference,4 as well as over-provisioning of resources caused by dynamic request loads and diverse model characteristics.
[0012] Heterogeneity management is complex: Integrating heterogeneous hardware such as central processing units (CPUs), graphics processing units (GPUs), and field-programmable gate arrays (FPGAs)18 can significantly optimize AI performance and energy efficiency, but it also introduces enormous complexity in system design, software development, and runtime scheduling. The characteristics of different hardware architectures (such as FLOPs, memory bandwidth, and interconnect technologies) vary greatly, making it difficult to conduct unified computing power assessment, management, and optimization20.
[0013] Prediction and Adaptation Lag: While existing performance modeling tools can make performance predictions, they are mostly based on runtime tracing or benchmarking for specific scenarios. This results in a lack of a universal mechanism to accurately assess the computing power requirements of an LLM model on different heterogeneous hardware before it is launched. Therefore, it is difficult to select the optimal hardware combination before model deployment, often requiring trial and error or empirical judgment, which increases the uncertainty and cost of deployment.
[0014] Rapid iteration of new models and hardware: The scale of large language models is growing exponentially, and new model architectures and new hardware are constantly emerging.11 Existing systems struggle to adapt to this change quickly, requiring extensive and repetitive benchmarking and manual optimization, which significantly increases the cost and time of LLM deployment.
[0015] Limitations of existing technologies in dynamic pre-boot evaluation and optimal hardware combination recommendation
[0016] The core pain point addressed by this patent lies in the fundamental lack of existing technologies in terms of dynamic evaluation and optimal hardware combination recommendation capabilities during the LLM pre-startup phase.
[0017] Lack of dynamic pre-startup evaluation capabilities: Existing schedulers primarily allocate and adjust resources during LLM runtime, failing to dynamically predict and recommend optimal hardware before startup. This means LLM may start on suboptimal hardware, leading to initial performance bottlenecks and resource waste. The fundamental challenge of this runtime reactive optimization strategy lies in its inherent lag—addressing problems only after they occur, and the fact that dynamic adjustments themselves introduce additional overhead and complexity.
[0018] Insufficient versatility: Existing hardware adaptation methods focus on compiling models to specific hardware, rather than dynamic evaluation and recommendation. While performance modeling tools can predict performance, their output is typically performance metrics, not directly providing a decision on "computing the most suitable combination of hardware resources" from a heterogeneous GPU pool. A unified mechanism cannot be provided to accurately assess the computational requirements of a model on different heterogeneous hardware before model initiation.
[0019] Lack of a unified abstraction of "virtual operators": Existing technologies optimize at the operator level, but lack a unified concept of "virtual operators." Such a concept could abstract the computational characteristics of any LLM and build a universal computing power demand judgment table based on it, thereby enabling dynamic evaluation of "any" existing and future LLM.
[0020] The quantification and recommendation mechanism for "best-fit hardware" is lacking: While existing technologies offer hardware selection guidelines, these are mostly rules of thumb or static matrices, failing to provide precise mathematical matching and recommendations for "best-fit hardware" based on the dynamic computing power requirements of LLMs and the real-time computing power reserves of GPU pools. This results in users often having to rely on manual experience or repeated trial and error to select hardware in actual deployments, leading to inefficiency and difficulty in achieving optimal configurations. Summary of the Invention
[0021] In view of the above problems, the present invention is proposed to provide a method for computational power evaluation and resource allocation for large language model reasoning that overcomes or at least partially solves the above problems.
[0022] According to one aspect of the present invention, a method for computational power assessment and resource allocation for large language model inference is provided, the method comprising:
[0023] Step S101: Input the large language model to be evaluated;
[0024] Step S102: Parse the large language model and identify virtual operators;
[0025] Step S103: Analysis of the computing power requirements of the virtual operator;
[0026] Step S104: Generate a computing power usage determination table;
[0027] Step S105: Collect heterogeneous GPU pool hardware capabilities in real time and obtain heterogeneous hardware capability parameters;
[0028] Step S106: Standardize the heterogeneous hardware capability parameters;
[0029] Step S107: Mathematical matching of computing power requirements and hardware capabilities;
[0030] Step S108: Dynamically recommend and select the best compatible hardware combination;
[0031] Step S109: Launch the large language model based on the recommended optimal hardware combination.
[0032] Optionally, step S101: inputting the large language model to be evaluated specifically includes: the user or system taking the large language model to be deployed and evaluated for computing power requirements as input.
[0033] Optionally, step S102: parsing the large language model and identifying virtual operators specifically includes:
[0034] After receiving the LLM model input, the system parses it in a secure sandbox context and extracts the intermediate representation (IR) of its computation graph.
[0035] By using a predefined rule base and / or machine learning model, identify subgraphs or operation sequences in the model that correspond to specific computational patterns of LLM, and abstract and map them into unique virtual operators.
[0036] Optionally, step S103: the virtual operator computing power requirement analysis specifically includes:
[0037] For each identified virtual operator, a detailed analysis of computing power requirements is performed;
[0038] This includes estimating the number of floating-point operations (FLOPs) required for its execution, the demand for video memory and memory bandwidth, its inherent parallelism characteristics, its sensitivity to latency, and its power consumption characteristics under different loads.
[0039] Optionally, step S104: generating a computing power usage determination table specifically includes:
[0040] The computing power demand characteristics of the virtual operators obtained in step S103 are structured to generate a computing power utilization judgment table.
[0041] The computing power usage decision table records the detailed computing power requirement characteristics of each virtual operator under different LLM models or their different operating modes, serving as a standardized input for subsequent hardware matching.
[0042] Optionally, step S105: Real-time acquisition of heterogeneous GPU pool hardware capabilities to obtain heterogeneous hardware capability parameters specifically includes:
[0043] The system monitors and collects detailed capability parameters of all heterogeneous hardware in the GPU pool in real time, including GPU model, memory capacity, memory bandwidth, computing throughput, interconnect bandwidth, static energy consumption parameters, and most importantly, the current real-time load, idle memory, and dynamic computing power reserves of idle computing units.
[0044] Optionally, step S106: standardizing the heterogeneous hardware capability parameters specifically includes: standardizing the heterogeneous hardware capability parameters collected in step S105 and converting them into unified quantitative indicators.
[0045] Optionally, step S107: mathematical matching of computing power requirements and hardware capabilities specifically includes:
[0046] The LLM computing power requirement data generated in step S104 and the parameterized and standardized heterogeneous hardware capability data in step S106 are used as inputs to construct a multi-objective optimization function.
[0047] This optimization problem is solved using advanced mathematical optimization algorithms, enabling precise mapping and allocation between virtual operator requirements and hardware capabilities.
[0048] Optionally, step S108: dynamically recommending and selecting the best compatible hardware combination specifically includes:
[0049] Based on the optimization results of step S107, the system first generates a recommended combination sorting table, which contains multiple hardware combinations that meet the LLM computing power requirements and have high efficiency, and sorts them according to the comprehensive optimization objectives.
[0050] The platform obtains and selects the optimal hardware resources in real time from the recommended combination ranking table.
[0051] This invention provides a method for computational power evaluation and resource allocation for large language model inference. The method includes: Step S101: Inputting the large language model to be evaluated; Step S102: Parsing the large language model and identifying virtual operators; Step S103: Analyzing the computational power requirements of the virtual operators; Step S104: Generating a computational power usage judgment table; Step S105: Real-time acquisition of heterogeneous GPU pool hardware capabilities to obtain heterogeneous hardware capability parameters; Step S106: Standardizing the heterogeneous hardware capability parameters; Step S107: Mathematically matching computational power requirements with hardware capabilities; Step S108: Dynamically recommending and selecting the best suitable hardware combination; Step S109: Starting the large language model according to the recommended best suitable hardware combination. Through a dynamic evaluation mechanism, the computational power resources required by the model are accurately predicted, and the most suitable hardware combination is intelligently recommended and selected from the heterogeneous GPU resource pool accordingly. This realizes a paradigm shift from passive runtime adjustment to proactive pre-startup optimization.
[0052] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a flowchart illustrating a method for evaluating computing power and allocating resources for large language model reasoning, provided as an embodiment of the present invention. Detailed Implementation
[0055] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0056] The terms "comprising" and "having," and any variations thereof, in the specification, embodiments, claims, and drawings of this invention are intended to cover non-exclusive inclusion, such as including a series of steps or units.
[0057] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0058] like Figure 1 As shown, a method for evaluating computing power and allocating resources for large language model inference includes:
[0059] Step 101: Input the large language model to be evaluated. The user or system inputs the large language model whose computational requirements need to be deployed and evaluated. This model can be any existing or future LLM that can run on the platform.
[0060] Step 102: LLM Model Parsing and Virtual Operator Identification. After receiving the LLM model input, the system first parses it in a secure sandbox context to extract the intermediate representation (IR) of its computation graph. Based on this, through a predefined rule base and / or machine learning model (e.g., LLM can be used to generate rules to extract semantic concepts, or a multilayer perceptron (MLP) can be used to predict kernel utilization), the system identifies the subgraphs or operation sequences in the model that correspond to specific computational patterns of LLM, and abstracts and maps them into "virtual operators" unique to this invention.
[0061] Step 103: Virtual Operator Computational Requirements Analysis. For each virtual operator identified in Step 102, a detailed computational requirements analysis is performed. This includes estimating the number of floating-point operations (FLOPs) required for its execution, its requirements for video memory capacity and memory bandwidth, its inherent parallelism characteristics, its sensitivity to latency, and its power consumption characteristics under different loads.
[0062] Step 104: Generate a computing power utilization determination table. The computing power requirement characteristics of the virtual operators analyzed in Step 103 are structured to generate a "computing power utilization determination table." This table records in detail the computing power requirement characteristics corresponding to each virtual operator under different LLM models or their different operating modes, serving as standardized input for subsequent hardware matching.
[0063] Step 105: Real-time acquisition of heterogeneous GPU pool hardware capabilities. The system monitors and acquires detailed capability parameters of all heterogeneous hardware in the GPU pool in real time. This includes static parameters such as GPU model, memory capacity, memory bandwidth, computing throughput (TFLOPS), interconnect bandwidth, and power consumption, as well as the most critical dynamic computing power reserves such as current real-time load, idle memory, and idle computing units.
[0064] Step 106: Heterogeneous Hardware Capability Parameterization and Standardization. The heterogeneous hardware capability parameters collected in Step 105 are standardized and converted into unified quantitative indicators. This allows hardware of different types and from different manufacturers to be compared and evaluated on the same dimension, providing unified hardware-side data for subsequent mathematical matching.
[0065] Step 107: Mathematical Matching of Computing Power Requirements and Hardware Capabilities. Using the LLM computing power requirement data from the "Computing Power Usage Judgment Table" generated in Step 104, and the parameterized and standardized heterogeneous hardware capability data from Step 106 as input, a multi-objective optimization function is constructed. This function comprehensively considers objectives such as performance, cost, energy consumption, utilization rate, and SLO. Then, advanced mathematical optimization algorithms (such as MILP) are used to solve this optimization problem, achieving precise mapping and allocation between virtual operator requirements and hardware capabilities.
[0066] Step 108: Dynamically recommend and select the best hardware combination. Based on the optimization results of Step 107, the system first generates a recommended combination ranking table. This ranking table contains multiple hardware combinations that meet the LLM computing power requirements and are highly efficient, and is ranked according to the comprehensive optimization objectives. Subsequently, the platform can obtain and select the optimal hardware resources in real time from this ranking table to reflect the latest load and availability of hardware in the GPU pool.
[0067] Step 109: Start the large language model. Based on the optimal hardware combination recommended in Step 108, the platform starts the large language model to ensure it runs in the optimal computing environment.
[0068] 1. Definition of Virtual Operator: The virtual operator is a core concept in this invention used to abstract and quantify the computational features of large language models. It differs from the specific low-level operators in traditional deep learning frameworks (such as convolution and matrix multiplication), instead providing a higher-level, hardware-independent abstraction of key computational patterns within the LLM (e.g., QKV projection, self-attention computation, and multilayer perceptron (MLP) layers in the Transformer layer). Each virtual operator represents the computational characteristics and resource requirements of a certain type or series of operations during the inference process of the LLM. This abstraction enables the system to understand the computational requirements of different LLM models in a unified way.
[0069] 2. Methods for generating virtual operators:
[0070] LLM Model Parsing and Graph Structure Extraction: First, the input LLM model file is parsed to obtain its intermediate representation (IR) of the computational graph. This is similar to how existing hardware adaptation frameworks obtain model IRs.
[0071] Computational Pattern Recognition and Virtual Operator Mapping: Based on a predefined rule base and / or machine learning model (e.g., LLM can be used to generate rules to extract semantic concepts22, or a multilayer perceptron (MLP) can be used to predict kernel utilization17), subgraphs or operation sequences in the IR corresponding to specific computational patterns of the LLM are identified and mapped to corresponding virtual operators. For example, a complete Transformer layer can be abstracted into one or a few virtual operators instead of being decomposed into dozens of underlying CUDA kernels.
[0072] Virtual computing power requirement analysis: For each identified virtual operator, perform offline or lightweight online analysis of its computing power requirements.
[0073] Includes: Computational complexity (FLOPs): Estimated number of floating-point operations required to execute it 15.
[0074] Memory access patterns and bandwidth requirements: Estimate its requirements for video memory capacity (VRAM) and memory bandwidth (HBMbandwidth), especially key-value (KV) cache 15 and model weight loading 8.
[0075] Parallelism characteristics: Evaluate its inherent parallelism, such as whether it is suitable for large-scale data parallelism or model parallelism.
[0076] Delay sensitivity: Assess its sensitivity to delay. For example, the pre-padding stage may be sensitive to computational delay, and the decoding stage may be sensitive to single-step delay.
[0077] Power consumption characteristics: Estimated energy consumption under different loads 18.
[0078] Computing power utilization decision table generation: The above analysis results are structured to generate a "computing power utilization decision table". It records the detailed computing power requirement characteristics of each virtual operator under different LLM models or their different operating modes (such as different context lengths and batch sizes).
[0079] Table 1: Example of Virtual Operator Computing Power Usage Judgment
[0080]
[0081]
[0082] This computing power approach uses a decision table to standardize and quantify complex LLM computing power requirements, presenting them in a structured form. This serves as input for subsequent "mathematical matching," enabling the system to understand the resource footprint of any LLM in a unified and computable manner. This solves the problem of vague descriptions of LLM computing power requirements and the difficulty in directly correlating them with hardware capabilities in traditional methods, laying the foundation for precise matching.
[0083] Heterogeneous hardware capabilities are mathematically matched and the hardware pool with the highest current computing power reserves is selected.
[0084] Another key aspect of this invention is to mathematically match the computing power requirements of LLM with the real-time capabilities of a heterogeneous GPU pool, thereby selecting the best-fit hardware.
[0085] Heterogeneous hardware capability parameterization:
[0086] Hardware Information Acquisition: Real-time acquisition of detailed capability parameters for each heterogeneous hardware component in the GPU pool. These parameters include, but are not limited to: GPU model (e.g., NVIDIA A100, H100, L40S)20, VRAM capacity20, memory bandwidth (HBMBandwidth)20, computational throughput (TFLOPS, especially FP16 / FP8 Tensor Core performance)15, interconnect bandwidth (e.g., NVLink, PCIe)20, power consumption (TDP)20, specific operator performance (some hardware may be optimized for specific types of DL operators)17, and most importantly, real-time load and available resources (dynamically monitoring the actual load, idle VRAM, idle computing units, and other real-time computing power reserves of each GPU in the current hardware pool)2.
[0087] Parameter standardization: Standardize the heterogeneous parameters of different hardware into a unified quantitative index for mathematical comparison.
[0088] Table 2: Examples of heterogeneous hardware capability parameters
[0089]
[0090]
[0091] This table quantifies and standardizes the capabilities of heterogeneous and complex hardware, and updates its computing power reserves in real time. This provides accurate and actionable hardware-side data for subsequent mathematical matching, enabling the system to comprehensively consider various performance dimensions of the hardware and, in conjunction with real-time load information, make more accurate recommendations.
[0092] Mathematical matching algorithms and optimal choices:
[0093] Matching objective function construction: Construct a multi-objective optimization function that comprehensively considers the computing power requirements of LLM (from the decision table) and the capability parameters of heterogeneous hardware (from hardware information acquisition), as well as other optimization objectives, such as: minimizing inference latency13, maximizing throughput4, minimizing total cost of ownership (TCO) or energy consumption19, maximizing GPU utilization2, and satisfying specific service level objectives (SLO)1.
[0094] Optimization Algorithm: Advanced optimization algorithms are employed to solve this optimization problem. Possible algorithms include, but are not limited to, Mixed Integer Linear Programming (MILP), heuristic algorithms (such as greedy algorithms and genetic algorithms), or reinforcement learning. The algorithm mathematically maps and allocates virtual operator requirements to hardware capabilities, considering the performance of strategies such as model parallelism, data parallelism, and pipeline parallelism on different hardware combinations.
[0095] Dynamic Recommendation and Selection: Based on the optimization results, the system dynamically recommends and selects the hardware pool or combination with the highest current computing power reserves (i.e., the one that best meets the current LLM requirements and is most efficient). This process is real-time and reflects the latest load and availability of the hardware in the GPU pool.
[0096] Beneficial effects:
[0097] Significantly improves resource utilization and throughput for large language model inference:
[0098] This invention ensures that the model runs on the most suitable GPU combination by precisely matching its computational requirements with hardware capabilities before LLM startup. For example, for computationally intensive models, the system allocates GPUs with high TFLOPS; for memory-intensive models (such as those requiring large KV caches), it allocates GPUs with high VRAM and high bandwidth. This pre-optimization avoids bottlenecks caused by resource mismatches at runtime, thereby significantly improving the actual utilization of GPUs. In the prior art, the utilization rate of GPUs in data centers is generally low, with average and median utilization rates of only 52% and 10%, respectively. This is partly because models may run on suboptimal hardware. The precise matching of this patent means that the computational tasks of the LLM (represented by virtual operators) are allocated to the hardware units that can execute them most efficiently. For example, if a virtual operator of an LLM (such as the Attention layer) has limited memory bandwidth, the system will preferentially select a GPU with high HBM bandwidth. This "tailor-made" allocation ensures that the GPU's computing cores and memory bandwidth are fully utilized, reducing idle resources and directly improving overall resource utilization and the amount of inference tasks completed per unit time (throughput).
[0099] Effectively reduce operating costs and energy consumption:
[0100] By intelligently selecting the most suitable hardware combination, unnecessary hardware upgrades or over-configuration are avoided, reducing hardware procurement costs. Simultaneously, LLM tasks are scheduled to the most energy-efficient hardware, reducing unnecessary energy consumption. LLM deployment is costly and consumes a large amount of energy.
[0101] By optimizing hardware selection, tasks are assigned to the "most suitable" and "energy-efficient" hardware. This avoids using overly high-performance but underutilized hardware, or hardware with insufficient performance causing excessively long task execution times. By selecting hardware with the optimal performance / power ratio for a specific LLM task, unnecessary energy waste is significantly reduced. For example, for some inference tasks, a low-power edge NPU or CPU may be more cost-effective than a large GPU. This fine-grained management directly translates into reduced operating costs and the realization of green computing.
[0102] Enhance the platform's dynamic adaptability to new and future large language models:
[0103] By employing the abstraction layer of "virtual operators," the inherent computational characteristics of any LLM can be described, rather than merely the fixed parameters of a specific model. This means that even with entirely new LLM architectures or variants, as long as their core computational patterns can be mapped to a defined or scalable set of virtual operators, the system can accurately assess their computational power and recommend hardware.22 This significantly reduces the adaptation and tuning workload required for new models, improving the platform's flexibility and forward-looking capabilities.23 LLM models iterate rapidly, making it difficult for existing systems to quickly adapt to new models.17 Virtual operators, as a universal language for LLM computational characteristics, eliminate the need for the system to perform benchmarking and performance analysis from scratch for each new LLM model. As long as a new model can be decomposed or mapped to these virtual operators, the system can evaluate it using existing computational power decision tables and hardware matching logic. This abstraction mechanism enables the platform to have "dynamic adaptability" to "future" LLMs, significantly shortening the development-to-deployment cycle of new models and ensuring long-term operational efficiency and competitiveness.
[0104] Provides accurate and real-time hardware resource recommendations to optimize user experience:
[0105] Users receive explicit recommendations on the optimal hardware combination before even launching an LLM, eliminating guesswork and trial-and-error in hardware selection. The system considers the real-time computing power reserves of the GPU pool, ensuring that the recommended hardware is the most available and highest-performing available. This precise and real-time feedback greatly simplifies LLM deployment and management, improving user efficiency and satisfaction. Traditional hardware selection may be based on experience or limited benchmarking, leading users to spend considerable time trying different configurations or discovering poor performance after deployment. This patented pre-launch recommendation mechanism directly provides a mathematically optimized and real-time-considered optimal solution, significantly reducing the user's decision-making burden and trial-and-error costs. This "what you see is what you get" precise recommendation directly improves the efficiency and experience of deploying an LLM, allowing users to focus more on the application of the model itself rather than the complexity of the underlying infrastructure.
[0106] The above specific embodiments further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for evaluating computing power and allocating resources for reasoning in large language models, characterized in that, The computing power assessment and resource allocation methods include: Step S101: Input the large language model to be evaluated; Step S102: Parse the large language model and identify virtual operators; Step S103: Analysis of the computing power requirements of the virtual operator; Step S104: Generate a computing power usage determination table; Step S105: Collect heterogeneous GPU pool hardware capabilities in real time and obtain heterogeneous hardware capability parameters; Step S106: Standardize the heterogeneous hardware capability parameters; Step S107: Mathematical matching of computing power requirements and hardware capabilities; Step S108: Dynamically recommend and select the best compatible hardware combination; Step S109: Launch the large language model based on the recommended optimal hardware combination.
2. The method for computational power evaluation and resource allocation for large language model reasoning according to claim 1, characterized in that, Step S101: Inputting the large language model to be evaluated specifically includes: the user or system taking the large language model to be deployed and evaluated for computing power requirements as input.
3. The method for computational power evaluation and resource allocation for large language model reasoning according to claim 1, characterized in that, Step S102: Parsing the large language model and identifying virtual operators specifically includes: After receiving the LLM model input, the system parses it in a secure sandbox context and extracts the intermediate representation (IR) of its computation graph. By using a predefined rule base and / or machine learning model, identify subgraphs or operation sequences in the model that correspond to specific computational patterns of LLM, and abstract and map them into unique virtual operators.
4. The method for computational power evaluation and resource allocation for large language model reasoning according to claim 1, characterized in that, Step S103: The virtual operator computing power requirement analysis specifically includes: For each identified virtual operator, a detailed analysis of computing power requirements is performed; This includes estimating the number of floating-point operations (FLOPs) required for its execution, the demand for video memory and memory bandwidth, its inherent parallelism characteristics, its sensitivity to latency, and its power consumption characteristics under different loads.
5. The method for computational power evaluation and resource allocation for large language model reasoning according to claim 1, characterized in that, Step S104: Generating the computing power usage determination table specifically includes: The computing power demand characteristics of the virtual operators obtained in step S103 are structured to generate a computing power utilization judgment table. The computing power usage decision table records the detailed computing power requirement characteristics of each virtual operator under different LLM models or their different operating modes, serving as a standardized input for subsequent hardware matching.
6. The method for computational power evaluation and resource allocation for large language model reasoning according to claim 1, characterized in that, Step S105: Real-time acquisition of heterogeneous GPU pool hardware capabilities and obtaining heterogeneous hardware capability parameters specifically includes: The system monitors and collects detailed capability parameters of all heterogeneous hardware in the GPU pool in real time, including GPU model, memory capacity, memory bandwidth, computing throughput, interconnect bandwidth, static energy consumption parameters, and most importantly, the current real-time load, idle memory, and dynamic computing power reserves of idle computing units.
7. The method for computational power evaluation and resource allocation for large language model reasoning according to claim 1, characterized in that, Step S106: Standardizing the heterogeneous hardware capability parameters specifically includes: standardizing the heterogeneous hardware capability parameters collected in step S105 and converting them into unified quantitative indicators.
8. The method for computational power evaluation and resource allocation for large language model reasoning according to claim 1, characterized in that, Step S107: Mathematical matching of computing power requirements and hardware capabilities specifically includes: The LLM computing power requirement data generated in step S104 and the parameterized and standardized heterogeneous hardware capability data in step S106 are used as inputs to construct a multi-objective optimization function. This optimization problem is solved using advanced mathematical optimization algorithms, enabling precise mapping and allocation between virtual operator requirements and hardware capabilities.
9. The method for computational power evaluation and resource allocation for large language model reasoning according to claim 1, characterized in that, Step S108: Dynamically recommending and selecting the best compatible hardware combination specifically includes: Based on the optimization results of step S107, the system first generates a recommended combination sorting table, which contains multiple hardware combinations that meet the LLM computing power requirements and have high efficiency, and sorts them according to the comprehensive optimization objectives. The platform obtains and selects the optimal hardware resources in real time from the recommended combination ranking table.
Citation Information
Patent Citations
Hardware adaptation device and method based on deep learning
CN114186678A
Cited By
Construction method of reasoning simulation model, data processing method and related products
CN121880035A
GPU computing power consumption value and reasoning Token value-based computing power metering method and device for intelligent computing cloud platform
CN122332237A