Small sample constructs big model inference energy consumption time delay map method, device and product
Patent Information
- Application Number
- CN202611014617.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-08-18
AI Technical Summary
[0004]本申请的主要目的在于提供一种小样本构建大模型推理能耗时延地图方法、设备及产品,旨在解决测量成本高、精准细分各计算内核能耗难、跨设备跨模型适配效果差的问题
[0017] One or more technical solutions proposed in this application have at least the following technical effects: defining a large model inference deployment stack and a multi-dimensional inference configuration space; based on the multi-dimensional inference configuration space, selecting representative measurement points covering the entire load range; according to the representative measurement points, synchronously collecting high-frequency execution information and low-frequency power sampling data to construct an initial high-frequency power trajectory; performing hardware sampling path aggregation on the initial high-frequency power trajectory, and combining it with the low-frequency power sampling data to perform trajectory calibration, obtaining a high-precision high-frequency power trajectory; based on the high-precision high-frequency power trajectory, performing fine-grained energy consumption attribution and kernel family decomposition, constructing a kernel family-level load scaling law small-sample generalization model; using the scaling law small-sample generalization model to complete the configuration space data, generating a complete energy consumption-latency map. Therefore, this application improves the accuracy of high-frequency energy consumption trajectory measurement by equivalent aggregation of hardware sampling paths and error calibration, realizes refined kernel-level energy consumption attribution, completes configuration space data by relying on small sample scaling law modeling, reduces the cost of traditional full exhaustive measurement, improves the generalization ability under different hardware and inference scenarios, and efficiently and accurately completes the construction of large model inference energy consumption-latency map, providing reliable technical support for inference energy-saving scheduling and configuration optimization.
Smart Images

Figure CN122594742A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of energy consumption perception technology for large model inference scheduling, and in particular to a method, device and product for constructing a large model inference energy consumption delay map with a small sample. Background Technology
[0002] With the rapid iteration of artificial intelligence technology and the large-scale deployment of large-model inference services, large language model inference services have become the fastest-growing computing load in data centers. Inference services need to continuously respond to massive online requests over long periods of time, while also considering throughput, latency metrics, and service level constraints. The continuous expansion of model size, context length, and concurrent clusters has made the GPU energy consumption, cooling costs, and carbon emission pressures of inference scenarios increasingly prominent.
[0003] To address the energy consumption and latency issues in large model inference, the industry needs to optimize inference configurations and schedule energy-saving operations based on energy consumption-latency maps. However, current mainstream construction methods have significant shortcomings. Relying on exhaustive full-configuration measurement is extremely costly and cannot achieve fine-grained kernel-level energy consumption attribution, making it difficult to construct configuration space energy consumption-latency maps at low cost and with high accuracy. Summary of the Invention
[0004] The main purpose of this application is to provide a method, device and product for constructing energy consumption and latency maps of large model inference using small samples, aiming to solve the problems of high measurement costs, difficulty in accurately subdividing the energy consumption of each computing kernel, and poor cross-device and cross-model adaptation.
[0005] To achieve the above objectives, this application proposes a method for constructing a large-scale model inference energy consumption and latency map using small samples, the method comprising: Define the large model inference deployment stack and the multidimensional inference configuration space; Based on the multidimensional inference configuration space, representative measurement points covering the entire load range are selected; Based on the representative measurement points, high-frequency execution information and low-frequency power sampling data are collected synchronously to construct an initial high-frequency power trajectory. Hardware sampling path aggregation is performed on the initial high-frequency power trajectory, and trajectory calibration is performed in combination with the low-frequency power sampling data to obtain a high-precision high-frequency power trajectory; Based on the high-precision high-frequency power trajectory, fine-grained energy consumption attribution and kernel family decomposition are performed to construct a kernel family-level load scaling law small-sample generalization model. The scaling law small sample generalization model is used to complete the configuration space data and generate a complete energy consumption-delay map.
[0006] In one embodiment, the step of selecting a small number of representative measurement points covering the entire load range based on the multidimensional inference configuration space includes: Representative configurations covering low, medium, and high load ranges are selected from the multidimensional inference configuration space using stratified sampling, boundary point sampling, active learning, uncertainty sampling, or historical experience selection as representative measurement points.
[0007] In one embodiment, the step of simultaneously acquiring high-frequency inference execution information and low-frequency power sampling data based on the representative measurement points to construct an initial high-frequency power trajectory includes: Fine-grained high-frequency inference execution information and low-frequency power sampling data are collected from the representative measurement points. The hardware metrics in the fine-grained high-frequency inference execution information are aligned to the same time axis to obtain time-synchronized standardized hardware feature data. Based on the standardized hardware feature data, a high-frequency power estimation model is constructed by combining the device operating frequency, hardware temperature, and inference execution status. The initial high-frequency power trajectory is output based on the high-frequency power estimation model.
[0008] In one embodiment, the step of performing hardware sampling path aggregation on the initial high-frequency power trajectory and combining it with the low-frequency power sampling data to perform trajectory calibration to obtain a high-precision high-frequency power trajectory includes: The initial high-frequency power trajectory is equivalently transformed according to the native sampling processing mechanism of the low-frequency power interface to generate a comparable sequence that matches and compares with the low-frequency power sampling data; The error between the comparable sequence and the real low-frequency power sampling data is calculated, and a loss function is constructed by introducing a regularization constraint term. The parameters of the high-frequency power estimation model are then calibrated to obtain the calibrated high-frequency power estimation model. Based on the calibrated high-frequency power estimation model, a high-precision high-frequency power trajectory is obtained.
[0009] In one embodiment, the step of performing an equivalent transformation on the initial high-frequency power trajectory according to the native sampling processing mechanism of the low-frequency power interface to generate a comparable sequence that matches and compares with the low-frequency power sampling data includes: The hardware mechanisms of window averaging, filtering, time delay, and smoothing of the native sampling of the low-frequency power interface are replicated to obtain the replicated native sampling processing mechanism of the low-frequency power interface. Using the replicated low-frequency power interface native sampling processing mechanism as input, the initial high-frequency power trajectory is equivalently converted into a comparable sequence matching the low-frequency power sampling data through the sampling path aggregation operator.
[0010] In one embodiment, the step of calculating the error between the comparable sequence and the actual low-frequency power sampling data, and introducing regularization constraints to calibrate the high-frequency power estimation model parameters to obtain the calibrated high-frequency power estimation model includes: Calculate the error between the comparable sequence and the low-frequency power sampling data; Based on the calculated error, a loss function with regularization constraints is constructed using the low-frequency true power sampling value as a benchmark. The high-frequency power estimation model is obtained by iteratively calibrating the parameters of the high-frequency power estimation model by minimizing the loss function.
[0011] In one embodiment, the step of performing fine-grained energy consumption attribution and kernel family decomposition based on the high-precision high-frequency power trajectory to construct a kernel family-level load scaling law small-sample generalization model includes: The high-precision high-frequency power trajectory is aligned with the computation kernel timestamp, inference stage boundary, request window, batch processing window and kernel family label to obtain aligned power trajectory data carrying multi-dimensional inference information. Based on the aligned power trajectory data, power integral calculations are performed on various time-series windows to obtain fine-grained energy consumption labels corresponding to single kernel, kernel family, inference stage, single request, and single batch. Based on the fine-grained energy consumption labels, kernel family decomposition and scaling law modeling are performed to construct a kernel family-level load scaling law small-sample generalization model.
[0012] In one embodiment, the step of performing power integral calculations on various time-series windows based on the aligned power trajectory data to obtain fine-grained energy consumption labels corresponding to single kernels, kernel families, inference stages, single requests, and single batches includes: Based on the aligned power trajectory data, time-series integration windows of different durations are defined; Perform power integration calculations for each time-series integration window and output fine-grained energy consumption labels for single kernel, kernel family, inference stage, single request, and single batch.
[0013] In one embodiment, the step of performing kernel family decomposition and scaling law modeling based on the fine-grained energy consumption label to construct a kernel family-level load scaling law small-sample generalization model includes: Based on the fine-grained energy consumption labels, computation kernel names, operator semantics, execution features, and trajectory metadata, the inference kernels are clustered and split to obtain multiple original kernel groups; The original kernel groups were categorized and organized to identify typical kernel families, including attention, matrix multiplication, key-value caching, normalization, activation function, and rotation position encoding. The load variation patterns of similar computations were unified to obtain standardized kernel family classification results. Based on the standardized kernel family classification results, an energy consumption-latency scaling law model is constructed for the pre-filling, decoding inference stages and each kernel family. The energy consumption-latency scaling law model includes load-related shared slope coefficients and deployment-specific intercept parameters. The shared slope coefficients, which are universal across hardware and models, are reused to infer the inherent laws of the load. The deployment-specific intercept parameters are calibrated using a small amount of target deployment measurement point data. Combined with the two parameters, a scaling law small sample generalization model is obtained.
[0014] In one embodiment, the step of using the scaling law small-sample generalization model to complete the configuration space data and generate a complete energy consumption-latency map includes: Based on the scaling law small sample generalization model, the power consumption and latency of each kernel family under unmeasured configuration are predicted, and the power consumption and latency data of the complete inference task are obtained by accumulating layer by layer, thus completing the configuration space data. Generate a complete energy consumption-latency map covering the configuration space based on the configuration space data.
[0015] Furthermore, to achieve the above objectives, this application also proposes a device for constructing a large model inference energy consumption and latency map using a small sample size. The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the method for constructing a large model inference energy consumption and latency map using a small sample size as described above.
[0016] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the method for constructing a large model inference energy consumption and latency map using small samples as described above.
[0017] One or more technical solutions proposed in this application have at least the following technical effects: defining a large model inference deployment stack and a multi-dimensional inference configuration space; based on the multi-dimensional inference configuration space, selecting representative measurement points covering the entire load range; according to the representative measurement points, synchronously collecting high-frequency execution information and low-frequency power sampling data to construct an initial high-frequency power trajectory; performing hardware sampling path aggregation on the initial high-frequency power trajectory, and combining it with the low-frequency power sampling data to perform trajectory calibration, obtaining a high-precision high-frequency power trajectory; based on the high-precision high-frequency power trajectory, performing fine-grained energy consumption attribution and kernel family decomposition, constructing a kernel family-level load scaling law small-sample generalization model; using the scaling law small-sample generalization model to complete the configuration space data, generating a complete energy consumption-latency map. Therefore, this application improves the accuracy of high-frequency energy consumption trajectory measurement by equivalent aggregation of hardware sampling paths and error calibration, realizes refined kernel-level energy consumption attribution, completes configuration space data by relying on small sample scaling law modeling, reduces the cost of traditional full exhaustive measurement, improves the generalization ability under different hardware and inference scenarios, and efficiently and accurately completes the construction of large model inference energy consumption-latency map, providing reliable technical support for inference energy-saving scheduling and configuration optimization. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A flowchart illustrating the first embodiment of the method for constructing a large model inference energy consumption and latency map from a small sample in this application; Figure 2 A flowchart illustrating the overall architecture of the energy consumption-latency map method for building large-model inference with small samples in this application; Figure 3 This is a flowchart of the high-frequency power trajectory reconstruction process in this application; Figure 4 This is a timeline diagram showing the energy consumption attribution in this application; Figure 5 This is a schematic diagram of the shared slope calibration in this application; Figure 6 This is a schematic diagram illustrating the recovery of a small sample energy consumption-delay map in this application; Figure 7 Configure the selection and incremental calibration closed-loop diagram for this application; Figure 8 A schematic diagram of the module structure of the device for constructing a large model inference energy consumption and time delay map using small samples in the embodiments of this application; Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the method of building a large model inference energy consumption and latency map from a small sample in this application embodiment.
[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0024] The main solution of this application embodiment is to define a large model inference deployment stack and a multi-dimensional inference configuration space; based on the multi-dimensional inference configuration space, select representative measurement points covering the entire load range; according to the representative measurement points, synchronously collect high-frequency execution information and low-frequency power sampling data to construct an initial high-frequency power trajectory; perform hardware sampling path aggregation on the initial high-frequency power trajectory, and perform trajectory calibration in combination with the low-frequency power sampling data to obtain a high-precision high-frequency power trajectory; based on the high-precision high-frequency power trajectory, perform fine-grained energy consumption attribution and kernel family decomposition to construct a kernel family-level load scaling law small-sample generalization model; use the scaling law small-sample generalization model to complete the configuration space data and generate a complete energy consumption-latency map. Therefore, this application achieves equivalent power trajectory conversion by replicating the native hardware sampling mechanism and improves the accuracy of fine-grained power consumption measurement by combining loss function calibration, which can accurately decompose the energy consumption contribution of different kernel families; it performs small sample modeling based on a small number of representative measurement points, eliminating the need for exhaustive testing of all configurations, thus reducing the computing power and power consumption of graphics processor testing; by mining the load scaling patterns of inference kernels, it improves the generalization prediction capability across hardware, models, and inference configurations, and can construct a configuration space energy consumption-latency map at low cost and high accuracy, providing effective support for energy-saving configuration selection and energy consumption-aware scheduling for large model inference.
[0025] This application embodiment takes into account that existing energy consumption-latency map construction schemes have obvious technical shortcomings, making it difficult to simultaneously meet the actual deployment requirements of low-cost measurement, fine-grained energy consumption analysis, and cross-scenario generalization. The specific defects are as follows: High measurement costs: Traditional solutions require exhaustive testing of all inference configurations. Measurement overhead increases significantly with the size of the configuration space, consuming a large amount of GPU computing power and power resources. Furthermore, after hardware, model, and inference engine iterations and updates, a full retest is required, resulting in high iteration costs.
[0026] Lack of fine-grained power consumption attribution capability: Low-frequency power sampling data from conventional devices undergoes internal hardware filtering, window averaging, and delay smoothing, and can only output device-level average power consumption. It cannot capture fine power consumption changes during short-term inference, making it difficult to achieve fine-grained power consumption breakdown and attribution at the kernel family and operator levels. External high-precision power acquisition devices are complex to deploy and costly, making it impossible to scale up for online inference scenarios.
[0027] Poor cross-scenario generalization ability: Traditional end-to-end fitting schemes cannot distinguish between kernel load scaling rules and kernel combination ratio changes, resulting in coarse modeling granularity; pure theoretical analysis schemes rely on a large number of dedicated hardware parameters and inference engines to implement details, making it difficult to adapt to new deployment environments, and resulting in insufficient prediction accuracy across devices, models, and configurations.
[0028] The practicality of the engineering implementation is weak: the full-scale test scheme has high testing costs and contradicts the goal of energy-saving inference; the theoretical analysis model has complex parameters and high debugging threshold, making it difficult to quickly adapt to the mainstream inference service architecture and support the need for rapid construction of energy consumption latency maps for large-scale online inference scenarios.
[0029] The technical solution provided in this application effectively solves the problems of high measured cost, difficulty in fine-grained energy consumption attribution, and weak cross-scenario generalization ability in the process of constructing large model inference energy consumption latency maps.
[0030] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device, a device or system for building a large model inference energy consumption and latency map from a small sample, etc. The following description uses a system for building a large model inference energy consumption and latency map from a small sample as an example to illustrate this embodiment and the following embodiments.
[0031] Based on this, embodiments of this application provide a method for constructing a large model inference energy consumption and latency map using a small sample size, referring to... Figure 1 , Figure 1 A flowchart illustrating the first embodiment of the method for constructing a large model inference energy consumption delay map using small samples in this application.
[0032] In this embodiment, the method for constructing a large model inference energy consumption and latency map using small samples includes steps S10 to S60. The following provides a detailed explanation of each step.
[0033] like Figure 1 As shown, the first embodiment of this application proposes a method for constructing a large model inference energy consumption and latency map using small samples, the method comprising: Step S10: Define the large model inference deployment stack and the multidimensional inference configuration space; Specifically, the combination of inference engine, AI accelerator, and large language model is defined as an independent inference deployment stack. Each uniquely matched inference engine, acceleration hardware, and large language model constitutes an independent deployment stack. Based on the deployment stack, candidate inference configurations are formed. The configuration information includes the corresponding deployment stack identifier, batch size, input length, and output length. In practical application scenarios, various dimensions that affect inference energy consumption and latency, such as the number of concurrent requests, hardware operating frequency, power limit, quantization accuracy, parallel execution method, cache scheduling strategy, and sampling parameters, can also be added.
[0034] Therefore, based on the standardized division of different inference operating environments, a complete configuration space covering multiple types of operating constraints and load characteristics is built, which defines a clear boundary range for subsequent selection of representative measurement points and the implementation of experimental measurement and modeling work.
[0035] Step S20: Based on the multidimensional inference configuration space, select representative measurement points covering the entire load range; Specifically, instead of exhaustively measuring all candidate configurations within the configuration space, a small number of representative configurations are selected from the complete multidimensional configuration space as actual measurement points. The selected measurement points can simultaneously cover all ranges of low, medium, and high load. The load in the pre-filling stage can be approximated by the product of batch size and input length, while the load in the decoding stage can be measured by load metrics such as the product of batch size, input length, and output length, or the cumulative context scan volume. Measurement point selection can employ any combination of fixed rules, stratified sampling, boundary point sampling, active learning, uncertainty sampling, and historical experience selection.
[0036] The purpose of this step is to cover various load conditions with a small number of measurement points, significantly reduce the workload of actual measurement, reduce hardware computing power and power consumption, and at the same time retain the complete load change characteristics within the configuration space, so as to provide sufficient and effective measured data for subsequent small sample modeling.
[0037] Step S30: Based on the representative measurement points, synchronously collect high-frequency execution information and low-frequency power sampling data to construct an initial high-frequency power trajectory; Specifically, inference tests were conducted for each representative measurement configuration, simultaneously collecting two types of data: firstly, fine-grained high-frequency execution information, including the computing kernel name, kernel start and end timestamps, inference stage boundaries, request window boundaries, activity of various computing cores, total memory read / write volume, cache access volume and hit rate at each level, hardware operating frequency, chip temperature, and batch processing running status; secondly, low-frequency power sampling data, taken from the hardware power reading interface opened by the device's underlying management driver. All high-frequency hardware indicators were aligned to the same timeline to build a high-frequency power estimation model. The model input included high-frequency hardware indicators, real-time operating frequency, chip temperature, and inference execution status. Linear models, multinomial models, piecewise models, tree models, neural networks, or physical heuristic models were used as power consumption fitting functions, and the output was the instantaneous power value under continuous time series, which was combined to generate a complete initial high-frequency power trajectory.
[0038] Therefore, fine-grained hardware runtime timing data and device-level real power consumption observation data are acquired simultaneously. Based on timing alignment modeling, a high-time-resolution continuous power consumption change curve is generated, which makes up for the shortcomings of low-frequency power sampling timing granularity being too coarse and unable to capture short-term inference dynamic power consumption changes. This provides complete basic timing data for subsequent hardware sampling path aggregation and model calibration.
[0039] Step S40: Perform hardware sampling path aggregation on the initial high-frequency power trajectory and perform trajectory calibration in combination with the low-frequency power sampling data to obtain a high-precision high-frequency power trajectory. Specifically, the low-frequency power sample value output by the device is not the instantaneous actual power consumption, but rather a statistical result after internal hardware sampling filtering, smoothing averaging, time delay, and window aggregation. This step does not directly compare the high-frequency instantaneous predicted value with the low-frequency sample point. Instead, it uses a preset sampling path aggregation operator to perform moving average, exponential smoothing, time delay compensation, integral averaging, and sampling window alignment operations on the initial high-frequency power trajectory, converting the high-temporal-resolution high-frequency power trajectory into a comparable sequence matching the low-frequency sampling time series. Subsequently, using the low-frequency actual power sample value as a benchmark, the high-frequency power estimation model is calibrated overall by minimizing the overall error between the aggregated high-frequency prediction sequence and the low-frequency power sample data, and by iteratively optimizing the model parameters with regularization constraints, resulting in a high-precision high-frequency power trajectory.
[0040] This eliminates timing deviations and system errors caused by hardware sampling mechanisms, allowing the high-frequency power trajectory to be highly matched with the actual power consumption observation results of the device after being mapped through a unified sampling path. Ultimately, it outputs a high-precision high-frequency power trajectory with accurate timing and controllable errors, providing a reliable data foundation for subsequent fine-grained energy consumption integration and kernel attribution.
[0041] Step S50: Based on the high-precision high-frequency power trajectory, perform fine-grained energy consumption attribution and kernel family decomposition to construct a kernel family-level load scaling law small-sample generalization model. Specifically, high-precision, high-frequency power trajectories are precisely aligned with the computation kernel timestamps, inference stage boundaries, request windows, and kernel family labels. Energy consumption is calculated for different windows using time-series integration, resulting in fine-grained energy consumption labels for single kernels, single inference stages, and single requests. Simultaneously, based on kernel execution characteristics, operator semantics, and trajectory metadata, the underlying computation kernels are divided into various kernel families, such as attention, matrix computation, cache read / write, and normalization, ensuring consistent computation and memory access patterns for kernels of the same type. Building upon this, a kernel family-level energy consumption and latency scaling law model, incorporating load slope coefficients and deployment intercept terms, is constructed for each inference stage and kernel family, forming a small-sample generalization model with generalization capabilities.
[0042] This enables refined energy consumption tracing and kernel regularization classification. By using a structured scaling law model to uncover stable load change patterns and modeling based on a small number of measured samples, the predictive generalization ability of the model across configurations and deployment environments is effectively improved.
[0043] Step S60: Use the scaling law small sample generalization model to complete the configuration space data and generate a complete energy consumption-delay map.
[0044] Specifically, a kernel family-level load scaling law small-sample generalization model is used, reusing the load slope parameter common to different deployment stacks, and calibrating the target deployment stack's specific intercept parameter with a small amount of measured data to predict energy consumption and latency for massive unmeasured configurations in the configuration space. By accumulating the prediction components of each kernel family, energy consumption and latency data for the complete inference process are obtained, ultimately forming an energy consumption-latency map that covers the configuration space and can accurately map the correspondence between inference configurations and energy consumption and latency.
[0045] Therefore, a full-dimensional energy consumption-latency map is constructed based on a small number of representative measurement points, reducing the cost of exhaustive full measurement while ensuring prediction accuracy and cross-scenario generalization ability, providing accurate and comprehensive data support for energy-saving configuration optimization and energy consumption perception scheduling for large model inference.
[0046] The main solution of this application embodiment is as follows: Define a large model inference deployment stack and a multi-dimensional inference configuration space; based on the multi-dimensional inference configuration space, select representative measurement points covering the entire load range; according to the representative measurement points, synchronously collect high-frequency execution information and low-frequency power sampling data to construct an initial high-frequency power trajectory; perform hardware sampling path aggregation on the initial high-frequency power trajectory and combine it with the low-frequency power sampling data for trajectory calibration to obtain a high-precision high-frequency power trajectory; based on the high-precision high-frequency power trajectory, perform fine-grained energy consumption attribution and kernel family decomposition to construct a kernel family-level load scaling law small-sample generalization model; use the scaling law small-sample generalization model to complete the configuration space data and generate a complete energy consumption-latency map. Therefore, this application achieves equivalent power trajectory conversion by replicating the native hardware sampling mechanism and improves the accuracy of fine-grained power consumption measurement by combining loss function calibration, which can accurately decompose the energy consumption contribution of different kernel families; it performs small sample modeling based on a small number of representative measurement points, eliminating the need for exhaustive testing of all configurations, thus significantly reducing the computing power and power consumption of graphics processor testing; at the same time, by mining the load scaling pattern of inference kernels, it improves the generalization prediction capability across hardware, models, and inference configurations, and can construct a configuration space energy consumption-latency map at low cost and high accuracy, providing effective support for energy-saving configuration selection and energy consumption-aware scheduling for large model inference.
[0047] This embodiment provides a complete method for constructing a large-scale model inference energy consumption-latency map, encompassing multi-dimensional configuration space construction, representative measurement point selection, synchronous acquisition of high and low frequency data, high-frequency power trajectory calibration, fine-grained energy consumption attribution, and small-sample generalization completion. This method optimizes the accuracy of high-frequency power consumption measurement through hardware sampling path aggregation and loss function calibration, achieving refined kernel family energy consumption decomposition. It completes the full configuration data through small-sample modeling based on a small number of representative measurement points, generating a high-precision energy consumption-latency map at low cost. This forms an automated energy consumption measurement and configuration optimization system adaptable to multiple hardware and models, overcoming many shortcomings of traditional solutions, such as high actual measurement costs, difficulty in fine-grained energy consumption attribution, weak cross-scenario generalization ability, difficulty in large-scale deployment, and poor engineering feasibility.
[0048] In one feasible implementation, step S10, defining the large model inference deployment stack and the multidimensional inference configuration space, includes steps S101-S102: Step S101: Define the standardized large model inference deployment stack.
[0049] Specifically, this system is deployed on a cloud-based large language model inference service platform, and an inference deployment stack is defined as follows: ,in Indicates the reasoning engine. Indicates GPU or other AI accelerators. It represents a large language model; multiple deployment stacks with different combinations can exist within the platform, supporting heterogeneous deployment scenarios where the same model adapts to multiple GPUs and different models share the same type of GPU.
[0050] Therefore, the platform can standardize the division of various inference runtime environments, distinguish the boundaries of different runtime environments, and provide a clear basis for building an energy consumption-latency map for each deployment stack.
[0051] Step S102: Construct a multidimensional inference configuration space.
[0052] Specifically, a corresponding candidate inference configuration space is built for each deployment stack. A candidate inference configuration space can be represented as follows: ,in: Indicates the reasoning engine. Indicates GPU or other AI accelerators. Representing a large language model, For batch size, For the input length, The output length can be configured as follows: concurrent request count, frequency, power limit, quantization accuracy, parallel mode, key-value caching strategy, sampling parameters, or other parameters that affect energy consumption and latency. For example, the batch size can be set to 1, 2, 4, 8, 16, 32, or 64, the input length can be divided into multiple levels from 32 to 2048, and the output length can be divided into multiple levels from 32 to 512. Each parameter level can be flexibly adjusted in combination with model type, business request distribution, service quality constraints, and hardware resources.
[0053] Therefore, a complete configuration space covering multiple loads and constraints is built, fully covering all operating parameter ranges of online inference, providing a complete configuration sample basis for subsequent selection of representative test points and small sample modeling.
[0054] This embodiment uses the above-described scheme to standardize the division of the heterogeneous inference environment in the cloud and build a full-dimensional configuration space. It eliminates the need to exhaustively test all candidate configurations and conducts subsequent energy consumption and latency calculation modeling for each deployment stack, thereby reducing the overall measurement workload.
[0055] In one feasible implementation, step S20, based on the multidimensional inference configuration space, filters representative measurement points covering the entire load range, including step S201: Step S201: Select representative configurations covering low, medium and high load ranges from the multidimensional inference configuration space as representative measurement points by using stratified sampling, boundary point sampling, active learning, uncertainty sampling or historical experience selection.
[0056] Specifically, the system does not exhaustively measure all configurations, but instead selects a small number of representative configurations from the configuration space as measurement points. Preferably, the measurement points cover low-load, medium-load, and high-load regions. For example, the load during the pre-filling phase can be used... Approximately, the load of the decoding stage is available The cumulative context scan volume or other load metrics are approximations. Measurement point selection can employ fixed rules, stratified sampling, boundary point sampling, active learning, uncertainty sampling, or historical experience.
[0057] Therefore, a small number of measurement points can be used to represent all load conditions, reducing the workload of actual measurement, while simultaneously collecting the multi-dimensional measured data required for modeling.
[0058] This embodiment avoids the high hardware and power costs associated with exhaustive full-configuration measurements by using the above-described scheme. It retains complete load characteristics by relying on a small number of measurement points that cover the entire load, providing sufficient and effective measured samples for subsequent high-frequency power modeling and fine-grained energy consumption attribution.
[0059] In one feasible implementation, step S30, which involves simultaneously acquiring high-frequency execution information and low-frequency power sampling data based on the representative measurement points to construct an initial high-frequency power trajectory, includes steps S301-S304: Step S301: Collect fine-grained high-frequency inference execution information and low-frequency power sampling data from the representative measurement points; Specifically, during each measurement run, the system collects fine-grained high-frequency inference execution information, including compute kernel name, compute kernel start timestamp, end timestamp, inference phase boundary, request window boundary, streaming multiprocessor activity, tensor core activity, compute core activity, video memory read / write activity, L2 cache / DRAM access activity, cache hit rate, operating frequency, temperature, batch processing status, etc. The system also collects low-frequency power sampling data, including power sampling values provided by the graphics card management program interface, data center graphics card management tool interface, driver interface, or device management interface.
[0060] Therefore, by synchronously collecting multi-dimensional high-frequency hardware operation data at representative measurement points, the hardware operation details at each stage of inference can be fully restored; by synchronously collecting low-frequency real power sampling values, real benchmark data can be provided for subsequent power model calibration, thereby improving the fitting accuracy of high-frequency power trajectory.
[0061] Step S302: Align the hardware metrics in the fine-grained high-frequency inference execution information to the same time axis to obtain time-synchronized standardized hardware feature data. Specifically, based on representative measurement points, the inference task is run multiple times to synchronously collect various high-frequency hardware indicators such as kernel execution trajectory, hardware activity, memory access volume, and temperature, as well as low-frequency power sampling data. All high-frequency hardware indicators are then time-aligned using timestamps as a benchmark to obtain time-synchronized standardized hardware feature data.
[0062] This enables the unification of timing data from multi-source heterogeneous high-frequency hardware, outputting standardized hardware features with synchronized timing and unified format, eliminating timing offset issues between various hardware indicators, and providing regular input data for subsequent power model construction.
[0063] Step S303: Based on the standardized hardware feature data, a high-frequency power estimation model is constructed by combining the device operating frequency, hardware temperature, and inference execution status. Specifically, the system aligns high-frequency hardware metrics to a unified time axis and establishes a high-frequency power estimation model. This model can be expressed as follows: ,in: Indicates high-frequency hardware specifications. Indicates the operating frequency. Indicates temperature. Indicates the execution state, inference phase, or kernel family state. This represents the power consumption estimation function to be learned or calibrated. Linear models, multinomial models, piecewise models, tree models, neural network models, or physics-inspired models can be used.
[0064] Therefore, by integrating multi-dimensional hardware operating variables to establish a power consumption mapping relationship, the key hardware operating elements that affect instantaneous power consumption are fully covered, ensuring that the power model fits the actual operating mechanism of the hardware.
[0065] Step S304: Output the initial high-frequency power trajectory according to the high-frequency power estimation model.
[0066] Specifically, the high-frequency power estimation model is run and the instantaneous power prediction value is output time-by-time, and the values are continuously spliced to form a complete initial high-frequency power trajectory.
[0067] This generates a continuous power consumption change curve with high time resolution, which makes up for the defect of coarse timing granularity in low-frequency power sampling and provides complete timing power consumption data for subsequent sampling path aggregation and model calibration.
[0068] This embodiment uses the above-described scheme to synchronously collect raw data from all dimensions of inference operation and perform time-series standardization processing. It integrates multiple types of hardware operating parameters to construct a power consumption fitting model, outputs a fine-grained continuous initial high-frequency power trajectory, and fully restores the instantaneous power consumption dynamic changes during inference. This provides reliable basic time-series data support for subsequent low-frequency sampling calibration and fine-grained energy consumption attribution.
[0069] In one feasible implementation, step S40, which involves hardware sampling path aggregation of the initial high-frequency power trajectory and trajectory calibration combined with the low-frequency power sampling data to obtain a high-precision high-frequency power trajectory, includes steps S401-S403: Step S401: Perform an equivalent transformation on the initial high-frequency power trajectory according to the native sampling processing mechanism of the low-frequency power interface to generate a comparable sequence that matches and compares with the low-frequency power sampling data; Specifically, the initial high-frequency power trajectory is equivalently transformed according to the native sampling processing mechanism of the low-frequency power interface to generate a comparable sequence that matches and compares with the low-frequency power sampling data, including... Step S4011: Replicate the hardware mechanisms of window averaging, filtering, time delay, and smoothing of the native sampling of the low-frequency power interface to obtain the replicated native sampling processing mechanism of the low-frequency power interface. Specifically, the low-frequency power sample value is usually not the instantaneous real power consumption, but is obtained by a smoothing hardware mechanism after internal sampling, averaging, filtering, time delay and window aggregation. This step replicates the native processing logic of this type of hardware, restores the actual sampling operation rules of the device's underlying layer, and accurately reproduces the native sampling processing mechanism of the low-frequency power interface.
[0070] This allows for precise matching of the inherent sampling characteristics of the hardware, eliminating comparison biases caused by different processing mechanisms for high and low frequency data, and providing a unified computational standard for subsequent sequence alignment.
[0071] Step S4012: The original sampling processing mechanism of the replicated low-frequency power interface is used as input, and the initial high-frequency power trajectory is equivalently converted into a comparable sequence that matches the low-frequency power sampling data through the sampling path aggregation operator.
[0072] Specifically, the inference process is divided into equally subdivided time slices, and complete hardware operation metrics are collected within each time slice. Based on the replicated hardware sampling mechanism, the sampling path aggregation operator is then used. The high-frequency power trajectory is converted into a comparable low-frequency sequence with completely consistent sampling timing and statistical rules. The sampling path aggregation operator evenly divides the inference time to collect high-frequency hardware indicators. After smoothing, delay compensation, and window alignment correction, it is integrated and aggregated hierarchically according to request, stage, and operator, transforming the high-frequency fluctuating power consumption trajectory into a uniform and comparable low-frequency energy consumption sequence. The single-point sampling value of this comparable low-frequency sequence is... All single-point sampled values are combined to form a complete low-frequency comparable sequence, where: For low-frequency sampling time points, It can include moving average, exponential smoothing, time delay compensation, integral averaging, and sampling window alignment.
[0073] This achieves dimensional alignment and rule unification between high-frequency prediction data and low-frequency measured data, reducing the matching distortion problem caused by direct instantaneous comparison.
[0074] Step S402: Calculate the error between the comparable sequence and the real low-frequency power sampling data, introduce a regularization constraint term to construct a loss function, calibrate the parameters of the high-frequency power estimation model, and obtain the calibrated high-frequency power estimation model. Specifically, the error between the comparable sequence and the actual low-frequency power sampling data is calculated, and a loss function is constructed by introducing a regularization constraint term. The parameters of the high-frequency power estimation model are then calibrated to obtain the calibrated high-frequency power estimation model, including: Step S4021: Calculate the error between the comparable sequence and the low-frequency power sampling data; Specifically, based on the actual low-frequency power sampling values collected by the equipment, the high-frequency comparable sequences obtained by time-series comparison and aggregation are used to quantify the overall deviation between the two sets of data and accurately calculate the mean square error between the model prediction and the hardware measurement. ,in: , This represents the low-frequency power sample value, and the model parameters are calibrated based on the error difference.
[0075] Therefore, the fitting deviation of the high-frequency power model can be accurately located, providing a clear error basis for parameter iterative optimization.
[0076] Step S4022: Based on the calculated error, construct a loss function with regularization constraints using the low-frequency true power sampling value as a benchmark; Specifically, by combining the time-series error values, a regularization constraint term is introduced to construct the complete loss function S= ,in: This is a regularization term.
[0077] Therefore, the model's fitting accuracy and generalization ability are balanced, ensuring stable fitting of multiple power consumption components and controllable parameter iteration.
[0078] Step S4023: Iteratively calibrate the parameters of the high-frequency power estimation model by minimizing the loss function to obtain the calibrated high-frequency power estimation model.
[0079] Specifically, with minimizing the loss function as the optimization objective, the power estimation model parameters and sampling path parameters are calibrated. The parameter update range includes multi-dimensional power consumption fitting parameters, sampling window length, time delay compensation amount and smoothing coefficient, so that the aggregated high-frequency sequence infinitely approximates the real low-frequency power sampling data, and the calibrated high-frequency power estimation model is obtained.
[0080] This allows for precise calibration of the overall model parameters, ensuring that the high-frequency trajectory aggregation results closely match the actual power consumption observations of the device.
[0081] Step S403: Based on the calibrated high-frequency power estimation model, a high-precision high-frequency power trajectory is obtained.
[0082] Specifically, after calibration, the high-frequency power trajectory, after being aggregated through the same sampling path, should be consistent with or have a small error compared to the low-frequency power sampling value. Based on the calibrated high-frequency power estimation model, a continuous and complete high-precision high-frequency power trajectory is refitted and generated. This trajectory retains the millisecond-level fine-grained instantaneous power consumption change characteristics, accurately captures the dynamic power consumption changes during short-term inference stages such as pre-filling, and is consistent with the device-level low-frequency power consumption observation results. At the same time, time-slice energy consumption can be allocated for overlapping kernel operations and data transfer operations.
[0083] This ensures compatibility between high-frequency trajectories and equipment-level low-frequency power observations, resulting in high-precision high-frequency power trajectories with fine timing, accurate scale, and controllable errors, thus meeting the energy consumption measurement needs of short-term inference scenarios.
[0084] This embodiment replicates the hardware's native low-frequency sampling mechanism for high-frequency trajectory aggregation and alignment using the above-described scheme. It eliminates model system errors through iterative calibration with regularization constraints, effectively solving the problems of large low-frequency sampling periods and inaccurate energy consumption calculations during short-term inference stages. At the same time, it reduces the repeated attribution defects caused by multi-core overlapping operation through a time-slice energy consumption amortization mechanism, significantly improving the overall accuracy and stability of fine-grained power consumption reconstruction, and providing accurate and reliable high-frequency power consumption data for subsequent multi-level energy consumption attribution and modeling.
[0085] In one feasible implementation, step S50, which involves performing fine-grained energy consumption attribution and kernel family decomposition based on the high-precision high-frequency power trajectory to construct a kernel family-level load scaling law small-sample generalization model, includes steps S501-S503: Step S501: Align the high-precision high-frequency power trajectory with the calculation kernel timestamp, inference stage boundary, request window, batch processing window and kernel family label to obtain aligned power trajectory data carrying multi-dimensional inference information; Specifically, the reconstructed high-precision high-frequency power time-series trajectory is precisely matched and aligned with the start and end timestamps of the computation kernel, the boundaries of the pre-filled and decoding inference stages, the single request execution window, the batch processing execution window, and the preset kernel family classification labels to obtain aligned power trajectory data carrying multi-dimensional inference information.
[0086] Therefore, a one-to-one correspondence is established between power consumption time-series data and inference service units, providing a precise time-series matching basis for multi-level, fine-grained energy consumption integral attribution.
[0087] Step S502: Based on the aligned power trajectory data, perform power integral calculation on various time windows to obtain fine-grained energy consumption labels corresponding to single core, kernel family, inference stage, single request, and single batch. Specifically, based on the aligned power trajectory data, power integral calculations are performed on various time-series windows to obtain fine-grained energy consumption labels corresponding to single kernels, kernel families, inference stages, single requests, and single batches, including: Step S5021: Delineate time-series integration windows of different durations based on the aligned power trajectory data; Specifically, the system aligns the calibrated high-frequency power trajectory with the calculated kernel timestamp, stage boundary, request window, and kernel family label, and integrates the aligned power trajectory data within a specified window to obtain the time-series integration window. The total energy consumption of the time-series integration window W can be expressed as... ,in: W represents the total energy consumption corresponding to the window, and W represents the time-series integral window. , The duration of the high-frequency subdivision time slice. Window W can be a single computing kernel, several similar computing kernels, a pre-filling stage, a decoding stage, a request window, a batch processing window, or a kernel family window.
[0088] This allows for the delineation of multi-dimensional, hierarchical time-series integration windows, providing a standard calculation range for refined energy consumption calculations.
[0089] Step S5022: Perform power integration calculation on each timing integration window and output fine-grained energy consumption labels for single core, kernel family, inference stage, single request, and single batch.
[0090] Specifically, the system categorizes the numerous computational kernels used in the large language model inference process into several kernel families and obtains kernel family-level energy consumption labels based on the reconstructed high-frequency power trajectories. After obtaining the kernel family labels, the system aligns the time window of each computational kernel with the high-frequency power trajectory. For a single computational kernel, the system can integrate the power between its start and end times to obtain the computational kernel-level energy consumption, forming a structured, fine-grained energy consumption label.
[0091] This breaks through the limitation of traditional device-level overall power consumption not being able to be broken down, and enables precise traceability of energy consumption at all levels, from total hardware power consumption to kernel, stage, request, and batch, providing high-precision fine-grained label support for subsequent kernel family modeling.
[0092] Step S503: Based on the fine-grained energy consumption label, perform kernel family decomposition and scaling law modeling to construct a kernel family-level load scaling law small-sample generalization model.
[0093] Specifically, based on the fine-grained energy consumption labels, kernel family decomposition and scaling law modeling are performed to construct a kernel family-level load scaling law small-sample generalization model, including: Step S5031: Based on the fine-grained energy consumption label, computation kernel name, operator semantics, execution features and trajectory metadata, the inference kernel is clustered and split to obtain multiple kernel original groups; Specifically, based on various fine-grained energy consumption labels, the kernel name, execution characteristics, operator semantics, execution trajectory metadata, or user-specified rules are calculated to classify the computing kernels into different kernel families, resulting in multiple original kernel groups.
[0094] Therefore, preliminary feature clustering of heterogeneous computing kernels is performed, which effectively sorts out the differences in the operational features of scattered kernels and provides a clear feature grouping basis for subsequent standardized kernel family regularization and classification.
[0095] Step S5032: The original kernel groups are classified and organized to identify typical kernel families, such as attention, matrix multiplication, key-value caching, normalization, activation function, and rotation position encoding. The load variation pattern of similar computations is unified to obtain standardized kernel family classification results. Specifically, the original kernel groups obtained from clustering are regularized, merged, and categorized. The preferred kernel families include attention, matrix multiplication, key-value caching, normalization, activation function, element-wise operation, and rotation position encoding. For communication-intensive inference or multi-GPU scenarios, kernel families such as communication, synchronization, memory management, and data transfer can also be added. According to the typical operation types of large model inference, they are uniformly divided into standard kernel families of attention, matrix multiplication, key-value caching, normalization, activation function, and rotation position encoding. The unified load change characteristics of each type of kernel are condensed to obtain standardized kernel family classification results.
[0096] Therefore, computational kernels with similar characteristics are uniformly classified, the random differences between individual kernels are weakened, and a stable, transferable, and reusable load scaling law is condensed, which greatly improves the stability and generalization of subsequent scaling law modeling.
[0097] Step S5033: Based on the standardized kernel family classification results, construct an energy consumption-latency scaling law model for the pre-filling, decoding inference stages and each kernel family. The energy consumption-latency scaling law model includes load-related shared slope coefficients and deployment-specific intercept parameters. Specifically, for each inference stage, each kernel family, and each prediction target q, based on the standardized kernel family classification results, energy consumption and latency scaling law models are built for the two major inference stages of pre-filling and decoding, as well as for each type of kernel family. This can be either time delay or energy consumption. A preferred form is... .in: Indicates kernel family, This represents the load-related slope coefficient shared across the same inference engine. This indicates the deployment of relevant intercept items. This indicates kernel family-related load characteristics. For batch size, For the input length, For the output length, This indicates that it has inherent fixed losses. This indicates offset compensation caused by the GPU or other AI accelerators. This represents the offset compensation brought about by the large language model. The model structure distinguishes between two types of core parameters: load-related shared slope coefficients and deployment environment-specific intercept parameters.
[0098] Furthermore, this form is only a preferred embodiment; in practice, non-logarithmic forms, piecewise scaling laws, generalized linear models, or other models with shareable load structures can also be used. This decouples general load patterns from deployment-specific parameters, separates shareable features from environment-specific features, and provides a reasonable model structure to support parameter reuse across deployment stacks and accurate calibration with small samples.
[0099] Step S5034: The shared slope coefficient, which is universal across hardware and models, is reused to infer the inherent law of the load. The deployment-specific intercept parameter is calibrated using a small amount of target deployment measurement point data. The two parameters are combined to form a scaling law small sample generalization model.
[0100] Specifically, multiple sets of measurement data are aggregated for various deployment stacks under the same inference engine. The shared slope coefficients related to the load for each kernel family are jointly calculated. Since the computation kernels, cache management, and batch processing logic of the same inference engine are similar, the load scaling variation patterns are universal across hardware and models. Only a small number of measurement points covering the low, medium, and high load ranges of the target deployment stack are selected. Based on the energy consumption and latency data of each kernel family stage output by the measurement points, deployment-specific intercept parameters are fitted to the current hardware computing power, memory bandwidth, model specifications, and operating temperature. Different kernel families can be paired with differentiated load characteristics to characterize load changes. The general slope coefficients and specific intercept parameters are integrated to complete model construction, resulting in a scaling-law small-sample generalization model.
[0101] Therefore, by completing the modeling of the new deployment stack with a small number of measurement points, without the need for full configuration and actual measurement, the measurement cost is significantly reduced while taking into account the model fitting accuracy and the generalization ability across hardware and models, thus achieving high-precision modeling effect with small samples.
[0102] This embodiment first achieves accurate alignment of multi-dimensional power trajectories and multi-level fine-grained energy consumption attribution through the above scheme. Then, it forms a standardized kernel family through kernel clustering and regularization. Based on the dual-parameter decoupled modeling method of load slope and deployment intercept, it realizes cross-deployment sharing of load patterns and accurate calibration of deployment differences with small samples. It only relies on a small number of measurement samples to perform high-precision kernel family scaling law modeling, reducing the experimental cost of configuration space modeling. At the same time, it ensures that the model has stable load fitting ability and excellent cross-deployment generalization performance, providing core model support for the accurate recovery of the complete energy consumption-latency map.
[0103] In one feasible implementation, step S60, which involves using the scaling law small sample generalization model to complete the configuration space data and generate a complete energy consumption-delay map, includes steps S601-S602: Step S601: Based on the scaling law small sample generalization model, predict the energy consumption and latency of each kernel family under unmeasured configuration, and accumulate the power consumption and latency data of the complete inference task layer by layer to complete the configuration space data. Specifically, a scale-law-based small-sample generalization model is trained, and the general load slope coefficient learned from the multi-source deployment stack is reused. This is combined with a dedicated intercept parameter calibrated using a small number of measurement points in the new deployment stack. For unmeasured configurations, the system first predicts the energy consumption and latency of each kernel family, then sums the kernel family components to obtain the energy consumption and latency for the inference phase or the complete request. The final output is an energy consumption-latency map, a mapping from candidate configurations to predicted energy consumption and latency. For batch prediction of all unmeasured inference configurations within the configuration space, the system first outputs the energy consumption and latency components for each kernel family at different batch sizes, input / output lengths, and hardware levels. Then, the prediction results for all kernel families are accumulated layer by layer to synthesize the overall power consumption and latency data for the complete inference phase and the single-request dimension, filling in the missing data in the multi-dimensional configuration space.
[0104] Therefore, the process of exhaustively testing with massive configurations is eliminated. Instead, data is supplemented across the entire space by modeling with only a small sample. This significantly reduces the computing power and time required for testing with new hardware, new models, and new deployment stacks. The general load variation pattern can stably maintain the calculation accuracy of all prediction data.
[0105] Step S602: Generate a complete energy consumption-latency map covering the configuration space based on the configuration space data.
[0106] Specifically, based on the complete full-coverage configuration space data, a precise mapping relationship is established between inference configuration parameters and energy consumption and latency indicators. A complete energy consumption-latency map is built to characterize the power consumption and performance characteristics obtained from inference of various loads and hardware configurations. Subsequently, local parameters can be updated by incremental measurement to correct the map and adapt it to scenarios such as hardware iteration, model updates, and changes in business load.
[0107] This generates a comprehensive, accurate, and long-term iterative energy consumption-latency panoramic map, providing robust and reliable underlying data support for online inference configuration filtering and energy consumption-aware scheduling.
[0108] Furthermore, this map can be used by schedulers or controllers, for example, given delay constraints. Choose the configuration with the lowest energy consumption. ,in: Indicates delay constraints, This represents the optimal inference configuration. This represents a candidate inference configuration. This indicates the energy consumption predicted by the corresponding model. This indicates the corresponding model prediction latency configuration. This indicates the constraints; alternatively, it can be combined with the real-time grid carbon intensity to select a configuration or operating time window with lower carbon emissions.
[0109] This embodiment utilizes the above-described scheme to complete small-sample, high-precision data using the kernel family's shared scaling law, overcoming the limitations of traditional schemes that require extensive field testing and modeling. New deployment environments can quickly build maps. The complete energy consumption-latency map can filter low-energy inference configurations that meet latency constraints, simultaneously optimizing inference service quality, computing power consumption, and carbon emission indicators. It adapts to business scenarios involving iterative updates and dynamic scheduling of heterogeneous cloud-based inference clusters, reducing the overall testing and maintenance costs of the platform.
[0110] The complete operational architecture of this embodiment can be found by referring to... Figure 2 As shown. Figure 2 The overall architecture flowchart of the energy consumption-delay map method for building a large model inference using small samples in this application is shown.
[0111] like Figure 2 As shown, the system input includes a high-frequency hardware counter, a calculated kernel timestamp, operating frequency, temperature, video memory access information, and low-frequency power sampling information. The data acquisition module collects raw data from multiple sources and performs time alignment. The high-frequency power reconstruction module generates an initial high-frequency power trajectory based on the high-frequency execution information. The low-frequency sampling path calibration module uses low-frequency power sampling constraints to correct the high-frequency power trajectory. The energy consumption attribution module outputs fine-grained energy consumption labels based on the calibration trajectory. The kernel family decomposition module extracts the kernel family feature structure, the scaling law modeling module builds a kernel family energy consumption and latency model, the small sample calibration module optimizes the model by deploying measurement points on a small number of targets, the energy consumption-latency map output module generates a complete energy consumption-latency map, and finally, the configuration selection interface module outputs the optimal inference configuration based on business constraints.
[0112] like Figure 3 The diagram shown is a flowchart of the high-frequency power trajectory reconstruction process described in this application. Figure 3 As shown, high-frequency hardware metrics such as core activity, memory read / write speed, operating temperature, and kernel timestamps are first collected, and a power estimation model is built based on these metrics to output initial high-frequency power prediction results. Subsequently, according to the sampling, filtering, delay, and window aggregation rules of the low-frequency power interface, the fine-grained high-frequency power estimates are aggregated to the low-frequency sampling time point. The aggregated sequence is aligned and calibrated with the low-frequency power sampling values output by the device management interface. The calibration error is used to iteratively correct the high-frequency power model and sampling path parameters to obtain a high-frequency power trajectory that matches the actual power consumption of the device. The reconstructed trajectory can be used for energy consumption integral calculation at the kernel, inference stage, and kernel family level.
[0113] like Figure 4 The figure shown is a timeline diagram of energy consumption attribution in this application. Figure 4As shown, the diagram presents the correspondence between request windows, inference stages, decoding rounds, computation kernel sequences, and reconstructed power trajectories using a unified timeline. The upper layer of the diagram divides various independent request windows, while the middle layer distinguishes between the pre-filling and decoding inference stages. The underlying computation kernel sets for the pre-filling and decoding stages are completely identical, containing seven types of operator kernels: attention, GEMM (Generalized Matrix Multiplication) matrix computation, KV (Key-Value) cache read / write, normalization, activation functions, element-wise operations, and RoPE (Rotation Position Encoding). Each complete model forward computation executes the entire set of kernels sequentially in a fixed order. The background curve represents the reconstructed high-frequency power trajectory. The system aligns the power trajectory with kernel timestamps, stage boundaries, request window boundaries, and kernel family labels. Integral operations are performed over any time interval to obtain energy consumption labels at the single kernel, inference stage, kernel family, and request window levels, enabling refined breakdown and tracing of the device's total power consumption down to the underlying computational units.
[0114] like Figure 5 The diagram shown is a schematic of the shared slope calibration in this application. Figure 5 The diagram illustrates the process of cross-deployment migration modeling between multiple source deployment stacks and a single target deployment stack. Multiple source deployment stacks are comprised of various acceleration hardware components and large language models associated with the same inference engine. The system learns a unified load change slope across kernel families within each source deployment stack. For a new target deployment stack, the system collects only a small number of representative measurement points and uses these points to calibrate the stack-specific intercept parameters or deployment-related scale parameters. Through this process, the target deployment stack can reuse the load growth trends learned from the source deployment stacks, while also absorbing the differences in absolute energy consumption or absolute latency caused by different GPUs, model sizes, memory bandwidth, operating environments, and deployment states. This modeling calibration logic can be used for both energy consumption and latency prediction targets.
[0115] like Figure 6 The diagram shown is a schematic representation of the energy consumption-delay map recovery for a small sample in this application. Figure 6 As shown, the system reconstructs the energy consumption and latency maps of the complete configuration space by measuring only a small number of configuration points. The multi-dimensional configuration space on the left retains only a small number of measured points covering low, medium, and high loads. The middle processing link sequentially performs kernel family decomposition, scaling law modeling, and small sample parameter calibration. The output on the right is a complete energy consumption and latency map covering all configurations. This solution does not require exhaustive testing of all candidate configurations. It relies on the kernel family load scaling law combined with a small number of measurement points to establish a complete mapping between configuration parameters and predicted energy consumption and latency. In addition to batch size, input length, and output length, the configuration dimensions can also be expanded to include parameters such as concurrent request count, hardware frequency, power limit, and quantization accuracy.
[0116] like Figure 7 The diagram shown is a closed-loop diagram of the configuration selection and incremental calibration in this application. Figure 7As shown in the figure, this illustrates how the energy consumption-latency map is used in runtime configuration selection and long-term maintenance. The system first reads or queries the energy consumption-latency map, and, combined with conditions such as latency constraints, carbon intensity, power limits, and dynamic frequency adjustment parameters, the scheduling unit selects low-energy-consumption inference configurations that meet service quality requirements for online operation. Online inference continuously collects measured data. If the error between the measured value and the map prediction exceeds a set threshold, the incremental calibration module updates the model parameters or local map data. Relying on this closed-loop mechanism, the map can be continuously iteratively corrected when hardware, models, and business loads change, maintaining stable prediction accuracy in the long term and preventing the map from becoming invalid after long-term use.
[0117] This embodiment addresses the problems of high measured costs, coarse granularity of power consumption measurement timing, poor generalization effect across hardware and models, and continuous decay of prediction accuracy over long-term operation in traditional energy consumption-latency map construction methods. It establishes an integrated architecture encompassing "full-link data acquisition, high-frequency trajectory reconstruction, hardware sampling path calibration, kernel family layered modeling, cross-deployment parameter reuse, and online incremental closed-loop updates." It replaces full-configuration traversal testing with a small number of representative measurement points, significantly reducing hardware computing power and power consumption. It calibrates power consumption curves to align with the underlying hardware sampling mechanism, achieving accurate multi-level energy consumption attribution. It adopts a modeling approach of globally shared load slope and individually calibrated deployment intercepts to improve adaptability to heterogeneous inference environments. A matching online self-updating closed loop adapts to long-term dynamic changes in business, providing stable and comprehensive data support for energy-saving scheduling of large-model inference platforms.
[0118] Compared with the prior art, the technical solution of this embodiment has the following advantages: The measured resource consumption is significantly reduced: a small number of configurations covering the entire load range are selected as measurement points, and exhaustive measurement of the entire space is abandoned; the newly added heterogeneous deployment stack only requires a small number of measurement points for modeling, reducing the computing power and power consumption caused by long-term hardware operation, and adapting to the batch deployment of large-scale inference clusters.
[0119] High precision in fine-grained power consumption measurement: Trajectory alignment calibration is performed by taking into account the smoothing, latency, and window aggregation characteristics of hardware low-frequency sampling to eliminate bias in high- and low-frequency data mechanisms; it supports multi-level energy consumption integration tracing at the single-core, inference stage, kernel family, and request window levels, making up for the shortcomings of the original low-frequency sampling timing being coarse and unable to finely decompose power consumption.
[0120] Excellent cross-deployment generalization capability: The kernel family load slope is uniformly shared under the same inference engine. Only a few test points are used to calibrate the exclusive intercept parameters corresponding to different hardware and models. There is no need to completely retest and build a new deployment stack map, which can adapt to diverse heterogeneous inference scenarios.
[0121] Long-term dynamic self-calibration with low maintenance costs: A closed-loop link of online inference, error detection, and incremental update is built. Map data is automatically corrected after hardware, model, and business load iterations, maintaining stable prediction accuracy over the long term and reducing manual iteration and maintenance workload.
[0122] Compatible with complex configuration parameters of multiple dimensions: The modeling system is compatible with various parameters that affect energy consumption and latency, such as batch size, text length, concurrency, hardware frequency, and quantization accuracy, and is suitable for diverse and highly dynamic online inference business conditions.
[0123] It can directly support the implementation of energy-saving scheduling on the platform: The complete energy consumption-latency map can automatically select the optimal low-energy-consumption inference configuration by combining latency constraints and carbon intensity indicators, taking into account both business service quality and computing power energy-saving needs, and has strong practicality for engineering implementation.
[0124] Furthermore, the present invention is not limited to the specific embodiments described above. Without departing from the overall technical concept of the present invention, there are various equivalent, alternative, and adaptable implementation methods to expand the scope of protection of the present invention. These include the following alternative solutions: (1) Hardware Ecosystem Alternatives: The hardware is not limited to NVIDIA GPUs (NVIDIA Graphics Processing Units). Although NVIDIA ecosystem tools such as NVML (NVIDIA Management Library), DCGM (Data Center GPU Manager), and Nsight Systems can be used in the preferred embodiments, this invention is also applicable to AMD GPUs (AMD Graphics Processing Units), TPUs (Tensor Processing Units), NPUs (Neural Processing Units), AI ASICs (AI Application Specific Integrated Circuits), or other AI accelerators. As long as the hardware or driver can provide low-frequency power sampling and a certain number of high-frequency execution metrics, similar high-frequency power reconstruction and small-sample map construction methods can be used.
[0125] (2) Alternative Sampling Interfaces: The low-frequency power sampling interface is not limited to NVML. Low-frequency power values can come from DCGM, ROCm SMI (ROC System Management Interface), onboard sensors, server BMC (Baseboard Management Controller), power management interfaces, rack-mount power interfaces, or other device management interfaces. External power meters can be used as a verification or enhancement method, but are not a necessary condition for this invention.
[0126] (3) Alternatives for Metric Collectors: High-frequency metrics are not limited to collection by a single tool. High-frequency execution information can come from Nsight Systems, CUPTI (CUDA Profiling Tools Interface), hardware performance counters, inference engine logs, driver execution traces, runtime API (Application Programming Interface) hooks, kernel startup interceptors, or operating system performance events. The names and granularity of metrics collected by different tools may differ, but they can all be mapped to information such as compute activity, memory access activity, frequency, temperature, and time boundaries.
[0127] (4) Replacement of kernel families: Kernel family categories are not limited to attention, general matrix multiplication, key-value caching, normalization, activation, element-wise operation, and rotation position encoding. For hybrid expert models, expert routing and expert multilayer perceptron kernel families can be added; for multi-GPU inference, communication, synchronization, and data transport kernel families can be added; when the inference engine adopts a new fusion computing kernel, a new kernel family can be created based on naming rules, execution features, or manual annotation.
[0128] (5) Alternatives to scaling-law models: Scaling-law models are not limited to log-linear models. As long as the model retains the ideas of kernel family decomposition, shared load-related structures across deployment stacks, and a small number of target deployment calibrations, nonlinear regression, piecewise functions, spline functions, tree models, neural networks, Bayesian models, or physically inspired models can be used.
[0129] (6) Alternatives to configuration space: Configuration space is not limited to batch size, input length, and output length. Actual systems can also include parameters such as frequency, voltage, power limit, concurrency, key-value cache allocation strategy, quantization precision, parallel strategy, inference engine version, CUDA version, driver version, temperature range, fan strategy, and scheduling strategy.
[0130] (7) Alternatives to map building methods: Map building methods are not limited to offline one-time building. The system can incrementally update the map online. When the model version, inference engine version, graphics card driver, CUDA version, ambient temperature or business load changes, the intercept can be recalibrated or the model can be locally updated based on a small number of new measurement points.
[0131] (8) Service Alternatives: This invention can be extended to deep learning inference tasks that are not large language models, such as diffusion models, recommendation models, transformer coding models, visual language models, and multimodal inference. For these tasks, the high-frequency power reconstruction and few-sample map building framework can be used simply by redefining the corresponding kernel families and load features.
[0132] (9) System form substitution: The present invention can be implemented as a method, system, device, server, cluster management component, inference service platform module, computer program product or computer-readable storage medium. Each functional module can be deployed on the same device or distributed on performance profiling nodes, monitoring nodes, scheduling nodes or control nodes.
[0133] Furthermore, the advancements achieved by this invention can be seen in ten dimensions: In power observation, it balances high-frequency timing accuracy without requiring additional power acquisition hardware; in fine-grained energy consumption attribution, it achieves precise energy consumption decomposition at multiple levels (kernel, kernel family, stage, request), solving the problem of difficulty in subdividing single-kernel energy consumption; in small-sample map construction, it relies on shared load slope requiring only a small number of measurement points, significantly reducing GPU measurement costs; in cross-deployment migration, it reuses general load patterns to adapt to heterogeneous hardware and multiple models; online operation has low computational overhead and can support real-time energy consumption scheduling; multi-layer calibration structure improves prediction robustness; measured data shows that only 3 measurement points are needed to complete full-configuration modeling, with a prediction error of approximately 9.6% and a 15-fold reduction in measurement energy consumption; it also reduces online and offline computing power and manpower maintenance costs; it is compatible with existing monitoring cluster infrastructure, facilitating engineering implementation; and it can also support low-carbon green inference scheduling in data centers, reducing overall carbon emissions.
[0134] Furthermore, assuming that the aforementioned basic technical process can fully realize small-sample map construction, the following preferred implementation methods can be used to further improve the power reconstruction accuracy, fine-grained attribution accuracy, model generalization ability, and long-term system stability. The specific preferred settings are as follows: (1) Preferably, the equivalent sampling path of the low-frequency power interface employs configurable or learnable parameters. Different GPUs, driver versions, device management interfaces, and sampling periods may correspond to different internal averaging, delay, smoothing, and time alignment behaviors. The system can estimate the average window length, delay, smoothing coefficient, or time offset in the sampling path using calibration data, so that the high-frequency power trajectory, after being aggregated through the sampling path, is consistent with the low-frequency power interface sampling values. This preferred scheme can improve the consistency between the high-frequency power reconstruction results and the low-frequency power observations.
[0135] (2) Preferably, a phase-aware energy consumption attribution method is adopted for short time windows. The pre-filling phase is usually short, and the low-frequency power sampling samples may be insufficient. Directly integrating the low-frequency power values can easily produce large errors. The system can combine adjacent idle segments, decoding segments, batch processing boundaries, computation kernel timestamps, and request phase boundaries to constrain and calibrate the power of short windows. For example, the baseline power can be estimated using adjacent idle windows, the energy consumption scale of the same deployment stack can be calibrated using longer decoding windows, and the energy consumption attribution range of the pre-filling phase can be limited using computation kernel timestamps, thereby reducing the energy consumption attribution error of short windows.
[0136] (3) Preferably, the kernel family composition, load characteristics, energy consumption label, and latency label within each stage are constructed according to the boundary of the inference stage. The pre-filling stage usually performs a forward computation on the complete input sequence, and the computational load is closely related to the input length and batch size; the decoding stage mainly uses word-by-word loops, and each word generation requires access to the existing key-value cache. Since the execution mode, computational density, memory access mode, and duration of the two stages are different, it is preferable to establish separate models for the latency of the pre-filling stage, the energy consumption of the pre-filling stage, the latency of the decoding stage, and the energy consumption of the decoding stage, thereby avoiding mixing the load scaling rules of different stages in the same model.
[0137] (4) Preferably, modeling is performed at the kernel family level, rather than just establishing a global stage model. The global stage model is easily affected by changes in the composition ratio of the kernel family; the kernel family level model can capture the load scaling rules of different execution modes such as attention, general matrix multiplication, key-value caching, normalization, activation functions, element-wise operations, and rotational position encoding. The system predicts the energy consumption and latency for each kernel family separately, and then recovers the stage-level, request-level, or configuration-level energy consumption and latency by summing. This preferred scheme can improve the model's generalization ability to different combinations of input length, output length, and batch size.
[0138] (5) Preferably, load characteristics matching the execution mode are set for different kernel families. For the attention kernel family in the decoding stage, the cumulative context access volume can be introduced as a load characteristic, which can be determined by the input length, output length, and the position of the generated tokens. For the batch size, the batch size itself, the logarithm of the batch size, the interaction between the batch size and the context length, the interaction between the batch size and the input length, or the interaction between the batch size and the output length can be introduced to characterize the impact of the batch size on the total computation, parallel efficiency, memory access, cache behavior, and fixed overhead amortization. This preferred scheme can more accurately describe the changes in energy consumption and latency under long context, long output, and different batch sizes.
[0139] (6) Preferably, each source deployment stack selects a small number of representative measurement points covering low-load, medium-load, and high-load areas. Low-load points are used to characterize fixed overhead and small-scale execution behavior, medium-load points are used to anchor common operating areas, and high-load points are used to characterize load growth trends under long context, large batch, or long output conditions. In a preferred embodiment, each deployment stack selects 3 representative measurement points; when the measurement budget is more sufficient, more measurement points can be selected to further reduce map recovery errors; in new deployment migration or extremely low measurement budget scenarios, one or a small number of target deployment measurement points can also be used for calibration.
[0140] (7) Preferably, for new GPUs, new models, or new deployment stacks, only the deployment-related intercept terms are recalibrated. When multiple source deployment stacks are available to learn a shared load slope, the new deployment stack can reuse the load-related slope coefficients in the existing kernel family scaling law model and calibrate the absolute power consumption and latency scales using only a small number of target measurement points. If the system detects significant changes in the kernel family implementation, inference engine version, compute kernel combination, or hardware execution characteristics of the new deployment stack, it can choose to locally update part of the slope, relearn a specific kernel family model, or prompt for additional measurement points.
[0141] (8) Preferably, the system monitors the reconstruction error and map prediction error. When the error of the high-frequency power trajectory aggregated through the low-frequency sampling path exceeds a preset threshold, the system can re-estimate the sampling path parameters or supplement calibration data. When the prediction error of the energy consumption-delay map at newly added measurement points exceeds a preset threshold, the system can determine whether the error source is an overall scale offset, a specific kernel family offset, or a specific load area offset, and trigger intercept updates, local slope updates, or the acquisition of new measurement points, respectively. This preferred scheme can improve the stability and freshness of the map during long-term operation.
[0142] Furthermore, this invention comprises seven interconnected and progressively layered core innovative technologies, all of which are key protective points that distinguish this invention from existing technologies and possess inventiveness, as detailed below: (1) High-Frequency Power Trajectory Reconstruction Method Based on Low-Frequency Power Interface Constraints. This proposal suggests a high-frequency power trajectory reconstruction method for the inference process of large language models. The system collects high-frequency execution information during the inference process, including the execution trajectory of the computation kernel, hardware performance counters, computation unit activity, memory access behavior, operating frequency, temperature information, and the running status of the inference engine, and estimates the high-frequency power trajectory of the device during the inference process based on the above information. At the same time, the system uses the low-frequency power sampling values provided by the device management interface to constrain and calibrate the estimated high-frequency power trajectory. Thus, without relying on an external high-frequency power meter, the system can recover the power change process with a time resolution higher than that of the device's low-frequency power interface, providing a basis for energy consumption analysis at the short time window, computation kernel, inference stage, and request granularity. Compared with schemes that only use low-frequency power sampling values for integration or regression, this scheme combines high-frequency execution information with low-frequency power observation for joint constraints, which can alleviate the problem that low-frequency sampling cannot reflect short-term power fluctuations.
[0143] (2) Equivalent Modeling and Joint Calibration Method for Low-Frequency Power Sampling Path. This scheme further performs equivalent modeling of the low-frequency power sampling process in the device management interface. Instead of directly comparing the high-frequency power estimate with the low-frequency power sampling points point by point, the system models the window averaging, filtering, smoothing, sampling delay, and time alignment offset that may exist within the low-frequency power interface as an equivalent sampling path. Specifically, the system first inputs the estimated high-frequency power trajectory into the equivalent sampling path to generate a low-frequency sequence that is comparable to the low-frequency power sampling value of the device management interface in terms of time scale and sampling semantics; then, based on the difference between the low-frequency sequence and the actual low-frequency power sampling value, the high-frequency power model parameters and the equivalent sampling path parameters are jointly calibrated. Through the above method, the system can reduce the systematic error caused by the filtering, delay, and time misalignment of the device's low-frequency power interface to the power trajectory reconstruction. This method differs from schemes that directly fit the low-frequency power sampling points, directly integrate the low-frequency power sampling value, or simply downsample the high-frequency estimation result.
[0144] (3) Fine-grained energy consumption attribution method based on reconstructed high-frequency power trajectories. This scheme performs fine-grained attribution of energy consumption during the inference process of a large language model based on the calibrated high-frequency power trajectory. The system maps the high-frequency power trajectory with the computation kernel timestamp, inference stage boundary, request window boundary, batch processing window boundary, and kernel family label onto a unified time axis, and performs power integration within the corresponding time interval to obtain energy consumption estimates at the computation kernel level, inference stage level, kernel family level, request window level, or batch processing window level. In some implementations, the inference stage boundary includes the pre-filling stage, decoding stage, or other execution stages defined by the inference engine. The system can perform power integration, kernel family statistics, and energy consumption attribution separately within different inference stages to avoid mixing execution modes, load scaling rules, and kernel composition ratios of different stages. In this way, the system can obtain fine-grained energy consumption labels that are difficult to provide directly by the low-frequency power interface in the absence of an external high-frequency power meter. The fine-grained energy consumption labels can be used for subsequent kernel family modeling, stage-level energy consumption analysis, request-level energy consumption analysis, and energy consumption-latency map construction.
[0145] (4) Energy-Latency Map Construction Method for Large Language Model Inference Kernel Families. After obtaining fine-grained energy consumption and latency labels, this scheme divides the computational kernels in the large language model inference process into several kernel families according to execution semantics, hardware behavior, memory access characteristics, computational characteristics, or performance counter characteristics, and establishes kernel family-level energy consumption and latency models respectively. The system predicts the energy consumption and latency components corresponding to each kernel family based on batch size, input length, output length, number of concurrent requests, inference stage type, and kernel family composition relationship under different candidate configurations; then, the components of each kernel family are synthesized to obtain stage-level, request-level, batch processing window-level, or configuration-level energy consumption and latency estimation results, thereby forming a mapping relationship from candidate configuration to energy consumption and latency. Through kernel family-level modeling, this scheme can separate the load scaling rules within the kernel family from the changes in the kernel family composition ratio under different request lengths, batch sizes, stage types, and deployment conditions. Compared with the approach of directly performing black-box modeling of overall configuration-level energy consumption or latency, this method can improve the interpretability, transferability and small sample recovery capability of the energy consumption-latency map.
[0146] (5) A method for recovering a small-sample energy consumption-latency map based on kernel family scaling laws. This solution addresses the high cost of measuring the entire configuration space by proposing a method for recovering a small-sample energy consumption-latency map based on kernel family scaling laws. The system first learns the energy consumption and latency scaling laws of different kernel families as load varies on the source deployment stack. Load factors include batch size, input length, output length, request concurrency, stage type, or other inference configuration parameters. On the target deployment stack, the system selects a small number of representative configurations for actual measurement and uses these measurement results to calibrate the model parameters in the target deployment stack. The calibration can be applied to energy consumption scale parameters, latency scale parameters, bias parameters, local correction parameters, or other deployment-related parameters. After calibration, the system predicts the energy consumption and latency of each kernel family in the unmeasured configurations of the target deployment stack and further synthesizes an energy consumption-latency map in the complete configuration space. This method does not rely solely on fitting a small number of samples to the overall configuration-level model. Instead, it combines fine-grained energy consumption attribution, kernel family-level scaling laws, cross-deployment migration, and calibration with a small number of measurements on the target deployment to achieve energy consumption and latency prediction for unmeasured configurations. This significantly reduces the cost of constructing energy consumption-latency maps for large language model inference.
[0147] (6) A migration modeling mechanism for sharing load-related structures and calibrating deployment-related scales across deployment stacks. This scheme further proposes a migration modeling mechanism across deployment stacks. For different GPU-model deployment stacks under the same inference engine, the system shares load-related structures in the kernel family model and calibrates deployment-related scales through a small number of measurement points on the target deployment stack. The deployment stack can be determined by one or more factors among inference engine, model type, model size, GPU type, memory bandwidth, operating frequency, temperature status, software environment, and deployment status. The load-related structure is used to describe the trend of kernel family power consumption or latency changes with load factors such as batch size, input length, output length, and concurrency; the deployment-related scale is used to absorb the absolute power consumption and absolute latency differences caused by different GPUs, model sizes, memory bandwidth, operating environments, and deployment statuses. In this way, the system can utilize the kernel family load scaling rules already learned in the source deployment stack to reduce the number of measurement points required for the target deployment stack. Compared to the approach of directly remeasuring the entire configuration space on the target deployment stack, this method can reduce measurement costs; compared to the approach of directly migrating the overall configuration-level model, this method can preserve the load structure at the kernel family granularity and calibrate the absolute scale at the deployment granularity, thereby improving the stability of cross-deployment prediction.
[0148] (7) Incremental update and error-triggered calibration mechanism for energy consumption-delay map. This solution also supports incremental maintenance of the energy consumption-delay map during long-term operation. The system continuously compares the errors between the reconstructed power trajectory, low-frequency power sampling values, fine-grained energy consumption attribution results, and map prediction results. When the low-frequency aggregation error, fine-grained attribution error, or map prediction error exceeds a preset threshold, the system determines the source of the error based on the location, duration, inference stage, kernel family type, and load area, and triggers the corresponding update operation. In some implementations, when the error is mainly manifested as a systematic offset between the low-frequency sampling sequence and the equivalent sampling path output, the system updates the sampling path parameters; when the error is mainly manifested as an overall energy consumption or delay scale offset of the target deployment stack, the system updates the deployment-related scale parameters; when the error is concentrated in a specific kernel family, the system updates the energy consumption model or delay model of that kernel family; when the error is concentrated in a specific batch size, input length, output length, temperature range, or frequency range, the system supplements the collection of measurement points in the corresponding load area and performs incremental calibration on the local energy consumption-delay map. Through the above mechanism, the system can continuously correct the energy consumption-latency map when the model version, inference engine version, hardware status, operating environment, or load distribution changes, avoiding long-term map inaccuracies and reducing the overhead required to re-measure the entire configuration space. In some implementations, the energy consumption-latency map generated by the system can be provided to modules for inference configuration selection, batch processing control, energy-aware scheduling, carbon-aware scheduling, dynamic voltage and frequency adjustment, or power limit control. These modules can select the inference configuration based on the predicted energy consumption and predicted latency corresponding to the candidate configurations, under the condition of satisfying service quality constraints, latency constraints, throughput constraints, or energy consumption constraints.
[0149] Furthermore, to achieve the above objectives, this application also proposes a device for constructing a large-scale model inference energy consumption and time delay map using a small sample size, such as... Figure 8 As shown, the device includes: Deployment stack configuration definition module 10 is used to define the large model inference deployment stack and multi-dimensional inference configuration space; The measurement point filtering module 20 is used to filter representative measurement points covering the entire load range based on the multidimensional inference configuration space. The initial power trajectory construction module 30 is used to construct an initial high-frequency power trajectory by synchronously collecting inference high-frequency execution information and low-frequency power sampling data based on the representative measurement points. The power trajectory calibration module 40 is used to perform hardware sampling path aggregation on the initial high-frequency power trajectory and perform trajectory calibration in combination with the low-frequency power sampling data to obtain a high-precision high-frequency power trajectory. The scaling law modeling module 50 is used to perform fine-grained energy consumption attribution and kernel family decomposition based on the high-precision high-frequency power trajectory, and to construct a kernel family-level load scaling law small-sample generalization model. The energy consumption and time delay map generation module 60 is used to complete the configuration space data using the scaling law small sample generalization model to generate a complete energy consumption and time delay map.
[0150] The small-sample energy consumption and latency map construction device for large-model inference provided in this application adopts the small-sample energy consumption and latency map construction method for large-model inference provided in the above embodiments, which can solve the problems of high measurement cost, difficulty in accurately subdividing the energy consumption of each computing core, and poor cross-device and cross-model adaptation. Compared with the prior art, the beneficial effects of the small-sample energy consumption and latency map construction device for large-model inference provided in this application are the same as the beneficial effects of the small-sample energy consumption and latency map construction method for large-model inference provided in the above embodiments, and other technical features in the small-sample energy consumption and latency map construction device for large-model inference are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.
[0151] Furthermore, to achieve the above objectives, this application also proposes a device for constructing a large model inference energy consumption and latency map using a small sample size. The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The computer program is configured to implement the steps of the method for constructing a large model inference energy consumption and latency map using a small sample size as described above.
[0152] The following is for reference. Figure 9 This document illustrates a structural diagram suitable for implementing a small-sample large-model inference energy consumption and latency map device according to embodiments of this application. The small-sample large-model inference energy consumption and latency map device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The small-sample device for building large-model inference energy consumption and latency maps shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0153] like Figure 9As shown, the device for building a large-scale model inference energy consumption and latency map from a small sample size may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the small-sample large-model inference energy consumption and latency mapping device to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows small-sample large-model inference energy consumption and latency mapping devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.
[0154] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0155] The small-sample energy consumption and latency map construction device for large-model inference provided in this application adopts the small-sample energy consumption and latency map construction method for large-model inference provided in the above embodiments, which can solve the problems of high measurement cost, difficulty in accurately subdividing the energy consumption of each computing core, and poor cross-device and cross-model adaptation. Compared with the prior art, the beneficial effects of the small-sample energy consumption and latency map construction device for large-model inference provided in this application are the same as the beneficial effects of the small-sample energy consumption and latency map construction method for large-model inference provided in the above embodiments, and other technical features in the small-sample energy consumption and latency map construction device for large-model inference are the same as the features disclosed in the previous embodiment method, and will not be repeated here.
[0156] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0157] In addition, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method for constructing a large model inference energy consumption and latency map using small samples as described above.
[0158] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0159] The aforementioned computer-readable storage medium may be included in the device for constructing a large model inference energy consumption and latency map from a small sample; or it may exist independently and not be assembled into the device for constructing a large model inference energy consumption and latency map from a small sample.
[0160] The aforementioned computer-readable storage medium carries one or more programs. When these programs are executed by the small-sample large-model inference energy consumption and latency map building device, the small-sample large-model inference energy consumption and latency map building device performs the following: defines a large-model inference deployment stack and a multi-dimensional inference configuration space; based on the multi-dimensional inference configuration space, it selects representative measurement points covering the entire load range; according to the representative measurement points, it synchronously collects high-frequency execution information and low-frequency power sampling data to construct an initial high-frequency power trajectory; it performs hardware sampling path aggregation on the initial high-frequency power trajectory and performs trajectory calibration in conjunction with the low-frequency power sampling data to obtain a high-precision high-frequency power trajectory; based on the high-precision high-frequency power trajectory, it performs fine-grained energy consumption attribution and kernel family decomposition to construct a kernel family-level load scaling law small-sample generalization model; and it uses the scaling law small-sample generalization model to complete the configuration space data and generate a complete energy consumption-latency map. Therefore, this application achieves equivalent power trajectory conversion by replicating the native hardware sampling mechanism and improves the accuracy of fine-grained power consumption measurement by combining loss function calibration, which can accurately decompose the energy consumption contribution of different kernel families; it performs small sample modeling based on a small number of representative measurement points, eliminating the need for exhaustive testing of all configurations, thus reducing the computing power and power consumption of the graphics processor test; by mining the load scaling pattern of the inference kernel, it improves the generalization prediction capability across hardware, models, and inference configurations, and can construct a configuration space energy consumption-latency map at low cost and high accuracy, providing effective support for energy-saving configuration selection and energy consumption-aware scheduling for large model inference.
[0161] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0162] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0163] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0164] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described method for constructing a large-model inference energy consumption and latency map using a small sample sample. This method can solve the problems of high measurement costs, difficulty in accurately subdividing the energy consumption of each computing core, and poor cross-device and cross-model adaptation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the method for constructing a large-model inference energy consumption and latency map using a small sample sample provided in the above embodiments, and will not be elaborated upon here.
[0165] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the method for constructing a large model inference energy consumption and latency map using small samples as described above.
[0166] The computer program product provided in this application can solve the problems of high measurement costs, difficulty in accurately subdividing the energy consumption of each computing core, and poor cross-device and cross-model adaptation. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the method for constructing large model inference energy consumption and latency maps using small samples provided in the above embodiments, and will not be repeated here.
[0167] One or more technical solutions proposed in this application have at least the following technical effects: defining a large model inference deployment stack and a multi-dimensional inference configuration space; based on the multi-dimensional inference configuration space, selecting representative measurement points covering the entire load range; according to the representative measurement points, synchronously collecting high-frequency execution information and low-frequency power sampling data to construct an initial high-frequency power trajectory; performing hardware sampling path aggregation on the initial high-frequency power trajectory, and combining it with the low-frequency power sampling data to perform trajectory calibration, obtaining a high-precision high-frequency power trajectory; based on the high-precision high-frequency power trajectory, performing fine-grained energy consumption attribution and kernel family decomposition, constructing a kernel family-level load scaling law small-sample generalization model; using the scaling law small-sample generalization model to complete the configuration space data, generating a complete energy consumption-latency map. Therefore, this application achieves equivalent power trajectory conversion by replicating the native hardware sampling mechanism and improves the accuracy of fine-grained power consumption measurement by combining loss function calibration, which can accurately decompose the energy consumption contribution of different kernel families; it performs small sample modeling based on a small number of representative measurement points, eliminating the need for exhaustive testing of all configurations, thus reducing the computing power and power consumption of the graphics processor test; at the same time, by mining the load scaling pattern of the inference kernel, it improves the generalization prediction capability across hardware, models, and inference configurations, and can construct a configuration space energy consumption-latency map at low cost and high accuracy, providing effective support for energy-saving configuration selection and energy consumption-aware scheduling for large model inference.
[0168] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for constructing a large model inference energy consumption and time delay map using small samples, characterized in that, The method includes: Define the large model inference deployment stack and the multidimensional inference configuration space; Based on the multidimensional inference configuration space, representative measurement points covering the entire load range are selected; Based on the representative measurement points, high-frequency execution information and low-frequency power sampling data are collected synchronously to construct an initial high-frequency power trajectory. Hardware sampling path aggregation is performed on the initial high-frequency power trajectory, and trajectory calibration is performed in combination with the low-frequency power sampling data to obtain a high-precision high-frequency power trajectory; Based on the high-precision high-frequency power trajectory, fine-grained energy consumption attribution and kernel family decomposition are performed to construct a kernel family-level load scaling law small-sample generalization model. The scaling law small sample generalization model is used to complete the configuration space data and generate a complete energy consumption-delay map.
2. The method as described in claim 1, characterized in that, The selection of a small number of representative measurement points covering the entire load range based on the multidimensional inference configuration space includes: Representative configurations covering low, medium, and high load ranges are selected from the multidimensional inference configuration space using stratified sampling, boundary point sampling, active learning, uncertainty sampling, or historical experience selection as representative measurement points.
3. The method as described in claim 1, characterized in that, The step of simultaneously collecting high-frequency execution information and low-frequency power sampling data based on the representative measurement points to construct an initial high-frequency power trajectory includes: Fine-grained high-frequency inference execution information and low-frequency power sampling data are collected from the representative measurement points. The hardware metrics in the fine-grained high-frequency inference execution information are aligned to the same time axis to obtain time-synchronized standardized hardware feature data. Based on the standardized hardware feature data, a high-frequency power estimation model is constructed by combining the device operating frequency, hardware temperature, and inference execution status. The initial high-frequency power trajectory is output based on the high-frequency power estimation model.
4. The method as described in claim 1, characterized in that, The step of performing hardware sampling path aggregation on the initial high-frequency power trajectory and combining it with the low-frequency power sampling data for trajectory calibration to obtain a high-precision high-frequency power trajectory includes: The initial high-frequency power trajectory is equivalently transformed according to the native sampling processing mechanism of the low-frequency power interface to generate a comparable sequence that matches and compares with the low-frequency power sampling data; The error between the comparable sequence and the real low-frequency power sampling data is calculated, and a loss function is constructed by introducing a regularization constraint term. The parameters of the high-frequency power estimation model are then calibrated to obtain the calibrated high-frequency power estimation model. Based on the calibrated high-frequency power estimation model, a high-precision high-frequency power trajectory is obtained.
5. The method as described in claim 4, characterized in that, The step of performing an equivalent transformation on the initial high-frequency power trajectory based on the native sampling processing mechanism of the low-frequency power interface to generate a comparable sequence that matches and compares with the low-frequency power sampling data includes: The hardware mechanisms of window averaging, filtering, time delay, and smoothing of the native sampling of the low-frequency power interface are replicated to obtain the replicated native sampling processing mechanism of the low-frequency power interface. Using the replicated low-frequency power interface native sampling processing mechanism as input, the initial high-frequency power trajectory is equivalently converted into a comparable sequence matching the low-frequency power sampling data through the sampling path aggregation operator.
6. The method as described in claim 4, characterized in that, The calculation of the error between the comparable sequence and the actual low-frequency power sampling data, and the introduction of regularization constraints to calibrate the high-frequency power estimation model parameters, to obtain the calibrated high-frequency power estimation model, includes: Calculate the error between the comparable sequence and the low-frequency power sampling data; Based on the calculated error, a loss function with regularization constraints is constructed using the low-frequency true power sampling value as a benchmark. The high-frequency power estimation model is obtained by iteratively calibrating the parameters of the high-frequency power estimation model by minimizing the loss function.
7. The method as described in claim 1, characterized in that, Based on the high-precision, high-frequency power trajectory, fine-grained energy consumption attribution and kernel family decomposition are performed to construct a kernel family-level load scaling law small-sample generalization model, including: The high-precision high-frequency power trajectory is aligned with the computation kernel timestamp, inference stage boundary, request window, batch processing window and kernel family label to obtain aligned power trajectory data carrying multi-dimensional inference information. Based on the aligned power trajectory data, power integral calculations are performed on various time-series windows to obtain fine-grained energy consumption labels corresponding to single kernel, kernel family, inference stage, single request, and single batch. Based on the fine-grained energy consumption labels, kernel family decomposition and scaling law modeling are performed to construct a kernel family-level load scaling law small-sample generalization model.
8. The method as described in claim 7, characterized in that, The step of calculating power integrals for various time-series windows based on the aligned power trajectory data to obtain fine-grained energy consumption labels corresponding to single kernels, kernel families, inference stages, single requests, and single batches includes: Based on the aligned power trajectory data, time-series integration windows of different durations are defined; Perform power integration calculations for each time-series integration window and output fine-grained energy consumption labels for single kernel, kernel family, inference stage, single request, and single batch.
9. The method as described in claim 7, characterized in that, The step of performing kernel family decomposition and scaling law modeling based on the fine-grained energy consumption labels to construct a kernel family-level load scaling law small-sample generalization model includes: Based on the fine-grained energy consumption labels, computation kernel names, operator semantics, execution features, and trajectory metadata, the inference kernels are clustered and split to obtain multiple original kernel groups; The original kernel groups were categorized and organized to identify typical kernel families, including attention, matrix multiplication, key-value caching, normalization, activation function, and rotation position encoding. The load variation patterns of similar computations were unified to obtain standardized kernel family classification results. Based on the standardized kernel family classification results, an energy consumption-latency scaling law model is constructed for the pre-filling, decoding inference stages and each kernel family. The energy consumption-latency scaling law model includes load-related shared slope coefficients and deployment-specific intercept parameters. The shared slope coefficients, which are universal across hardware and models, are reused to infer the inherent laws of the load. The deployment-specific intercept parameters are calibrated using a small amount of target deployment measurement point data. Combined with the two parameters, a scaling law small sample generalization model is obtained.
10. The method as described in claim 1, characterized in that, The step of using the scaling law small-sample generalization model to complete the configuration space data and generate a complete energy consumption-latency map includes: Based on the scaling law small sample generalization model, the power consumption and latency of each kernel family under unmeasured configuration are predicted, and the power consumption and latency data of the complete inference task are obtained by accumulating layer by layer, thus completing the configuration space data. Generate a complete energy consumption-latency map covering the configuration space based on the configuration space data.
11. A device for constructing a large model for inference using small samples, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the method for constructing a large model inference energy consumption delay map using small samples as described in any one of claims 1 to 10.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the method for constructing a large model inference energy consumption and latency map using small samples as described in any one of claims 1 to 10.