Resource metering method and device, electronic equipment and storage medium
By combining the sliding window aggregation algorithm with the pre-trained pricing model, the billing rules are dynamically adjusted to solve the problem of deviation between resource usage and billing amount, achieve accurate billing of heterogeneous resources, and improve the credibility and fairness of the billing system.
Patent Information
- Application Number
- CN202511246702.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-02
AI Technical Summary
There is a significant deviation between resource usage and billing amount in the existing billing system, which affects the credibility and fairness of the billing system.
A sliding window aggregation algorithm is used to dynamically merge the execution timestamps and register states of heterogeneous resources, and a pre-trained pricing model is combined to determine real-time billing rules. Dynamic price adjustment signals are generated based on real-time resource supply and demand and user characteristics to achieve accurate billing.
It improves the accuracy of resource usage measurement, enhances the credibility and fairness of the billing system, and ensures the authenticity and real-time nature of billing results.
Smart Images

Figure CN120750683A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a resource metering method, device, electronic device, and storage medium. Background Art
[0002] With the continuous development of cloud computing technology, billing systems, as the core support module for cloud service operations, are widely used in areas such as resource scheduling, cost control, and service pricing.
[0003] Related technologies utilize a collaborative monitoring system, rule engine, and pricing model to build a complete technical system from resource collection and feature extraction to billing generation. While this billing method can achieve a certain degree of resource usage accounting, its minute-by-minute sampling mechanism can lead to significant discrepancies between resource usage and billing amounts, impacting the credibility and fairness of the billing system. Summary of the Invention
[0004] The present application provides a resource metering method, device, electronic device and storage medium to at least solve the problem in related technologies of significant deviation between resource usage and billing amount, which affects the credibility and fairness of the billing system.
[0005] This application provides a resource measurement method, including: Obtain execution timestamps and register states of heterogeneous resources; Based on a sliding window aggregation algorithm, dynamically merge the resource usage features in adjacent windows of the execution timestamps and register states of the heterogeneous resources with a preset step size and a preset duration, and extract metering information of the burst load; Determining real-time billing rules based on a pre-trained pricing model; wherein the pre-trained pricing model is trained based on real-time resource supply and demand and user characteristics; The charging for the heterogeneous resources is determined based on the real-time charging rule.
[0006] Optionally, obtaining the execution timestamp and register status of the heterogeneous resources includes: Obtain the start and end timestamps of GPU kernel functions based on the performance analysis interface, and obtain register status to reflect the real-time usage of hardware resources; Based on the performance monitoring subsystem and preset probes, lightweight monitoring is performed on virtualized resources; wherein the virtualized resources include at least one of a central processing unit, memory, and network bandwidth.
[0007] Optionally, the method of dynamically merging the resource usage characteristics within adjacent windows of the execution timestamps and register states of the heterogeneous resources with a preset step size and a preset duration based on a sliding window aggregation algorithm and extracting metering information of the burst load further includes: Dynamically adjusting the preset step size based on real-time resource usage characteristics; In response to a sudden load fluctuation, the preset step size is reduced.
[0008] Optionally, obtaining the execution timestamp and register status of the heterogeneous resources further includes: Inserting monitoring instructions into the graphics processor through instruction instrumentation to capture resource usage events at the graphics processor kernel function level; The preset probe is used to collect the usage status of virtualized resources in real time in the operating system kernel, and perform timestamp alignment processing.
[0009] Optionally, determining the real-time billing rules based on the pre-trained pricing model further includes: Based on the pre-trained pricing model, the regional resource supply-demand ratio and user consumption behavior characteristics are calculated in real time to generate dynamic price adjustment signals; Distributed training based on the policy gradient algorithm is used to optimize the price adjustment strategy.
[0010] Optionally, the method further includes: Collecting execution timestamps and register states of the heterogeneous resources based on the hardware layer; The resource usage characteristics are obtained based on the software layer, and metering information of the burst load is extracted.
[0011] Optionally, after obtaining the resource usage characteristics based on the software layer and extracting the metering information of the burst load, the method further includes: The raw data collected by the hardware layer is integrated with the feature data extracted by the software layer to generate resource measurement results across time granularity.
[0012] Optionally, determining the charging for the heterogeneous resources based on the real-time charging rule further includes: The intermediate data generated during the billing process is encrypted to generate a billing voucher with a timestamp, which is used by the user end to verify the authenticity of the billing result.
[0013] Optionally, after reducing the preset step size in response to a sudden load fluctuation, the method further includes: In response to the burst load returning to a stable state, the preset step size is gradually increased to an initial value.
[0014] Optionally, the heterogeneous resources include edge computing resources and cloud computing resources, and the determining of real-time billing rules based on the pre-trained pricing model further includes: The billing coefficient is adjusted based on the communication loss characteristics of the edge computing resources and the cloud computing resources.
[0015] Optionally, the lightweight monitoring of virtualized resources based on the performance monitoring subsystem and the preset probes further includes: Dynamically adjust monitoring granularity based on user permission levels; open fine-grained resource usage details to high-privilege users, and display aggregated metering results to ordinary users.
[0016] Optionally, after generating the resource measurement results across time granularities, the method further includes: The measurement results are synchronized to the distributed ledger according to the preset period to form an unalterable billing record chain, supporting multi-node reconciliation and audit tracing.
[0017] The present application also provides a resource metering device, comprising: An acquisition unit, used to obtain the execution timestamp and register status of heterogeneous resources; an extraction unit, configured to dynamically merge resource usage features within adjacent windows of execution timestamps and register states of the heterogeneous resources with a preset step size and a preset duration based on a sliding window aggregation algorithm, and extract metering information of the burst load; A determination unit, configured to determine a real-time charging rule based on a pre-trained pricing model; wherein the pre-trained pricing model is trained based on real-time resource supply and demand and user characteristics; The determining unit is further configured to determine charging for the heterogeneous resources based on the real-time charging rule.
[0018] Optionally, the acquiring unit is further configured to: Obtain the start and end timestamps of GPU kernel functions based on the performance analysis interface, and obtain register status to reflect the real-time usage of hardware resources; Based on the performance monitoring subsystem and preset probes, lightweight monitoring is performed on virtualized resources; wherein the virtualized resources include at least one of a central processing unit, memory, and network bandwidth.
[0019] Optionally, the extraction unit is further used for: Dynamically adjusting the preset step size based on real-time resource usage characteristics; In response to a sudden load fluctuation, the preset step size is reduced.
[0020] Optionally, the acquiring unit is further configured to: Inserting monitoring instructions into the graphics processor through instruction instrumentation to capture resource usage events at the graphics processor kernel function level; The preset probe is used to collect the usage status of virtualized resources in real time in the operating system kernel, and perform timestamp alignment processing.
[0021] Optionally, the determining unit is further configured to: Based on the pre-trained pricing model, the regional resource supply-demand ratio and user consumption behavior characteristics are calculated in real time to generate dynamic price adjustment signals; Distributed training based on the policy gradient algorithm is used to optimize the price adjustment strategy.
[0022] Optionally, the device further includes: A collection unit, configured to collect execution timestamps and register states of the heterogeneous resources based on a hardware layer; The collection unit is further configured to obtain the resource usage characteristics based on the software layer and extract metering information of the burst load.
[0023] Optionally, the device further includes: The fusion unit is used to fuse the original data collected by the hardware layer with the feature data extracted by the software layer after the collection unit obtains the resource usage characteristics based on the software layer and extracts the metering information of the burst load, so as to generate resource metering results across time granularity.
[0024] Optionally, the determining unit is further configured to: The intermediate data generated during the billing process is encrypted to generate a billing voucher with a timestamp, which is used by the user end to verify the authenticity of the billing result.
[0025] Optionally, the device further includes: The extraction unit is further configured to gradually increase the preset step size to an initial value in response to the burst load recovering to a stable state after the extraction unit reduces the preset step size in response to the burst load fluctuation.
[0026] Optionally, the heterogeneous resources include edge computing resources and cloud computing resources, and the determining unit is further configured to: The billing coefficient is adjusted based on the communication loss characteristics of the edge computing resources and the cloud computing resources.
[0027] Optionally, the acquiring unit is further configured to: Dynamically adjust monitoring granularity based on user permission levels; open fine-grained resource usage details to high-privilege users, and display aggregated metering results to ordinary users.
[0028] Optionally, the device further includes: The synchronization unit is used to synchronize the resource measurement results across time granularity generated by the fusion unit to the distributed ledger according to a preset period, forming an unalterable billing record chain and supporting multi-node reconciliation and audit tracing.
[0029] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned resource metering methods when executing the computer program.
[0030] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned resource metering methods are implemented.
[0031] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned resource metering methods when executed by a processor.
[0032] Through this application, due to the use of a sliding window aggregation algorithm to merge the usage characteristics of heterogeneous resources with dynamic step size and duration, the metering information of burst loads can be accurately captured. At the same time, the pre-trained pricing model based on real-time resource supply and demand and user characteristics is combined to determine the billing rules, thereby achieving more refined resource usage metering and adaptive billing than minute-level sampling. Therefore, it can solve the technical problem that the existing billing method, which adopts a minute-level sampling mechanism, causes a significant deviation between resource usage and billing amount, affecting the credibility and fairness of the billing system, and achieves the technical effect of improving the accuracy of heterogeneous resource billing and enhancing the credibility and fairness of the billing system. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0034] Figure 1 A schematic diagram of a flow chart of a resource metering method provided in an embodiment of the present disclosure; Figure 2 A schematic diagram of the structure of a resource metering device provided in an embodiment of the present disclosure; Figure 3 A schematic diagram of the structure of another resource metering device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0035] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0036] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0037] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0038] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the resource metering method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0039] The embodiments of the present application provide a resource metering method, and the method is described in detail in conjunction with the execution process of the resource metering method. Figure 1 A flowchart of a resource metering method provided in an embodiment of the present disclosure.
[0040] like Figure 1 As shown, the method comprises the following steps: Step 101, obtaining the execution timestamp and register status of the heterogeneous resources; Heterogeneous resources refer to the various types of computing resources involved in scenarios such as cloud computing and artificial intelligence (AI) reasoning. These include but are not limited to hardware resources with different architectures or functions, such as graphics processing units (GPUs), central processing units (CPUs), and memory. These resources often exhibit burst load characteristics at the millisecond or even microsecond level in scenarios such as large-scale AI model reasoning and real-time stream processing. Traditional minute-level sampling makes it difficult to accurately capture their actual usage status.
[0041] To accurately capture the execution timestamps and register states of heterogeneous resources, a hardware-level monitoring interface is integrated with a system-level performance monitoring mechanism. Specifically, for heterogeneous hardware resources such as GPUs, the CUDAProfiling API (Compute Unified Device Architecture Profiling API) is used to directly monitor the hardware execution process. This API provides in-depth awareness of key nodes such as the startup, execution, and termination of GPU kernel functions. Simultaneously, the Linux performance monitoring subsystem provides system-level hardware event tracing capabilities to ensure comprehensive resource status monitoring. Through the collaborative design of these two, the system can utilize the lightweight event marking mechanism CUDAEvent to record precise timestamps and register states during the execution of heterogeneous resources. The timestamp accuracy can reach 100μs, while the register state contains hardware status information about the resource during execution, such as register occupancy, instruction execution progress, and other key data.
[0042] In actual execution, when heterogeneous resources (such as GPUs) start kernel functions for data processing, the system registers CUDA runtime callback functions to capture the kernel function's start time (startNs) and end time (endNs) in real time, and simultaneously records the register state parameters during this period. These timestamp information is accurate to the nanosecond level and can accurately reflect the actual execution time of the kernel function, while the register state provides hardware-level raw data for subsequent analysis of resource load intensity, computational density, and other characteristics. For example, in the real-time inference scenario of the DeepSeek large model, when a sudden load occurs during the model decoding phase where the GPU utilization soars from 15% to 89% within 50ms, the high-precision timestamps and register states obtained through this step can fully record the details of the resource fluctuations in this short period of time, avoiding the measurement deviation caused by insufficient time granularity in traditional sampling methods, providing real and accurate raw data support for the subsequent dynamic billing engine, and is the prerequisite for accurate billing in millisecond-level load scenarios.
[0043] Step 102: Based on a sliding window aggregation algorithm, dynamically merge the resource usage features in adjacent windows of the execution timestamps and register states of the heterogeneous resources with a preset step size and a preset duration, and extract burst load measurement information; The sliding window aggregation algorithm is a real-time processing method for time series data. By setting a fixed time window length (preset duration) and the step size of the window movement (preset step size), it performs segmented aggregation analysis on continuously collected resource usage data. In this solution, the selection of the preset duration and preset step size is based on the load characteristics of heterogeneous resources (especially GPUs in AI inference scenarios): considering that the burst load of GPUs in AI inference tasks usually lasts 10-200ms, in order to fully capture 1-4 complete fluctuation cycles while avoiding the increase in CPU overhead caused by excessively small windows due to excessive aggregation calculation frequency, or the reduction in measurement accuracy due to excessively large windows, the preset window length (preset duration) is set to 50ms and the preset step size is set to 10ms. This setting allows the window to slide at intervals of 10ms, continuously covering the time series data of resource usage, ensuring continuous tracking of load changes over a short period of time.
[0044] The process of dynamically merging resource usage features within adjacent windows is as follows: the system maintains resource usage data within the window through an incremental update mechanism. When new resource usage information (based on the execution timestamp and register state extracted in step 101, including features such as resource utilization and load intensity) is input, it first determines whether the current window is not full. If not, the valid data count is increased. The contribution of old data to the window sum, which is about to be overwritten by the new data, is then removed. The new data is stored in the window and the sum is updated. At the same time, the write pointer is moved through circular logic to achieve dynamic sliding of the window. In this way, resource usage features within adjacent windows are organically merged, preserving the load details of the short period of time while reducing the redundancy of the original data through aggregation, ensuring a balance between data processing efficiency and accuracy.
[0045] Extracting metering information for burst loads on this basis means identifying fluctuations beyond the normal load range by calculating the aggregation results of resource usage characteristics within the window (such as the average value and weighted sum of resource utilization within the window). For example, in a large-model real-time inference scenario, when GPU utilization soars from 15% to 89% within 50ms, through sliding aggregation with a 50ms window and a 10ms step size, the system can accurately capture key metering information such as the start time, peak intensity, and duration of this sudden fluctuation. Traditional minute-level sampling can only record average values (such as 30.06%) due to insufficient time granularity, resulting in metering deviations. Therefore, step 102 provides the subsequent dynamic billing engine with a metering basis that can truly reflect the burst load status of heterogeneous resources through dynamic processing of the sliding window aggregation algorithm, and is the core technical link for solving the problem of inaccurate metering of millisecond-level burst loads.
[0046] Step 103: determining a real-time charging rule based on a pre-trained pricing model; wherein the pre-trained pricing model is trained based on real-time resource supply and demand and user characteristics; The pre-trained pricing model is trained based on real-time resource supply and demand and user characteristics. It is the core link in implementing pricing strategies to respond to resource fluctuations and user needs in real time. Its function is to convert real-time resource status and user behavior characteristics into dynamically adjusted billing rules through the pre-trained model, solving the lag problem of traditional fixed rates that cannot reflect real-time supply and demand changes.
[0047] The pre-trained pricing model is an intelligent pricing framework that integrates real-time data and historical patterns. Its training process uses real-time resource supply and demand information and user characteristics as core inputs. Real-time resource supply and demand information primarily includes the current demand (D(t)) and supply (S(t)) of heterogeneous resources within a region (such as GPUs, CPUs, and memory). Specifically, the regional supply-demand ratio (D(t) / S(t)) is calculated using data collected by the multimodal perception layer, including resource utilization, remaining capacity, and peak load. This ratio directly reflects resource scarcity. When demand far exceeds supply, the supply-demand ratio increases, indicating a shortage of resources. Conversely, when the supply-demand ratio decreases, resources are in a state of abundance. User characteristics include key information such as cumulative consumption (Cu(t)), consumption frequency, historical billing records, and credit rating. Cumulative consumption reflects a user's dependence on resources and their spending power, and serves as a crucial basis for the model to balance price elasticity and user experience. During the pre-training phase, a large amount of historical data is used to learn the mapping relationship between these features and reasonable pricing, optimize the learnable parameters in the model, and enable the model to output reasonable price adjustment rules under different supply and demand scenarios and user characteristics.
[0048] The process of determining real-time billing rules based on the pre-trained pricing model is essentially the model's dynamic response and calculation to real-time input. When the system is running, the pre-trained pricing model receives real-time resource supply and demand data (such as remaining regional GPU capacity and current utilization) and user characteristics (such as a user's current cumulative spending) from the multimodal perception layer. It then dynamically calculates the price using a built-in price generation function. For example, the model first calculates a base price adjustment coefficient based on the real-time supply-demand ratio. When the supply-demand ratio increases (resource shortage), the coefficient increases to reflect a resource premium. Simultaneously, the model adjusts the price elasticity using an exponential function based on the user's cumulative spending, appropriately reducing the premium for high-spending users to balance resource allocation efficiency and user stickiness. For example, the price calculation process involves first calculating the supply-demand ratio using (ds) / s, and then generating a real-time price based on the ratio of the user's cumulative spending to their total spending.
[0049] To ensure the real-time and accuracy of billing rules, the pre-trained pricing model also supports a dynamic update mechanism, adjusting model parameters online every hour based on the latest resource supply and demand data and user behavior feedback (such as tuning and verifying parameters such as beta through the control variable method). This allows the generated billing rules to continuously adapt to fluctuations in the resource market. When it detects that the supply and demand ratio of GPU resources in a region exceeds a preset threshold due to sudden AI inference tasks, the model can trigger a gradient price adjustment rule (such as adjusting prices by 3%-6% every minute); and when resource supply and demand stabilize, the rules can automatically fall back to the baseline level. This pre-trained pricing model, trained based on real-time resource supply and demand and user characteristics, transforms complex market fluctuations into quantifiable and dynamically adjustable billing rules, realizing the transformation of billing strategies from "static presets" to "real-time responses," providing an accurate pricing basis for the dynamic billing engine.
[0050] Step 104: Determine charging for the heterogeneous resources based on the real-time charging rule.
[0051] Heterogeneous resources encompass a variety of hardware resources involved in cloud computing and AI inference scenarios, including but not limited to GPUs, CPUs, memory, and network bandwidth. These resources, due to differences in architecture and functionality, have varying usage characteristics and metering methods, necessitating unified billing mechanisms. Real-time billing rules are dynamic policies generated by the pre-trained pricing model in step 103, combining real-time resource supply and demand with user characteristics. These policies include key elements such as billing mode, unit price, and adjustment coefficients. Examples include operator-granular billing for bursty AI inference loads and price fluctuation rules based on regional supply-demand ratios.
[0052] The process of determining billing based on real-time billing rules is mainly implemented by the three-layer architecture of the dynamic billing engine. First, through resource normalization processing, the heterogeneous resource usage characteristics extracted in step 102 (such as GPU utilization, CPU occupancy, memory consumption, etc.) are converted into unified billing units. Specifically, the resource normalization model (RNM) uses a pre-trained classification algorithm to map multi-dimensional resource feature vectors (such as register status, execution time, etc.) to a unified measurement standard. For example, the GPU computing power consumption and memory storage occupancy are converted into "resource units" that can be directly accounted for, ensuring that the usage of different types of resources can be measured in the same dimension.
[0053] The billing rule executor dynamically loads and executes real-time billing rules based on the Drools rule engine. When the system detects a sudden load on heterogeneous resources (such as a short-term surge in GPU utilization in AI reasoning tasks) through the metering information of step 102, the rule engine automatically triggers the preset rules and switches the billing mode from the default time granularity to the operator granularity, that is, accurate billing is based on the number of executions and resource consumption of specific calculation operators; and when the load returns to normal, the rule engine can automatically switch back to the default mode to ensure that the billing mode matches the resource usage status in real time. For example, in the DeepSeek large model inference scenario, when millisecond-level GPU load fluctuations are detected in the decoding stage, the rule engine activates the "real-time text generation" billing strategy, calculating the cost based on the execution time and resource occupancy of each decoding operator, avoiding metering deviations caused by traditional time-granularity billing.
[0054] Real-time billing processing uses a distributed processing pipeline built with Apache Flink to achieve end-to-end low-latency computing. The system combines normalized resource usage with parameters such as the unit price and adjustment factor in the real-time billing rules to calculate resource costs within each time window in real time. Flink's state backend stores the user's cumulative usage and cost data. For example, based on the real-time price generated in step 103 (e.g., a 6% increase in the GPU unit price due to tight supply and demand), the GPU resource unit consumption within a 50ms window is multiplied by the unit price to determine the billing amount for that window, which is then added to the user's cumulative costs.
[0055] The billing process also incorporates the dynamic adjustment mechanism within real-time billing rules to adjust fees in real time. For example, when the regional resource supply-demand ratio exceeds a preset threshold, the fee for the current billing cycle is adjusted accordingly based on the gradient price adjustment coefficient specified in the rule (e.g., a 3%-6% increase per minute). Furthermore, for users with high cumulative consumption, the premium is appropriately reduced based on the elasticity coefficient specified in the rule, balancing resource allocation efficiency and user experience. Finally, the system generates a real-time bill for users by integrating billing results from each time window, enabling precise billing for heterogeneous resources based on real-time billing rules.
[0056] In some embodiments, obtaining the execution timestamp and register status of the heterogeneous resources includes: Obtain the start and end timestamps of GPU kernel functions based on the performance analysis interface, and obtain register status to reflect the real-time usage of hardware resources; Based on the performance monitoring subsystem and preset probes, lightweight monitoring is performed on virtualized resources; wherein the virtualized resources include at least one of a central processing unit, memory, and network bandwidth.
[0057] For core hardware resources like GPUs, the system implements in-depth monitoring through a performance analysis interface. This interface directly senses the lifecycle of GPU kernel functions. When the GPU starts a kernel function to perform data processing tasks, it registers a runtime callback function to capture the kernel function's start timestamp (startNs) and end timestamp (endNs) in real time. These two timestamps are accurate to nanoseconds and accurately reflect the kernel function's actual execution duration. Simultaneously, the interface synchronously obtains the GPU's register status, including hardware-level information such as register occupancy, instruction execution progress, and data cache status. This status data directly reflects the GPU's real-time resource usage intensity during execution, providing the basis for subsequent judgment of load type (e.g., compute-intensive or general-purpose).
[0058] For virtualized resources, including CPU, memory, and network bandwidth, the system efficiently monitors them through a performance monitoring subsystem (such as the Linux performance monitoring subsystem) combined with pre-defined lightweight probes (such as the eBPF probe (Extended Berkeley Packet Filter)). The Linux performance monitoring subsystem provides system-level hardware event and software behavior tracking capabilities, collecting essential data such as CPU utilization, process scheduling information, memory allocation and release frequencies, and network packet transmission rates in real time. Pre-defined eBPF probes, as dynamic tracing tools, enable fine-grained detection of key virtualized resource behaviors without disrupting normal system operation. These probes capture details such as CPU process context switches, memory page swapping in and out, network connection establishment, and data transfer volume. This lightweight monitoring approach ensures real-time data collection (millisecond-level response) while minimizing the performance overhead of the monitoring process on the virtualized resources themselves. This ensures that captured data such as CPU load fluctuations, memory usage peaks, and network bandwidth utilization truly reflect actual virtualized resource usage. Through targeted monitoring of hardware resources such as GPUs and virtualized resources such as CPUs and memory, the system comprehensively obtains key data such as the execution timestamps and register status of heterogeneous resources, laying a data foundation for subsequent accurate metering and billing.
[0059] In some embodiments, the method of dynamically merging the resource usage characteristics within adjacent windows of the execution timestamps and register states of the heterogeneous resources with a preset step size and a preset duration based on a sliding window aggregation algorithm and extracting the metering information of the burst load further includes: Dynamically adjusting the preset step size based on real-time resource usage characteristics; In response to a sudden load fluctuation, the preset step size is reduced.
[0060] The preset step size refers to the interval at which the sliding window moves on the time series data (such as the initial setting of 10ms). Its size directly affects the window's tracking sensitivity to changes in resource load: the smaller the step size, the more intensive the window sliding, and the more precise the capture of load fluctuations in a short period of time, but it increases data processing overhead; the larger the step size, the higher the processing efficiency, but may miss key load details.
[0061] Dynamically adjusting the preset step size based on real-time resource usage characteristics essentially involves adaptively optimizing the step size parameters by analyzing the fluctuation characteristics of resource load in real time (such as the rate of change, amplitude, and duration of resource utilization). Real-time resource usage characteristics can be obtained through heterogeneous resource execution timestamps, register states, and aggregated utilization data collected by the multimodal perception layer. For example, if resource utilization changes smoothly within a continuous window (such as GPU utilization remaining stable at 20%-25%), it indicates that the load is stable. In this case, maintaining or appropriately increasing the preset step size (such as maintaining it at 10ms or adjusting it to 15ms) can ensure measurement accuracy while reducing computational overhead. However, if resource utilization fluctuations increase (such as a surge from 15% to 89% in a short period of time), the real-time resource usage characteristics are determined to be unstable, triggering the dynamic step size adjustment mechanism.
[0062] Reducing the preset step size in response to sudden load fluctuations is a precise optimization strategy for highly dynamic load scenarios. Sudden load fluctuations are usually manifested as drastic changes in the utilization of heterogeneous resources (such as GPUs in AI inference scenarios) within milliseconds. Such fluctuations are short-lived but affect measurement accuracy. If a fixed preset step size (such as 10ms) is still used, the onset, peak, and fallback of the fluctuation may not be fully captured due to the large window sliding interval, causing the aggregation result to smooth out key fluctuation details. At this time, reducing the preset step size (such as adjusting it from 10ms to 5ms) allows the sliding window to cover the time series data at a denser frequency, aggregating resource usage characteristics at smaller intervals, thereby more accurately recording the fluctuation period, peak intensity, and duration of the sudden load. For example, during the decoding phase of DeepSeek's large-scale real-time inference, when GPU utilization soared from 15% to 89% within 50ms, reducing the preset step size from 10ms to 5ms allowed the sliding window to capture twice as many load change nodes within the same timeframe. The aggregated metering information more accurately reflects the resource consumption of this sudden burst, avoiding the loss of fluctuation details caused by a fixed step size and providing more accurate intermediate data support for subsequent billing calculations. This dynamic adjustment mechanism balances efficiency and accuracy during stable loads and prioritizes accuracy during bursts, enabling the sliding window aggregation algorithm to adapt to complex resource scenarios.
[0063] In some embodiments, obtaining the execution timestamp and register status of the heterogeneous resources further includes: Inserting monitoring instructions into the graphics processor through instruction instrumentation to capture resource usage events at the graphics processor kernel function level; The preset probe is used to collect the usage status of virtualized resources in real time in the operating system kernel, and perform timestamp alignment processing.
[0064] To capture resource usage events at the kernel level of the GPU, the system uses instruction instrumentation technology to precisely insert lightweight monitoring instructions into the GPU's instruction execution stream. These monitoring instructions are deeply bound to the lifecycle of the GPU kernel function, enabling real-time perception of key events such as the kernel function's startup, runtime phase switching, and termination. They also synchronously record the precise timestamp corresponding to each event (e.g., the nanosecond point at which the kernel function begins execution, the switching moment between key computational phases), and register state changes (e.g., hardware characteristics such as register occupancy and instruction execution efficiency). This instrumentation approach does not interfere with the kernel function's normal computational flow, but it can capture kernel function internal resource usage details that are difficult to capture with traditional interfaces. This ensures a complete record of fine-grained GPU resource consumption, providing underlying data support for the subsequent identification of compute-intensive tasks or bursty loads.
[0065] To collect the usage status of virtualized resources, the system implements real-time monitoring at the operating system kernel level through pre-set probes (such as eBPF probes, or extended Berkeley Packet Filters). As a dynamic tracing tool, eBPF probes can be loaded into kernel space to non-invasively detect key behaviors of virtualized resources. For example, they collect real-time status data such as CPU process scheduling information, memory page allocation and release frequencies, network packet transmission volume and latency, and synchronously record the kernel timestamps when these states occur. Because the timestamps of hardware resources such as GPUs and the kernel timestamps of virtualized resources may come from different clock sources, timestamp alignment is required to ensure time consistency in subsequent resource usage feature aggregation analysis. The system uses a unified clock synchronization mechanism (such as calibration based on a high-precision clock source) to map the timestamps of GPU kernel function events and the kernel timestamps of virtualized resource status data to the same time base, eliminating deviations between different hardware and software clock sources. For example, when there is a microsecond deviation between the timestamp of the GPU kernel function execution event and the timestamp of the CPU usage peak, the alignment processing can accurately associate the corresponding relationship between the two in the time dimension, providing the subsequent sliding window aggregation algorithm with raw data with a consistent time base, avoiding misjudgment of resource usage characteristics due to time dislocation, thereby ensuring the accuracy and reliability of heterogeneous resource metering information.
[0066] In some embodiments, determining the real-time charging rules based on the pre-trained pricing model further includes: Based on the pre-trained pricing model, the regional resource supply-demand ratio and user consumption behavior characteristics are calculated in real time to generate dynamic price adjustment signals; Distributed training based on the policy gradient algorithm is used to optimize the price adjustment strategy.
[0067] Based on a pre-trained pricing model, the regional resource supply-demand ratio and user consumption behavior characteristics are calculated in real time to generate dynamic price adjustment signals. This relies on real-time data continuously collected by the multimodal perception layer. The calculation of the regional resource supply-demand ratio focuses on the real-time demand (D(t)) and supply (S(t)) of heterogeneous resources within the region. Demand is calculated by monitoring the utilization of resources such as GPUs and CPUs by currently active tasks, while supply is determined by subtracting idle resources from the region's total resource capacity. The ratio of the two (D(t) / S(t)) directly reflects resource scarcity. When the ratio is greater than 1, resource demand exceeds supply, triggering a price premium adjustment. When the ratio is less than 1, resource availability is high, and price reductions can be appropriately implemented to incentivize usage. User consumption behavior is characterized by cumulative consumption (Cu(t)), consumption frequency, and historical price sensitivity. Cumulative consumption is aggregated in real time by the dynamic billing engine's status backend, reflecting users' long-term resource dependence and spending power. The pre-trained pricing model calculates these features through a built-in price generation function, converting the supply-demand ratio into a basic price adjustment coefficient. It also adjusts price elasticity through an exponential function based on the user's cumulative consumption (e.g., users with high cumulative consumption have lower premiums). Ultimately, it generates a dynamic price adjustment signal that includes the real-time unit price, adjustment range, and effective period, providing a specific price basis for billing rules.
[0068] Distributed training based on policy gradient algorithms (such as TD3, the double-delayed deep deterministic policy gradient algorithm) to optimize price adjustment policies is key to ensuring that pre-trained pricing models continuously adapt to resource market dynamics. This distributed training framework uses resource supply and demand data from multiple regions, user consumption behavior data, and historical pricing performance data as training samples. The state space is composed of regional resource supply-demand ratios, user consumption elasticity coefficients, and the current price adjustment range. The price adjustment actions for the next period (such as premium rates and discount coefficients) are used as continuous outputs. The policy gradient algorithm learns the mapping between state, action, and reward. During training, the model aims to maximize resource utilization and user satisfaction. By calculating the long-term benefits of different price adjustment strategies (e.g., increasing revenue with a moderate premium when the supply-demand ratio is high, while avoiding excessive premiums that lead to user churn), it continuously optimizes the model's learnable parameters. For example, the beta parameter (the weight of supply and demand influence) is tuned using a control variable approach. The price adjustment effects are verified in test scenarios across different regions. When a specific beta value is found to improve resource utilization and reduce user complaint rates, this parameter is updated in the model. The distributed training architecture ensures that the model can process heterogeneous data in multiple regions simultaneously, preventing data bias in a single region from affecting the generalization of the strategy. Ultimately, the optimized price adjustment strategy can not only quickly respond to real-time resource fluctuations, but also accurately match user consumption characteristics, achieving dynamic adaptation and continuous optimization of billing rules.
[0069] In some embodiments, the method further comprises: Collecting execution timestamps and register states of the heterogeneous resources based on the hardware layer; The resource usage characteristics are obtained based on the software layer, and metering information of the burst load is extracted.
[0070] The hardware layer collects execution timestamps and register states of heterogeneous resources, relying on specialized monitoring technologies for different resource types to achieve in-depth data capture. For core hardware resources such as GPUs, the hardware layer integrates the CUDA Profiling API (Unified Compute Device Architecture Performance Profiling Interface) with the Linux performance monitoring subsystem to build a high-precision monitoring system. The CUDA Profiling API directly connects to the GPU hardware interface, enabling real-time perception of key nodes such as kernel function startup, execution, and termination. Through the CUDA Event mechanism, it records start timestamps (startNs) and end timestamps (endNs) accurate to 100μs. It also synchronously collects register state information (such as hardware-level parameters such as register occupancy and instruction execution progress), fully reflecting the real-time resource consumption status of the GPU during kernel function execution. For virtualized resources such as CPU, memory, and network bandwidth, the hardware layer interacts with the operating system kernel through preset lightweight probes (such as eBPF probes, the extended Berkeley Packet Filter). Without interfering with the normal operation of resources, it captures key data such as process scheduling timestamps, memory allocation and release times, and network data packet transmission timing of virtualized resources in real time. It also embeds monitoring logic in the hardware execution flow through instruction instrumentation technology to ensure the collection of full-level timestamp and status information of heterogeneous resources from physical hardware to the virtualization layer, providing original data support for subsequent analysis.
[0071] The software layer acquires resource usage characteristics and extracts metering information for burst loads, performing refined processing and analysis based on the raw data from the hardware layer. The software layer first receives heterogeneous resource execution timestamps, register states, and virtualized resource usage data transmitted from the hardware layer. Using a streaming computing engine (such as Apache Flink), it builds a real-time processing pipeline to aggregate the discrete raw data into time series. Subsequently, a sliding window aggregation algorithm is used to process the aggregated data: the software layer presets the window duration (e.g., 50ms) and step size (e.g., 10ms), and dynamically maintains resource usage data within the window through an incremental update mechanism. When new data is input, old data within the window that is about to expire is removed, the new data is stored, and the total resource utilization within the window is updated in real time. Continuous sliding is achieved by moving the window pointer through circular logic. During this process, the software layer extracts resource usage characteristics from aggregated data, including the average resource utilization rate, fluctuation amplitude, and rate of change. When monitoring resource utilization and experiencing dramatic fluctuations within a short period of time (e.g., GPU utilization soaring from 15% to 89% within 50ms), the software layer analyzes the peak intensity, fluctuation period, and duration of the aggregated results within the window to accurately identify the burst load and extract its metering information (e.g., burst start time, peak utilization rate, and duration). This layered design, with the hardware layer collecting raw data and the software layer analyzing and extracting features, not only ensures data collection accuracy but also enables precise measurement of burst loads through intelligent algorithms, providing a reliable data basis for subsequent dynamic billing.
[0072] In some embodiments, after obtaining the resource usage characteristics based on the software layer and extracting the metering information of the burst load, the method further includes: The raw data collected by the hardware layer is integrated with the feature data extracted by the software layer to generate resource measurement results across time granularity.
[0073] For hardware resources like GPUs, raw data includes hardware-level parameters such as kernel function start / end nanosecond timestamps and register occupancy, captured via the CUDA Profiling API. This accurately reflects the timing details of individual kernel function execution. For virtualized resources like CPUs and memory, raw data includes underlying events like microsecond-level process scheduling times and memory page allocation times collected by eBPF probes at the operating system kernel level, fully documenting resource fluctuations at fine-grained time. While this raw data can capture microscopic load variations below the millisecond level, it is large and discrete, making it difficult to directly support metering requirements in macro-billing scenarios when used alone.
[0074] The feature data extracted by the software layer aggregates and refines the raw data, with clear temporal granularity. Using a sliding window aggregation algorithm, the software layer extracts features such as average resource utilization, fluctuations, and burst load peaks from the raw data. These feature data focus on millisecond-level load patterns and clearly reflect overall resource usage trends over short periods of time. However, the aggregation process may smooth out some microscopic details. Fusion processing achieves the complementary integration of microscopic details and macroscopic trends by establishing a correlation mapping between the raw data from the hardware layer and the feature data from the software layer.
[0075] Specifically, the fusion process first eliminates clock skew between the two types of data through a timestamp alignment mechanism. This calibrates the nanosecond timestamps of the hardware-layer raw data with the millisecond-level window time base of the software-layer feature data, ensuring accurate correspondence between raw events and feature indicators within the same time period. Subsequently, the system uses a data association algorithm to bind key raw events at the hardware layer (such as the sudden launch of GPU kernel functions and the intensive scheduling of CPU processes) with the feature data of the corresponding software-layer windows (such as the peak GPU utilization within a 50ms window and the fluctuation amplitude of the CPU load), making the micro-events the "data anchors" of the macro-features. After fusion processing, the kernel function's contribution to the window peak can be clearly identified through the raw data, while the overall load level within the window can be grasped through the feature data.
[0076] The resulting cross-time granularity resource metering results cover both micro- and macro-level metering information: at the micro-granularity, they retain the key event timing and state parameters in the original data at the hardware layer to ensure accurate recording of sub-millisecond burst loads (such as instantaneous computing power consumption at the kernel function level); at the macro-granularity, they integrate the characteristic data of the software layer to form load statistical indicators at different time scales, such as the window level and minute level, to meet the metering requirements for overall resource usage trends in billing scenarios. This fusion processing effectively addresses the limitations of single-granularity data—it avoids metering fragmentation caused by the discretization of original data and compensates for the loss of details caused by the aggregation of characteristic data, allowing the generated resource metering results to maintain accuracy and consistency at different time granularities, such as nanoseconds, microseconds, and milliseconds, providing a comprehensive and reliable metering basis for the subsequent execution of dynamic billing rules.
[0077] In some embodiments, determining the charging for the heterogeneous resources based on the real-time charging rule further includes: The intermediate data generated during the billing process is encrypted to generate a billing voucher with a timestamp, which is used by the user end to verify the authenticity of the billing result.
[0078] Intermediate data in the billing process encompasses critical information from resource metering to fee calculation, including but not limited to real-time usage of heterogeneous resources (such as GPU kernel execution duration, CPU utilization, and memory consumption), dynamic pricing parameters in real-time billing rules (such as the premium factor corresponding to the regional supply-demand ratio and the user consumption elasticity adjustment value), window-aggregated metering results (such as peak resource utilization within a 50ms window), and intermediate fee calculation results (such as the billing amount for each window and the accumulated fee). This data is directly related to the accuracy of billing results and requires encryption to prevent tampering or forgery.
[0079] Encryption is implemented using the Secure Hash Algorithm 3 (SHA-3), which generates a fixed-length hash value for input data of any length, offering strong collision resistance and high security. Specifically, the system consolidates each piece of intermediate data into a pre-defined format (e.g., resource type, timestamp, usage details, pricing parameters, etc.), calculates its hash value using the SHA-3-256 algorithm, and generates a unique digital fingerprint. For example, for billing intermediate data for a particular window, the system combines the "timestamp + GPU kernel function name + execution duration + real-time unit price + window billing amount" into a raw data string. This string is then hashed using the SHA-3-256 algorithm using the MessageDigest utility class, resulting in a fixed-length byte array that serves as the data's cryptographic identifier. This ensures that any data tampering during transmission or storage can be effectively identified.
[0080] Based on the generated cryptographic hash, the system simultaneously adds a precise timestamp to the billing voucher. This timestamp is consistent with the billing window time corresponding to the intermediate data (for example, accurate to the nanosecond level). This timestamp marks the timing of data generation and prevents verification failures caused by timing errors. In addition to the cryptographic hash value, the timestamped billing voucher also integrates metadata such as the resource type identifier, user ID, and billing rule version. It also appends a system-level digital signature (a signature generated by encrypting the voucher's core information using the system's private key) to form a complete, verifiable unit. For example, a voucher might include structured information such as "User ID: xxx | Timestamp: xxx | Resource Type: GPU | Cryptographic Hash: xxx | System Signature: xxx." The system signature verifies that the voucher was officially generated by the billing system, ensuring the authority of the source.
[0081] The generated billing vouchers will be combined with the Merkel-Patricia Tree (MPT) in chronological order to build a verification chain, and the hash values of adjacent vouchers will be aggregated layer by layer to form block-level batch vouchers, further improving the integrity and traceability of the data. The user end can independently verify the authenticity of the billing results through the verification API provided by the system: the user only needs to enter the time range, resource type or user ID to be verified, and the system will return the corresponding billing voucher. The user uses the local calculation tool to recalculate the SHA-3 hash value of the original data in the voucher (or the detailed data obtained from the API) and compare it with the encrypted hash value in the voucher, while verifying the rationality of the timestamp and the validity of the system signature. If the two are consistent and the signature verification is passed, it proves that the intermediate billing data has not been tampered with and the billing results are authentic and reliable, which effectively solves the problem of users relying on service providers to provide aggregated bills and lacking independent verification methods in the traditional billing model.
[0082] In some embodiments, after reducing the preset step size in response to a sudden load fluctuation, the method further includes: In response to the burst load returning to a stable state, the preset step size is gradually increased to an initial value.
[0083] The burst load stability state refers to the return of the usage characteristics of heterogeneous resources (such as GPUs in AI inference scenarios) from violent fluctuations to a smooth state. This can be determined through real-time monitoring of resource usage characteristics at the software layer: when the resource utilization fluctuation amplitude within multiple consecutive sliding windows (such as the difference between the maximum and minimum GPU utilization) drops below the preset threshold (such as the fluctuation amplitude is less than 10%), and the stable state persists for a certain period of time (such as more than 3 window cycles), the burst load is determined to have returned to a stable state.
[0084] In response to this stable state, the process of gradually increasing the preset step size to the initial value must follow a smooth transition principle to avoid sudden changes in the step size that could cause fluctuations in metering accuracy or a sudden increase in system overhead. The initial value is the preset step size (e.g., 10ms) before the sudden load fluctuation. During the gradual increase, the system adjusts the step size in stages based on the duration of the load stability. For example, if the step size decreases from 10ms to 5ms during a sudden burst, upon detecting a stable state, the system will first adjust the step size to 7ms. After running the window for one or two window cycles, it will monitor whether the resource usage characteristics remain stable. If so, it will be adjusted to 9ms, and after re-verifying the stable state, it will finally return to the initial value of 10ms. This phased adjustment ensures a smooth transition in the sliding window's tracking sensitivity to resource load, preventing possible secondary fluctuations from being missed due to a sudden increase in the step size. Furthermore, after the load stabilizes, increasing the step size reduces the computational overhead of data aggregation. Once the step size returns to the initial value, the window sliding frequency decreases, reducing unnecessary recalculation, thereby improving overall operational efficiency while maintaining metering accuracy.
[0085] The core of this mechanism is to dynamically adapt the step size parameters through real-time perception of resource load status, so that the sliding window aggregation algorithm can capture details with high sensitivity (small step size) during burst loads and operate with high efficiency (initial step size) in a stable state, achieving a dynamic balance between metering accuracy and system performance, and providing accurate and efficient intermediate data support for subsequent billing calculations.
[0086] In some embodiments, the heterogeneous resources include edge computing resources and cloud computing resources, and determining the real-time billing rules based on the pre-trained pricing model further includes: The billing coefficient is adjusted based on the communication loss characteristics of the edge computing resources and the cloud computing resources.
[0087] Edge computing resources are typically deployed close to end devices (such as 5G base stations and edge servers), offering low latency and highly localized processing. Cloud computing resources, on the other hand, are deployed in centralized cloud data centers, boasting massive computing power and storage capabilities. When these two resources work together, data transmission between edge and cloud computing resources (e.g., uploading partial computational results of edge tasks to the cloud, or sending model parameters from the cloud to the edge) can incur communication losses. These losses are primarily characterized by transmission latency (the time it takes for data to travel between two nodes), bandwidth usage (the network bandwidth consumed during transmission), packet loss rate (the percentage of data lost during transmission), and retransmission cost (the additional resource consumption caused by data retransmission due to packet loss). These communication losses directly impact the actual utilization efficiency of heterogeneous resources. For example, high communication latency can increase the time edge tasks wait for cloud responses, indirectly increasing the idle cost of edge resources. High bandwidth usage can also strain network resources, increasing overall resource consumption.
[0088] The pre-trained pricing model incorporates these communication loss characteristics as key inputs when determining real-time billing rules. The model uses lightweight probes in the multimodal perception layer (such as eBPF probes deployed at the network interface between edge and cloud computing resources) to collect real-time communication loss data, including average transmission delay, instantaneous bandwidth utilization peak, packet loss rate statistics, etc., and quantifies this data into communication loss coefficients (such as delay coefficient and bandwidth consumption coefficient). During the pre-training phase, the model uses historical data to learn the mapping relationship between communication loss characteristics and reasonable billing adjustments. For example, when the edge-to-cloud communication delay exceeds a preset threshold (such as 50ms), the corresponding edge computing resource billing coefficient needs to be appropriately adjusted upward to reflect the resource efficiency loss caused by the delay. When the bandwidth utilization rate exceeds a certain ratio (such as 80%), the cloud resource billing coefficient is adjusted to reflect the additional costs caused by network resource congestion.
[0089] The specific process for adjusting the billing coefficient based on communication loss characteristics involves the pre-trained pricing model generating basic billing rules (based on resource supply-demand ratios and user characteristics) and then integrating the communication loss coefficient with parameters such as unit price and premium margin in the basic billing rules through a built-in correction function. For example, in a scenario where edge computing resources upload inference data to the cloud, the model first calculates the basic billing unit price based on the real-time supply-demand ratio between the edge and cloud. The final billing rule is then generated using the formula "Adjusted unit price = basic unit price × (1 + communication loss coefficient × weight parameter)." The weight parameter, optimized by the pre-trained model, balances the impact of communication loss on different resource types (edge computing power, cloud storage, and network bandwidth). As communication loss decreases (e.g., transmission latency and bandwidth usage decrease), the billing coefficient decreases accordingly, ensuring that the billing rule dynamically adapts to changing communication conditions.
[0090] This billing coefficient adjustment mechanism, which combines communication loss characteristics, enables real-time billing rules to not only reflect the supply and demand relationship of the resources themselves, but also accurately quantify the hidden costs brought by cross-node communication, thereby achieving fair and reasonable pricing for edge computing resources and cloud computing resources, and providing support for efficient collaboration and cost optimization of heterogeneous resources.
[0091] In some embodiments, the lightweight monitoring of virtualized resources based on the performance monitoring subsystem and the preset probes further includes: Dynamically adjust monitoring granularity based on user permission levels; open fine-grained resource usage details to high-privilege users, and display aggregated metering results to ordinary users.
[0092] Monitoring granularity refers to the level of detail of virtualization resource monitoring data. Fine-grained monitoring covers real-time time series data of resource usage and consumption details at the specific process level (such as the utilization rate of each CPU core, the allocation and release frequency of memory pages, the specific flow of network data packets, etc.); and the aggregated measurement results are macro information that is summarized, averaged, or statistically analyzed by fine-grained data (such as the average CPU usage per minute, the total memory consumption at the hourly level, the network bandwidth peak, etc.).
[0093] User permission levels are usually set based on user type, service agreement, or configuration. For example, high-privilege users may include enterprise administrators, paid advanced users, or authorized auditors. These users have higher requirements for resource usage details (such as for cost accounting, performance optimization, or problem troubleshooting). Ordinary users are general individual users or basic service users. Their core requirement is to understand the overall resource consumption to confirm the rationality of billing, without the need to obtain underlying detailed data.
[0094] Dynamic adjustment of monitoring granularity relies on the collaborative working mechanism between the performance monitoring subsystem and pre-set probes (such as eBPF probes). When a user initiates a resource monitoring data request, the system first verifies the user's privilege level through the permission management module. For high-privilege users, the system triggers fine-grained monitoring logic. The pre-set probe maintains a high sampling frequency (e.g., microsecond or millisecond). The performance monitoring subsystem directly returns raw or low-aggregate resource usage data, such as real-time push notifications of the scheduling timestamp of each CPU process, the size and duration of each memory allocation, and the packet volume of each network connection. This ensures that high-privilege users can obtain resource consumption information accurate to the operation level. For ordinary users, the system automatically activates coarse-grained aggregation logic. The probe reduces the sampling frequency or only returns pre-aggregated results. The performance monitoring subsystem uses algorithms such as sliding window aggregation and time slice aggregation to process the fine-grained data into macro-statistics (e.g., average CPU usage over a five-minute period, memory peaks and valleys, and total network traffic), filtering out sensitive or redundant details such as process- and operation-level data.
[0095] This dynamic adjustment mechanism does not affect the core advantages of lightweight monitoring (the low overhead of pre-set probes and the efficient data collection capabilities of the performance monitoring subsystem), but also achieves differentiated data supply through the permission dimension: Fine-grained details are available to high-privilege users to meet their in-depth analysis needs; aggregated results are displayed to ordinary users, reducing data transmission volume and user comprehension costs while preventing the leakage of sensitive details. For example, as a high-privilege user, an enterprise administrator can use the system to obtain detailed CPU core usage for each AI inference task on a specific server to optimize task scheduling; while ordinary users can only view hourly statistics on the total memory consumption under their own account, which not only clarifies the billing basis but also eliminates the need to access the underlying technical details. This design enables lightweight monitoring of virtualized resources to adapt to the actual needs of different users, but also reduces system storage and transmission pressure by providing data granularity on demand, further enhancing the practicality and flexibility of the monitoring mechanism.
[0096] In some embodiments, after generating resource measurement results across time granularities, the method further includes: The measurement results are synchronized to the distributed ledger according to the preset period to form an unalterable billing record chain, supporting multi-node reconciliation and audit tracing.
[0097] Cross-time granularity resource metering results cover the nanosecond-level raw data features collected at the hardware layer and the millisecond-level / window-level metering information aggregated at the software layer. They include both microscopic resource execution details (such as GPU kernel function startup timestamps and register state changes) and macroscopic load statistics (such as resource utilization peaks within a 50ms window and minute-level cumulative usage). These results are the core basis for billing calculations, and their integrity and immutability directly affect billing credibility.
[0098] The preset period refers to the synchronization interval set based on business needs (e.g., every 5 minutes or every hour), ensuring that metering results are regularly consolidated after integrity verification. The synchronization process to the distributed ledger relies on encryption and consensus mechanisms. First, the system consolidates metering results across time granularities according to a preset format (e.g., concatenating timestamps, resource types, micro-metering details, and macro-aggregate values). A unique hash value is calculated using the SHA-3-256 encryption algorithm to generate a digital fingerprint of the data. This hash value is then associated with the corresponding metering result metadata (e.g., synchronization period identifier, node ID) to form the record to be synchronized. The distributed ledger utilizes a chained structure based on the Merkle Patricia Tree (MPT). The hash of the metering results for each preset period serves as the leaf node of the MPT. The hash values of adjacent leaf nodes are aggregated layer by layer to form a block-level root hash. The root hash of each block is linked to the root hash of the previous block, forming a continuous chain of billing records. This chained structure ensures that any tampering with historical metering results invalidates the hash values of subsequent blocks, thus ensuring the immutability of the records.
[0099] To further strengthen the consensus validity of the distributed ledger, the system introduces the BLS multi-signature aggregation protocol: multiple regional nodes (such as edge computing resources and cloud computing resources) independently verify the hash of the measurement results for the same period. After verification, each generates a digital signature, and then merges the multiple signatures into an aggregate signature through a multi-signature aggregation algorithm. The records synchronized to the distributed ledger not only contain the measurement result hash and MPT block information, but also attach this aggregate signature to ensure that the record has been confirmed by multi-node consensus and prevent the data deviation of a single node from affecting the overall credibility. For example, when synchronizing the measurement results of a certain period, edge computing resource A and cloud computing resource B respectively verify the data integrity and sign. The system aggregates the two signatures and stores them in the ledger. Subsequent reconciliation only requires verifying the validity of the aggregate signature to confirm that the record has been approved by multiple nodes.
[0100] The resulting tamper-proof billing record chain supports multi-node reconciliation and audit tracing: During multi-node reconciliation, each node can obtain a complete record chain through the ledger synchronization mechanism, and quickly verify data consistency by comparing the locally stored measurement result hash with the block hash in the ledger. If there is a deviation, the record of the specific period can be located for detailed inspection; during audit tracing, the auditor or user can enter search conditions such as time range and resource type through the ledger query interface to obtain the measurement result hash and MPT verification path for the corresponding period. By recalculating the hash value of the measurement result locally and comparing it with the ledger record, and verifying the validity of the multi-signature aggregate signature, the authenticity and integrity of the measurement result can be confirmed. For example, when corporate auditors need to trace the basis for GPU billing for a certain hour, they can obtain the measurement result hash, MPT root hash, and aggregate signature for that period from the distributed ledger. By verifying the hash consistency and signature validity, they can confirm that the measurement result has not been tampered with, thus achieving full-link traceability. This record chain design based on a distributed ledger upgrades the measurement results across time granularity from single-point storage to tamper-proof records based on multi-node consensus, providing underlying technical support for the credibility of billing data.
[0101] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0102] The embodiment of the present application also provides a resource metering device, Figure 2 A schematic diagram of a resource metering device according to an embodiment of the present disclosure is shown in FIG. Figure 2 Shown, including: An acquisition unit 21 is used to acquire the execution timestamp and register status of heterogeneous resources; An extraction unit 22 is configured to dynamically merge resource usage features within adjacent windows of execution timestamps and register states of the heterogeneous resources based on a sliding window aggregation algorithm with a preset step size and a preset duration, and extract metering information of the burst load; A determination unit 23 is configured to determine a real-time charging rule based on a pre-trained pricing model, wherein the pre-trained pricing model is trained based on real-time resource supply and demand and user characteristics; The determining unit 23 is further configured to determine charging for the heterogeneous resources based on the real-time charging rule.
[0103] Furthermore, in a possible implementation of the embodiment of the present disclosure, the acquiring unit 21 is further configured to: Obtain the start and end timestamps of GPU kernel functions based on the performance analysis interface, and obtain register status to reflect the real-time usage of hardware resources; Based on the performance monitoring subsystem and preset probes, lightweight monitoring is performed on virtualized resources; wherein the virtualized resources include at least one of a central processing unit, memory, and network bandwidth.
[0104] Furthermore, in a possible implementation of the embodiment of the present disclosure, the extraction unit 22 is further configured to: Dynamically adjusting the preset step size based on real-time resource usage characteristics; In response to a sudden load fluctuation, the preset step size is reduced.
[0105] Furthermore, in a possible implementation of the embodiment of the present disclosure, the acquiring unit 21 is further configured to: Inserting monitoring instructions into the graphics processor through instruction instrumentation to capture resource usage events at the graphics processor kernel function level; The preset probe is used to collect the usage status of virtualized resources in real time in the operating system kernel, and perform timestamp alignment processing.
[0106] Furthermore, in a possible implementation of the embodiment of the present disclosure, the determining unit 23 is further configured to: Based on the pre-trained pricing model, the regional resource supply-demand ratio and user consumption behavior characteristics are calculated in real time to generate dynamic price adjustment signals; Distributed training based on the policy gradient algorithm is used to optimize the price adjustment strategy.
[0107] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 3 As shown, the device also includes: A collection unit 24, configured to collect execution timestamps and register states of the heterogeneous resources based on a hardware layer; The collection unit 24 is further configured to obtain the resource usage characteristics based on the software layer and extract metering information of the burst load.
[0108] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 3 As shown, the device also includes: The fusion unit 25 is used to fuse the original data collected by the hardware layer with the feature data extracted by the software layer after the collection unit 24 obtains the resource usage characteristics based on the software layer and extracts the metering information of the burst load, so as to generate a resource metering result across time granularity.
[0109] Furthermore, in a possible implementation of the embodiment of the present disclosure, the determining unit 23 is further configured to: The intermediate data generated during the billing process is encrypted to generate a billing voucher with a timestamp, which is used by the user end to verify the authenticity of the billing result.
[0110] Furthermore, in a possible implementation of the embodiment of the present disclosure, the apparatus further includes: The extraction unit 22 is further configured to gradually increase the preset step size to an initial value in response to the burst load recovering to a stable state after the extraction unit 22 reduces the preset step size in response to the burst load fluctuation.
[0111] Furthermore, in a possible implementation of the embodiment of the present disclosure, the heterogeneous resources include edge computing resources and cloud computing resources, and the determining unit 23 is further configured to: The billing coefficient is adjusted based on the communication loss characteristics of the edge computing resources and the cloud computing resources.
[0112] Furthermore, in a possible implementation of the embodiment of the present disclosure, the acquiring unit 21 is further configured to: Dynamically adjust monitoring granularity based on user permission levels; open fine-grained resource usage details to high-privilege users, and display aggregated metering results to ordinary users.
[0113] Furthermore, in a possible implementation of the embodiment of the present disclosure, as Figure 3 As shown, the device also includes: The synchronization unit 26 is used to synchronize the resource measurement results across time granularity generated by the fusion unit 25 to the distributed ledger according to a preset period, forming an unalterable billing record chain and supporting multi-node reconciliation and audit tracing.
[0114] For the description of the features in the embodiment corresponding to the resource metering device, reference can be made to the relevant description of the embodiment corresponding to the resource metering method, which will not be repeated here.
[0115] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above resource metering method embodiments.
[0116] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned resource metering method embodiments when running.
[0117] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0118] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned resource metering method embodiments are implemented.
[0119] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned resource metering method embodiments are implemented.
[0120] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0121] The above is a detailed introduction to a resource metering method, device, electronic device and storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A resource metering method, characterized in that: include: Obtain execution timestamps and register states of heterogeneous resources; Based on a sliding window aggregation algorithm, dynamically merge the resource usage features in adjacent windows of the execution timestamps and register states of the heterogeneous resources with a preset step size and a preset duration, and extract metering information of the burst load; Determining real-time billing rules based on a pre-trained pricing model; wherein the pre-trained pricing model is trained based on real-time resource supply and demand and user characteristics; The charging for the heterogeneous resources is determined based on the real-time charging rule.
2. The method according to claim 1, characterized in that The obtaining of the execution timestamp and register status of the heterogeneous resources includes: Obtain the start and end timestamps of GPU kernel functions based on the performance analysis interface, and obtain register status to reflect the real-time usage of hardware resources; Based on the performance monitoring subsystem and preset probes, lightweight monitoring is performed on virtualized resources; wherein the virtualized resources include at least one of a central processing unit, memory, and network bandwidth.
3. The method according to claim 1, characterized in that The method of dynamically merging the resource usage features in adjacent windows of the execution timestamps and register states of the heterogeneous resources with a preset step size and a preset duration based on a sliding window aggregation algorithm and extracting the metering information of the burst load further includes: Dynamically adjusting the preset step size based on real-time resource usage characteristics; In response to a sudden load fluctuation, the preset step size is reduced.
4. The method according to claim 2, characterized in that The obtaining of the execution timestamp and register status of the heterogeneous resources further includes: Inserting monitoring instructions into the graphics processor through instruction instrumentation to capture resource usage events at the graphics processor kernel function level; The preset probe is used to collect the usage status of virtualized resources in real time in the operating system kernel, and perform timestamp alignment processing.
5. The method according to claim 1, wherein Determining the real-time charging rules based on the pre-trained pricing model further includes: Based on the pre-trained pricing model, the regional resource supply-demand ratio and user consumption behavior characteristics are calculated in real time to generate dynamic price adjustment signals; Distributed training based on the policy gradient algorithm is used to optimize the price adjustment strategy.
6. The method according to claim 1, characterized in that The method further comprises: Collecting execution timestamps and register states of the heterogeneous resources based on the hardware layer; The resource usage characteristics are obtained based on the software layer, and metering information of the burst load is extracted.
7. The method according to claim 6, characterized in that After obtaining the resource usage characteristics based on the software layer and extracting the metering information of the burst load, the method further includes: The raw data collected by the hardware layer is integrated with the feature data extracted by the software layer to generate resource measurement results across time granularity.
8. The method according to claim 1, wherein The determining of charging for the heterogeneous resources based on the real-time charging rule further includes: The intermediate data generated during the billing process is encrypted to generate a billing voucher with a timestamp, which is used by the user end to verify the authenticity of the billing result.
9. The method according to claim 3, wherein After reducing the preset step size in response to a sudden load fluctuation, the method further includes: In response to the burst load returning to a stable state, the preset step size is gradually increased to an initial value.
10. The method according to claim 1, wherein The heterogeneous resources include edge computing resources and cloud computing resources, and the determining of real-time billing rules based on the pre-trained pricing model further includes: The billing coefficient is adjusted based on the communication loss characteristics of the edge computing resources and the cloud computing resources.
11. The method according to claim 2, wherein: The lightweight monitoring of virtualized resources based on the performance monitoring subsystem and the preset probes further includes: Dynamically adjust monitoring granularity based on user permission levels; open fine-grained resource usage details to high-privilege users, and display aggregated metering results to ordinary users.
12. The method according to claim 7, wherein: After generating the resource measurement results across time granularities, the method further includes: The measurement results are synchronized to the distributed ledger according to the preset period to form an unalterable billing record chain, supporting multi-node reconciliation and audit tracing.
13. A resource metering device, characterized in that: include: An acquisition unit, used to obtain the execution timestamp and register status of heterogeneous resources; an extraction unit, configured to dynamically merge resource usage features within adjacent windows of execution timestamps and register states of the heterogeneous resources with a preset step size and a preset duration based on a sliding window aggregation algorithm, and extract metering information of the burst load; A determination unit, configured to determine a real-time charging rule based on a pre-trained pricing model; wherein the pre-trained pricing model is trained based on real-time resource supply and demand and user characteristics; The determining unit is further configured to determine charging for the heterogeneous resources based on the real-time charging rule.
14. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the resource metering method according to any one of claims 1 to 12 when executing the computer program.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the resource metering method according to any one of claims 1 to 12 when executed by a processor.
Citation Information
Patent Citations
Methods, systems, and apparatus for multi-purpose metering
CA2615262A1
Discontinuous reception for a two-step random access procedure
CA3155060A1
Resource charging method and apparatus
CN107360006A
Embedded online federated learning
CN114600106A
Cloud resource metering and charging device and method, equipment and storage medium
CN115134179A
Cited By
Multi-dimensional fine-grained charging method and system for GPU (Graphics Processing Unit) computing resources
CN121326583A
Process-level power consumption analysis method and device, electronic equipment and storage medium
CN121412076A
AI server management system based on GPU computing power optimization
CN121785788A
Performance data analysis system, method and device, electronic equipment and storage medium
CN121935122A