Method and device for determining hardware utilization rate, equipment, medium and program product

By updating task execution time through statistical analysis and dynamic compensation, the problem of inaccurate hardware utilization in heterogeneous computing systems is solved, enabling more precise resource allocation and energy consumption optimization.

CN121008903APending Publication Date: 2025-11-25SHANGHAI CAMBRICON INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202410659698.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-24
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately determine the hardware utilization of heterogeneous computing systems, leading to inaccurate allocation and scheduling of computing resources and an inability to meet the needs of multi-user scenarios.

Method used

The hardware utilization rate is calculated by statistically analyzing the total execution time of tasks performed on the hardware, updating the total execution time using dynamic compensation, and combining this with the task execution data within a specified time period.

Benefits of technology

It improves the accuracy of hardware utilization determination, realizes dynamic load adjustment and energy consumption optimization of heterogeneous computing systems, and supports high-precision computing resource allocation and scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121008903A_ABST
    Figure CN121008903A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of hardware utilization rate calculation, and discloses a method and device for determining the hardware utilization rate, equipment, a medium and a program product, and the method for determining the hardware utilization rate comprises the following steps: counting first total execution time and second total execution time of tasks executed on hardware, updating the first total execution time based on the first dynamic compensation time, and updating the second total execution time based on the second dynamic compensation time; the first dynamic compensation time and the second dynamic compensation time are used for compensating the execution time of the task executed on the hardware; and determining a hardware utilization rate in a specified time period by utilizing the updated first total execution time, the updated second total execution time and the specified time period. Compared with the prior art, the accuracy of determining the hardware utilization rate in the specified time period can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hardware utilization calculation technology, and specifically to methods, apparatus, equipment, media, and program products for determining hardware utilization. Background Technology

[0002] Heterogeneous computing systems typically comprise multiple computing units using different instruction sets and / or architectures. Common computing units include, but are not limited to, GPUs (Graphics Processing Units), CPUs (Central Processing Units), DSPs (Digital Signal Processors), FPGAs (Field Programmable Gate Arrays), and ASICs (Application Specific Integrated Circuits). Given the complexity of heterogeneous computing systems, the rational scheduling of computing tasks, the efficient utilization of computing resources, and the appropriate allocation of computing resources are particularly important. Whether it's precisely allocating computing resources or precisely scheduling computing tasks, the foundation for both relies on the utilization rate of hardware resources (i.e., hardware utilization). Therefore, accurately determining hardware utilization is crucial. Summary of the Invention

[0003] In view of this, the present invention provides a method, apparatus, device, medium, and program product for determining hardware utilization, in order to solve the problem of how to accurately determine hardware utilization.

[0004] In a first aspect, the present invention provides a method for determining hardware utilization, the method comprising:

[0005] The first total execution time and the second total execution time of the tasks executed on the hardware are calculated. The first total execution time is the sum of the execution times of the tasks executed on the hardware before a first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before a second specified time, where the second specified time occurs before the first specified time.

[0006] The first total execution time is updated based on the first dynamic compensation time, and the second total execution time is updated based on the second dynamic compensation time; the first dynamic compensation time and the second dynamic compensation time are used to compensate for the execution time of the tasks executed on the hardware.

[0007] Using the updated first total execution time, the updated second total execution time, and the specified time period, determine the hardware utilization rate within the specified time period; the specified time period is the time difference between the first specified time and the second specified time.

[0008] In a second aspect, the present invention provides an apparatus for determining hardware utilization, the apparatus comprising:

[0009] The statistics module is used to calculate the first total execution time and the second total execution time of the tasks executed on the hardware; wherein, the first total execution time is the sum of the execution times of the tasks executed on the hardware before a first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before a second specified time, the second specified time occurring before the first specified time;

[0010] The update module is used to update the first total execution time based on the first dynamic compensation time, and to update the second total execution time based on the second dynamic compensation time; the first dynamic compensation time and the second dynamic compensation time are used to compensate for the execution time of the tasks executed on the hardware.

[0011] The determination module is used to determine the hardware utilization rate within a specified time period by using the updated first total execution time, the updated second total execution time, and a specified time period; the specified time period is the time difference between the first specified time and the second specified time.

[0012] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method for determining hardware utilization described in the first aspect or any corresponding embodiment thereof.

[0013] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the method for determining hardware utilization described in the first aspect or any corresponding embodiment thereof.

[0014] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the method for determining hardware utilization described in the first aspect or any corresponding embodiment thereof.

[0015] Compared with related technologies, the present invention innovatively updates the sum of the execution times of tasks executed on the hardware before a first specified time by using a first dynamic compensation time, and updates the sum of the execution times of tasks executed on the hardware before a second specified time by using a second dynamic compensation time. By compensating for the sum of the above two execution times, the present invention can make the statistical total execution time closer to the actual total execution time, thereby improving the accuracy of determining the hardware utilization rate within a specified time period. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a schematic diagram of the board structure according to an embodiment of the present invention.

[0018] Figure 2 This is a schematic diagram of the structure of a combined processing device in a chip according to an embodiment of the present invention.

[0019] Figure 3 This is a schematic diagram of the internal structure of a single-core computing device according to an embodiment of the present invention.

[0020] Figure 4 This is a schematic diagram of the internal structure of a multi-core computing device according to an embodiment of the present invention.

[0021] Figure 5 This is a flowchart illustrating a method for determining hardware utilization according to an embodiment of the present invention.

[0022] Figure 6 This is a flowchart illustrating another method for determining hardware utilization according to an embodiment of the present invention.

[0023] Figure 7 This is a flowchart illustrating another method for determining hardware utilization according to an embodiment of the present invention.

[0024] Figure 8 This is a schematic diagram illustrating a task execution time greater than the sampling period according to an embodiment of the present invention.

[0025] Figure 9 This is a schematic diagram illustrating that the execution time of a task according to an embodiment of the present invention is less than the sampling period.

[0026] Figure 10This is a flowchart illustrating another method for determining hardware utilization according to an embodiment of the present invention.

[0027] Figure 11 This is a flowchart illustrating a method for calibrating hardware utilization according to an embodiment of the present invention.

[0028] Figure 12 This is a structural block diagram of an apparatus for determining hardware utilization according to an embodiment of the present invention.

[0029] Figure 13 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] Figure 1 A schematic diagram of the structure of a board 10 according to an embodiment of the present invention is shown. Figure 1 As shown, board 10 includes chip 101, which is a system-on-chip (SoC) integrating one or more combined processing units. These combined processing units are artificial intelligence computing units used to support various deep learning and machine learning algorithms, meeting the intelligent processing needs of complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in cloud intelligence. A significant characteristic of cloud intelligence applications is the large volume of input data, placing high demands on the platform's storage and computing capabilities. Board 10 in this embodiment is suitable for cloud intelligence applications, possessing massive off-chip storage, on-chip storage, and powerful computing capabilities. In some specific embodiments, the relationship between board 10 and the host is heterogeneous, and the system composed of board 10 and the host is a heterogeneous computing system.

[0032] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. The calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface forms, such as a PCIe interface.

[0033] The board 10 also includes a storage device 104 for storing data, which includes one or more memory cells 105. The storage device 104 is connected to and transmits data with the controller 106 and the chip 101 via a bus. The controller 106 in the board 10 is configured to regulate the state of the chip 101. Therefore, in one application scenario, the controller 106 may include a microcontroller (MCU).

[0034] Figure 2 This is a structural diagram illustrating the combined processing device in chip 101 of this embodiment. (As shown) Figure 2As shown, the combined processing device 20 includes a computing device 201, an interface device 202, a processing device 203, and a storage device 204. The computing device 201 is configured to execute user-specified operations, primarily implemented as a single-core or multi-core intelligent processor, for performing deep learning or machine learning calculations. It can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations. The interface device 202 is used to transmit data and control commands between the computing device 201 and the processing device 203. For example, the computing device 201 can obtain input data from the processing device 203 via the interface device 202 and write it to the on-chip storage device of the computing device 201. Further, the computing device 201 can obtain control commands from the processing device 203 via the interface device 202 and write them to the on-chip control cache of the computing device 201. Alternatively or optionally, the interface device 202 can also read data from the storage device of the computing device 201 and transmit it to the processing device 203. The processing device 203, as a general-purpose processing device, performs basic controls including but not limited to data transfer and starting / stopping the computing device 201. Depending on the implementation, the processing device 203 may be one or more types of processors, including but not limited to digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, the computing device 201 of this invention can be considered as having a single-core structure or a homogeneous multi-core structure. However, when the computing device 201 and the processing device 203 are considered together, they are considered to form a heterogeneous multi-core structure. Storage device 204 is used to store data to be processed. It may be DRAM or DDR memory, typically 16G or larger in size, and is used to store data of computing device 201 and / or processing device 203.

[0035] Figure 3The diagram illustrates the internal structure of the single-core computing device 201. The single-core computing device 301 processes input data for computer vision, speech, natural language processing, and data mining. It comprises three main modules: a control module 31, an arithmetic module 32, and a storage module 33. The control module 31 coordinates and controls the operation of the arithmetic module 32 and the storage module 33 to complete deep learning tasks. It includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The IFU 311 fetches instructions from the processing device 203, while the IDU 312 decodes the fetched instructions and sends the decoding results as control information to the arithmetic module 32 and the storage module 33. The arithmetic module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 performs vector operations, supporting complex operations such as vector multiplication, addition, and nonlinear transformations. The matrix operation unit 322 is responsible for the core calculations of the deep learning algorithm, namely matrix multiplication and convolution. The storage module 33 is used to store or move related data, including the neuron RAM (NRAM) 331, the weight RAM (WRAM) 332, and the direct memory access (DMA) module 333. The NRAM 331 stores the input neurons, the output neurons, and the intermediate results after calculation; the WRAM 332 stores the convolution kernels of the deep learning network, i.e., the weights; the DMA 333 is connected to the DRAM 204 via the bus 34 and is responsible for data transfer between the single-core computing device 301 and the DRAM 204.

[0036] Figure 4 A schematic diagram of the internal structure of the multi-core computing device 201 is shown. The multi-core computing device 41 adopts a hierarchical design. As a system-on-a-chip, the multi-core computing device 41 includes at least one cluster, and each cluster includes multiple processor cores. In other words, the multi-core computing device 41 is constructed in a hierarchical structure of system-on-a-chip, cluster, and processor cores. From the perspective of the system-on-a-chip hierarchy, as... Figure 4 As shown, the multi-core computing device 41 includes an external storage controller 401, a peripheral communication module 402, an on-chip interconnect module 403, a synchronization module 404, and multiple clusters 405.

[0037] There can be multiple external storage controllers 401; two are shown as an example in the figure. These controllers are used to access external storage devices, such as those issued by the processor core, in response to access requests from the processor core. Figure 2The DRAM 204 in the chip allows data to be read from or written to external memory. The peripheral communication module 402 receives control signals from the processing device 203 via the interface device 202, initiating the computing device 201 to execute tasks. The on-chip interconnect module 403 connects the external memory controller 401, the peripheral communication module 402, and multiple clusters 405 to transmit data and control signals between modules. The synchronization module 404 is a global barrier controller (GBC) used to coordinate the working progress of each cluster and ensure information synchronization. Multiple clusters 405 are the computing cores of the multi-core computing device 41. Four are shown exemplary in the figure; however, with hardware development, the multi-core computing device 41 of this invention can also include 8, 16, 64, or even more clusters 405. Clusters 405 are used to efficiently execute deep learning algorithms. From a cluster hierarchy perspective, such as... Figure 4 As shown, each cluster 405 includes multiple processor cores (IPU cores) 406 and one memory core (MEM core) 407.

[0038] Four processor cores 406 are shown as an example in the figure, but the present invention does not limit the number of processor cores 406. Its internal architecture is as follows: Figure 5 As shown. Each processor core 406 is similar to Figure 3 The single-core computing device 301 also includes three main modules: a control module 51, an arithmetic module 52, and a storage module 53. The functions and structures of the control module 51, arithmetic module 52, and storage module 53 are largely the same as those of the control module 31, arithmetic module 32, and storage module 33, and will not be described again. It should be noted that the storage module 53 includes an input / output direct memory access (IODMA) module 533 and a move direct memory access (MVDMA) module 534. The IODMA 533 controls the memory access of NRAM 531 / WRAM 532 and DRAM 204 via the broadcast bus 409; the MVDMA 534 controls the memory access of NRAM 531 / WRAM 532 and SRAM 408.

[0039] Back Figure 4The storage core 407 is primarily used for storage and communication, namely storing shared data or intermediate results among processor cores 406, and performing communication between cluster 405 and DRAM 204, communication between clusters 405, and communication between processor cores 406. In other embodiments, the storage core 407 has scalar operation capabilities for performing scalar operations. The storage core 407 includes an SRAM 408, a broadcast bus 409, a cluster direct memory access (CDMA) module 410, and a global direct memory access (GDMA) module 411. SRAM 408 acts as a high-performance data relay station. Data multiplexed between different processor cores 406 within the same cluster 405 does not need to be obtained from DRAM 204 by each processor core 406. Instead, it is relayed between processor cores 406 via SRAM 408. Storage core 407 only needs to quickly distribute the multiplexed data from SRAM 408 to multiple processor cores 406 to improve inter-core communication efficiency and greatly reduce on-chip and off-chip input / output access.

[0040] Broadcast bus 409, CDMA 410, and GDMA 411 are used to perform communication between processor cores 406, communication between clusters 405, and data transfer between cluster 405 and DRAM 204, respectively. These will be explained below. Broadcast bus 409 is used to complete high-speed communication between processor cores 406 within cluster 405. In this embodiment, broadcast bus 409 supports inter-core communication methods including unicast, multicast, and broadcast. Unicast refers to point-to-point (e.g., single processor core to single processor core) data transfer; multicast is a communication method that transfers data from SRAM 408 to several specific processor cores 406; and broadcast is a communication method that transfers data from SRAM 408 to all processor cores 406, a special case of multicast. CDMA 410 is used to control memory access to SRAM 408 between different clusters 405 within the same computing device 201. GDMA 411 works in conjunction with external memory controller 401 to control memory access from SRAM 408 to DRAM 204 in cluster 405, or to read data from DRAM 204 into SRAM 408. As previously described, communication between DRAM 204 and NRAM 431 or WRAM 432 can be achieved through two channels. The first channel is a direct connection between DRAM 204 and NRAM 431 or WRAM 432 via IODAM 433; the second channel involves first transmitting data between DRAM 204 and SRAM 408 via GDMA 411, and then transmitting data between SRAM 408 and NRAM 431 or WRAM 432 via MVDMA 534. Although the second channel appears to require more components and has a longer data flow, in some embodiments, the bandwidth of the second channel is significantly greater than that of the first channel. Therefore, communication between DRAM 204 and NRAM 431 or WRAM 432 may be more efficient via the second channel. The embodiments of the present invention can select the data transmission channel according to their own hardware conditions.

[0041] In other embodiments, the functions of GDMA 411 and IODMA 533 can be integrated into the same component. For ease of description, GDMA 411 and IODMA 533 are considered different components. For those skilled in the art, any component that performs similar functions and achieves similar technical effects to this invention falls within the scope of protection of this invention. Furthermore, the functions of GDMA 411, IODMA 533, CDMA 410, and MVDMA 534 can also be implemented by the same component.

[0042] Heterogeneous computing systems are widely used in related technologies, but their effectiveness and performance are limited by hardware resource utilization. Accurate hardware utilization allows for dynamic load adjustment, improving overall performance, reducing energy consumption, and optimizing task scheduling. Furthermore, given the high computing power of heterogeneous computing systems, virtualization and container technologies (such as Docker) can be used to partition computing power, precisely controlling the available computing power or resources for each user. However, this partitioning relies on high-precision, high-resolution hardware utilization determination. Higher resolution and accuracy allow for more precise adjustments to the operation of the heterogeneous computing system, better leveraging its performance.

[0043] While utilization rates can be obtained through hardware design, this increases hardware design complexity and chip area. Furthermore, processors in heterogeneous computing systems typically have multiple computing units, requiring real-time calculation of hardware utilization to be performed within each unit, necessitating independent hardware circuitry to track working and idle times, further increasing hardware complexity and chip area. Moreover, real-time output of utilization data is required during operation; higher sampling resolution results in more data output, placing greater demands on software processing capabilities and consuming additional resources. In practical applications, hardware utilization statistics based on time periods are insufficient. Hardware resources are often time-division multiplexed by multiple users. Utilization results based on time periods only represent the overall hardware utilization, not the precise amount of hardware resources used by each user within a specific timeframe. In other words, for scenarios where multiple users simultaneously utilize the same computing resources, the hardware utilization of each individual user cannot be statistically analyzed. Furthermore, hardware utilization statistics based on task granularity cannot meet the requirements. Task execution time is often uncontrollable and is not synchronized with software sampling and calculation time. For example, when the task execution time is longer than the sampling interval, the hardware utilization will be judged as 0 for a period of time because no task is completed for the software during this period. The utilization will be greater than 100% for another period of time because the task execution time is longer than the sampling time for the software. Obviously, the hardware utilization judgment results in related technologies are inaccurate.

[0044] According to an embodiment of the present invention, a method embodiment for determining hardware utilization is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0045] One embodiment of the present invention provides a method for determining hardware utilization, which can be used in the aforementioned controller 106. The controller 106 of the present invention can be a coprocessor. Optionally, a driver program can run on the controller 106. The driver program is a software program used to control the operation of the chip 101. The method for determining hardware utilization in this embodiment of the present invention can be part of the driver program. Figure 5 This is a flowchart of a method for determining hardware utilization according to an embodiment of the present invention, such as... Figure 5 As shown, the process includes the following steps:

[0046] Step S501: Calculate the first total execution time and the second total execution time of the tasks executed on the hardware; wherein, the first total execution time is the sum of the execution times of the tasks executed on the hardware before the first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before the second specified time, and the second specified time occurs before the first specified time.

[0047] like Figure 1 The schematic diagram of the board structure shown in this embodiment is an AI (Artificial Intelligence) chip inside the board in a heterogeneous computing system. The AI ​​chip can be the chip 101 mentioned above.

[0048] The tasks executed on the hardware include, but are not limited to, computational tasks and memory access tasks. The hardware can provide at least one task queue for storing tasks, and multiple tasks in each queue can be executed in a first-in, first-out (FIFO) order. For example, if the utilization rate of hardware resources by the tasks in a certain task queue is determined, the first total execution time and the second total execution time in this embodiment can be the time recorded by the software in the task queue.

[0049] In this embodiment of the invention, the first total execution time refers to the total execution time of tasks in the current task queue up to the first specified time, and the second total execution time refers to the total execution time of tasks in the current task queue up to the second specified time; wherein, the execution time of a task refers to the execution time of a single task; the first total execution time can be expressed as T. total1 The second total execution time can be expressed as T. total2 The first specified time point can be represented as t1, the second specified time point can be represented as t2, and the execution time of the task can be represented as T. task .

[0050] Step S502: Update the first total execution time based on the first dynamic compensation time, and update the second total execution time based on the second dynamic compensation time; the first dynamic compensation time and the second dynamic compensation time are used to compensate for the execution time of the task executed on the hardware.

[0051] The first dynamic compensation time and the second dynamic compensation time in this embodiment can be specifically determined based on software time compensation, for example, based on the aforementioned driver program.

[0052] Specifically, if the first total execution time is shorter than its corresponding actual total execution time, the first total execution time is increased through a first dynamic compensation time; if the first total execution time is longer than its corresponding actual total execution time, the first total execution time is decreased through a first dynamic compensation time, thereby achieving the goal of making the first total execution time nearly the same as its corresponding actual total execution time. Similarly, if the second total execution time is shorter than its corresponding actual total execution time, the second total execution time is increased through a second dynamic compensation time; if the second total execution time is longer than its corresponding actual total execution time, the second total execution time is decreased through a second dynamic compensation time, thereby achieving the goal of making the second total execution time nearly the same as its corresponding actual total execution time.

[0053] The first dynamic compensation time involved in this invention can be expressed as T. com1 The first dynamic compensation time can be determined based on the task's execution time. The second dynamic compensation time can be expressed as T. com2 The second dynamic compensation time can be determined based on the task's execution time.

[0054] Optionally, the first dynamic compensation time is the same as the second dynamic compensation time.

[0055] Step S503: Using the updated first total execution time, the updated second total execution time, and the specified time period, determine the hardware utilization rate within the specified time period; the specified time period is the time difference between the first specified time and the second specified time.

[0056] Specifically, in this embodiment, the difference between the updated first total execution time and the updated second total execution time can be determined, and then the hardware utilization rate within the specified time period can be determined based on the ratio of the difference to the specified time period; optionally, the ratio is used as the hardware utilization rate within the specified time period.

[0057] Specifically, the hardware utilization rate involved in this invention can be the real-time utilization rate of hardware resources; the specified time period can be represented as T. period In this embodiment, T period = t1 - t2 or T period = |t1-t2|.

[0058] In this embodiment,

[0059] This invention innovatively employs a first dynamic compensation time to sum the execution times T of tasks executed on the hardware before a first specified time. total1 The update is performed, and the sum of the execution times T of the tasks executed on the hardware before the second specified time is used as the second dynamic compensation time. total2 By updating and compensating for the sum of the two execution times mentioned above, this invention can make the statistical total execution time closer to the actual total execution time, thereby improving the accuracy of determining the hardware utilization rate within a specified time period.

[0060] One embodiment of the present invention provides a method for determining hardware utilization, which can be used in the aforementioned controller 106. The controller 106 of the present invention can be a coprocessor. Optionally, a driver program can run on the controller 106. The driver program is a software program used to control the operation of the chip 101. The method for determining hardware utilization in this embodiment of the present invention can be part of the driver program. Figure 6 This is a flowchart of a method for determining hardware utilization according to an embodiment of the present invention, such as... Figure 6 As shown, the process includes the following steps:

[0061] Step S601: Calculate the first total execution time and the second total execution time of the tasks executed on the hardware; wherein, the first total execution time is the sum of the execution times of the tasks executed on the hardware before the first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before the second specified time, and the second specified time occurs before the first specified time.

[0062] Specifically, step S601 includes steps S6011 and S6022, which are executed in parallel or sequentially.

[0063] Step S6011: If the execution time of the task executed on the hardware before the first specified time is greater than the sampling period, then every time a sampling period has elapsed and the task is being executed, the first total execution time and the first dynamic compensation time are both increased by one sampling period.

[0064] In cases where the execution time of a task is relatively small compared to the sampling period, this embodiment can update the first total execution time based on the execution time of the task executed on the hardware before the first specified time being less than or equal to the sampling period. The updated first total execution time is equal to the sum of the first total execution time before the update and the execution time of the task.

[0065] In specific implementation, the driver program of this embodiment can check the execution status of tasks on the hardware at preset time intervals to determine hardware utilization. The preset time interval is the sampling period, which can be expressed as T. sample The present invention can set the specific value of the sampling period according to actual needs, such as 3 seconds, 1 second or 0.5 seconds.

[0066] like Figure 8 and Figure 9 As shown, the time point at the end of each sampling period is indicated by a black arrow. To distinguish it from the previously mentioned task execution time, the task execution time T for different duration ranges is represented differently. task This embodiment will use a duration greater than the sampling time T. sample The task execution time is represented as T. task0 and the duration is less than or equal to the sampling T sample The task execution time is represented as T. task1 Regarding the task execution time T task0 Greater than the sampling period T sample The situation (T) task0 >T sample The total execution time T of the task is as follows: (1) After each sampling period (indicated by the black arrow), while the task is being executed. total =T total +T sample In this embodiment, T total1 =T total1 +T sample T com1 =T com1 +T sample .

[0067] like Figure 9 As shown, the time point at the end of the sampling period is indicated by a black arrow. The execution time T of the task... task1 Less than or equal to the sampling period T sample The situation (T) task1 ≤T sample When the sampling period ends, from the perspective of the software (e.g., the driver), the task is not currently executing, so the total task execution time T is... total =T total +T task1 In this embodiment, the first total execution time T total1 =T total1 +T task1 Specifically, for T task1 ≤T sample In this embodiment, the hardware reports the execution time of the current task to the coprocessor after the task is completed.

[0068] Step S6012: If the execution time of the task executed on the hardware before the second specified time is greater than the sampling period, then every time a sampling period has elapsed and the task is being executed, the second total execution time and the second dynamic compensation time are both increased by one sampling period.

[0069] In cases where the execution time of a task is relatively short compared to the sampling period, this embodiment can update the second total execution time based on the execution time of the task executed on the hardware before the second specified time being less than or equal to the sampling period. The updated second total execution time is equal to the sum of the previous second total execution time and the task execution time.

[0070] It should be understood that this embodiment does not limit the execution order of steps S6011 and S6012, and the steps S6011 and S6012 involved can be two steps executed in parallel.

[0071] like Figure 8 As shown, for the task execution time T task0 Greater than the sampling period T sample The situation (T) task0 >T sample After each sampling period (indicated by the black arrow), the total execution time T of the task is... total =T total +T sample In this embodiment, T total2 =T total2 +T sample T com2 =T com2 +T sample .

[0072] like Figure 9 As shown, for the task execution time T task1 Less than or equal to the sampling period T sample The situation (T) task1 ≤T sample When the sampling period ends, from a software perspective (e.g., the driver), the task is not currently executing; therefore, the total task execution time T is... total =T total +T task1 In this embodiment, the first total execution time T total2 =T total2 +T task1 Specifically, for T task1 ≤T sample In this embodiment, the hardware reports the execution time of the current task to the coprocessor after the task is completed.

[0073] This embodiment uses the relationship between the task execution time and the sampling period as the statistical basis for the total task execution time. When the sampling period is smaller than the task execution time, the granularity for increasing the total execution time is the sampling period; conversely, when the sampling period is larger than the task execution time, the actual reported execution time of the task after execution is used as the basis for increasing the total time. This method can more accurately calculate the first and second total execution times of tasks executed on the hardware. Furthermore, when updating the total execution time when the task execution time is greater than the sampling period, this embodiment also increases the dynamic compensation time with the sampling period as the granularity. This method achieves the accumulation of time, thereby providing data for the next comparison between the dynamic compensation time and the task execution time, and providing a basis for the accurate determination of subsequent hardware time utilization.

[0074] In some alternative implementations, the method of the present invention for determining hardware utilization further includes steps a1 and a2, which are performed sequentially.

[0075] In step a1, during the process of updating the first total execution time and the second total execution time, a third ratio is determined between the number of threads used by the task and the total number of threads configured on the hardware.

[0076] Typically, each task on a heterogeneous computing platform (heterogeneous computing processor) starts multiple threads to work in parallel. This embodiment uses threads... task This indicates the number of threads used by the task (the total number of threads used by the task), expressed as `thread`. total This indicates the total number of threads on a heterogeneous computing platform (the total number of threads configured on the hardware).

[0077] In this embodiment, a third ratio of the number of threads used by the task to the total number of threads configured on the hardware is a weight.

[0078] Step a2: Use the third ratio to calibrate the execution time and / or sampling period of the task. The execution time of the calibrated task is the product of the execution time of the task before calibration and the third ratio. The sampling period after calibration is the product of the sampling period before calibration and the third ratio.

[0079] In this embodiment, the execution time T' of the calibrated task task =T task *Weight, the calibrated sampling period T' sample =T sample *Weight. Optionally, this embodiment determines the first total execution time and the second total execution time based on the calibrated task execution time and / or the calibrated sampling period.

[0080] For modern heterogeneous computing platforms with multiple computing cores, this embodiment can also perform weighted processing on the execution time and sampling period (i.e., sampling interval) of tasks. Specifically, the weighting is based on the actual resources used during runtime. This method further improves the accuracy of the hardware utilization calculation results.

[0081] Step S602: Update the first total execution time based on the first dynamic compensation time, and update the second total execution time based on the second dynamic compensation time; the first and second dynamic compensation times are used to compensate for the execution time of tasks executed on the hardware. For details, please refer to... Figure 5 Step S502 of the illustrated embodiment will not be described again here.

[0082] Step S603: Using the updated first total execution time, the updated second total execution time, and the specified time period, determine the hardware utilization within the specified time period; the specified time period is the time difference between the first specified time and the second specified time. For details, please refer to [link to details]. Figure 5 Step S503 of the illustrated embodiment will not be described again here.

[0083] It should be understood that, Figure 6 One or more embodiments of the corresponding method for determining hardware utilization may be implemented in conjunction with, where feasible, a method for determining hardware utilization. Figure 5 One or more embodiments of the corresponding method for determining hardware utilization can be reasonably combined, and the combined embodiments and their reasonable modifications are also within the protection scope of this invention.

[0084] One embodiment of the present invention provides a method for determining hardware utilization, which can be used in the aforementioned controller 106. The controller 106 of the present invention can be a coprocessor. Optionally, a driver program can run on the controller 106. The driver program is a software program used to control the operation of the chip 101. The method for determining hardware utilization in this embodiment of the present invention can be part of the driver program. Figure 7 This is a flowchart of a method for determining hardware utilization according to an embodiment of the present invention, such as... Figure 7 As shown, the process includes the following steps:

[0085] Step S701: Calculate the first total execution time and the second total execution time of the tasks executed on the hardware; wherein the first total execution time is the sum of the execution times of the tasks executed on the hardware before a first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before a second specified time, where the second specified time occurs before the first specified time. For details, please refer to [link to relevant documentation]. Figure 6 Step S601 of the illustrated embodiment will not be described again here.

[0086] Step S702: Update the first total execution time based on the first dynamic compensation time, and update the second total execution time based on the second dynamic compensation time; the first dynamic compensation time and the second dynamic compensation time are used to compensate for the execution time of the task executed on the hardware.

[0087] like Figure 8 As shown, for the task execution time T task0 Greater than the sampling period T sample The situation (T) task0 >T sample Assuming the task starts T after the first sampling period... delta1 Execution begins after a certain time, and ends before the end of the third sampling period (T). delta2 If the task is completed after a certain time, the statistical deviation T of the total task execution time is directly calculated based on the sampling period. 偏差 =3T sample -T delta1 -T delta2 The present invention can achieve a monotonically increasing time record when tasks are continuously issued by introducing dynamic time compensation, as described in the preceding embodiments. total1 =T total1 +T sample T com1 =T com1 +T sample T total2 =T total2 +T sample T com2 =T com2 +T sample Then, the statistical bias is effectively eliminated through the following examples.

[0088] Specifically, in this embodiment, step S702 includes steps S7021 and S7022, which are executed in parallel or sequentially.

[0089] Step S7021: When each task executed on the hardware is completed, if the execution time of the task is greater than the first dynamic compensation time, the first total execution time is updated according to the first difference between the execution time of the task and the first dynamic compensation time. The updated first total execution time is equal to the sum of the first total execution time before the update and the first difference. The first dynamic compensation time is reset to the first preset value, which is the initial value of the first dynamic compensation time.

[0090] In this embodiment, the update process of the first total execution time and the reset process of the first dynamic compensation time can be executed synchronously.

[0091] Wherein, the first difference = T task -T com1The first preset value can be, for example, 0.

[0092] When each task executed on the hardware completes, the execution time T of the task is... task Greater than the first dynamic compensation time T com1 The situation (T) task >T com1 In this embodiment, the first total execution time is specifically corrected in the following way: T total1 =T total1 +(T task -T com1 ), and let T com1 =0.

[0093] In this embodiment, when each task executed on the hardware is completed, if the execution time of the task is less than or equal to the first dynamic compensation time, the first dynamic compensation time is updated using the second difference between the first dynamic compensation time and the execution time of the task.

[0094] Wherein, the second difference = T com1 -T task .

[0095] When each task executed on the hardware completes, the execution time T of the task is... task Less than or equal to the first dynamic compensation time T com1 The situation (T) task ≤T com1 In this embodiment, the first dynamic compensation time T can be corrected in the following way. com1 T com1 =T com1 -T task .

[0096] Step S7022: When each task executed on the hardware is completed, if the execution time of the task is greater than the second dynamic compensation time, the second total execution time is updated according to the third difference between the execution time of the task and the second dynamic compensation time. The updated second total execution time is equal to the sum of the previous second total execution time and the third difference. The second dynamic compensation time is reset to the second preset value, which is the initial value of the second dynamic compensation time.

[0097] In this embodiment, the update process of the first total execution time and the reset process of the first dynamic compensation time can be executed synchronously.

[0098] Wherein, the third difference = T task -T com2 The second preset value can be 0.

[0099] When each task executed on the hardware completes, the execution time T of the task is... taskGreater than the second dynamic compensation time T com2 The situation (T) task >T com2 In this embodiment, the second total execution time is specifically corrected in the following way: T total2 =T total2 +(T task -T com2 ), and T com2 =0.

[0100] In this embodiment, when each task executed on the hardware is completed, if the execution time of the task is less than or equal to the second dynamic compensation time, the second dynamic compensation time is updated using the fourth difference between the second dynamic compensation time and the execution time of the task.

[0101] Wherein, the fourth difference = T com2 -T task .

[0102] When each task executed on the hardware completes, the execution time T of the task is... task Less than or equal to the second dynamic compensation time T com2 The situation (T) task ≤T com2 In this embodiment, the second dynamic compensation time T can be corrected in the following way. com2 T com2 =T com2 -T task .

[0103] Step S703: Using the updated first total execution time, the updated second total execution time, and the specified time period, determine the hardware utilization rate within the specified time period; the specified time period is the time difference between the first specified time and the second specified time.

[0104] Specifically, the hardware utilization rate in this embodiment can be represented as follows:

[0105]

[0106] In this embodiment of the invention, each sampling period is taken as the execution time of the task within that time period, and these times are accumulated into the total task execution time and the dynamic compensation time. When the task is completed, this embodiment can check the dynamic compensation time by comparing the task execution time with the dynamic compensation time. If the accumulated dynamic compensation time is too long (T... task0 ≤T com1 If the total task execution time is short (T), then the total task execution time will not be increased. task0 >T com1If the time that is not accumulated is added to the total time of task execution, then the present invention can realize the calculation result of hardware utilization based on software time compensation, which greatly improves the accuracy of the calculation result of hardware utilization.

[0107] It should be understood that, Figure 7 One or more embodiments of the corresponding method for determining hardware utilization may be implemented in conjunction with, where feasible, a method for determining hardware utilization. Figure 6 , Figure 5 The invention may reasonably combine one or more embodiments of the method for determining hardware utilization corresponding to at least one of the accompanying drawings, and the combined embodiments and their reasonable modifications are also within the protection scope of the invention.

[0108] One embodiment of the present invention provides a method for determining hardware utilization, which can be used in the aforementioned controller 106. The controller 106 of the present invention can be a coprocessor. Optionally, a driver program can run on the controller 106. The driver program is a software program used to control the operation of the chip 101. The method for determining hardware utilization in this embodiment of the present invention can be part of the driver program. Figure 10 This is a flowchart of a method for determining hardware utilization according to an embodiment of the present invention, such as... Figure 10 As shown, the process includes the following steps:

[0109] Step S1001: Calculate the first total execution time and the second total execution time of the tasks executed on the hardware; wherein the first total execution time is the sum of the execution times of the tasks executed on the hardware before a first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before a second specified time, where the second specified time occurs before the first specified time. For details, please refer to [link to details]. Figure 7 Step S701 of the illustrated embodiment will not be described again here.

[0110] Step S1002: Update the first total execution time based on the first dynamic compensation time, and update the second total execution time based on the second dynamic compensation time; the first and second dynamic compensation times are used to compensate for the execution time of tasks executed on the hardware. For details, please refer to... Figure 7 Step S702 of the illustrated embodiment will not be described again here.

[0111] Step S1003: Using the updated first total execution time, the updated second total execution time, and the specified time period, determine the hardware utilization rate within the specified time period; the specified time period is the time difference between the first specified time and the second specified time.

[0112] Specifically, step S1003 includes steps S10031 and S10032, which are executed sequentially.

[0113] Step S10031: Determine the fifth difference between the updated first total execution time and the updated second total execution time.

[0114] In this embodiment, the fifth difference = T total1 -T total2 .

[0115] Step S10032: Determine the first ratio between the fifth difference and the specified time period, and determine the first ratio as the first hardware utilization rate of the tasks in the task queue during the specified time period; the hardware utilization rate includes the first utilization rate.

[0116] Specifically, in this embodiment, the task queue represents the first utilization of hardware by tasks within a specified time period. q This can be represented as follows:

[0117]

[0118] Based on the above embodiments, the present invention accurately records the cumulative execution time of tasks in each corresponding task queue, thereby enabling accurate and real-time calculation of the utilization rate of hardware resources of tasks over any given period of time.

[0119] It should be understood that, Figure 10 One or more embodiments of the corresponding method for determining hardware utilization may be implemented in conjunction with, where feasible, a method for determining hardware utilization. Figure 7 , Figure 6 , Figure 5 The invention may reasonably combine one or more embodiments of the method for determining hardware utilization corresponding to at least one of the accompanying drawings, and the combined embodiments and their reasonable modifications are also within the protection scope of the invention.

[0120] One embodiment of the present invention provides a method for determining hardware utilization, which can be used in the aforementioned controller 106. The controller 106 of the present invention can be a coprocessor. Optionally, a driver program can run on the controller 106. The driver program is a software program used to control the operation of the chip 101. The method for determining hardware utilization in this embodiment of the present invention can be part of the driver program. Figure 11 This is a flowchart of a method for determining hardware utilization according to an embodiment of the present invention, such as... Figure 11 As shown, the process includes the following steps:

[0121] Step S1101: Calculate the first total execution time and the second total execution time of the tasks executed on the hardware; wherein, the first total execution time is the sum of the execution times of the tasks executed on the hardware before a first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before a second specified time, where the second specified time occurs before the first specified time. For details, please refer to [link to relevant documentation]. Figure 10 Step S1001 of the illustrated embodiment will not be described again here.

[0122] Step S1102: Update the first total execution time based on the first dynamic compensation time, and update the second total execution time based on the second dynamic compensation time; the first and second dynamic compensation times are used to compensate for the execution time of tasks executed on the hardware. For details, please refer to... Figure 10 Step S1002 of the illustrated embodiment will not be described again here.

[0123] Step S1103: Using the updated first total execution time, the updated second total execution time, and the specified time period, determine the hardware utilization rate within the specified time period; the specified time period is the time difference between the first specified time and the second specified time. For details, please refer to [link to details]. Figure 10 Step S1003 of the illustrated embodiment will not be described again here.

[0124] Step S11031: Determine the fifth difference between the updated first total execution time and the updated second total execution time. In this embodiment, the task is a task in the task queue. For details, please refer to... Figure 10 Step S10031 of the illustrated embodiment will not be described again here.

[0125] Step S11032: Determine the ratio of the fifth difference to the first ratio over the specified time period. This first ratio is then defined as the first hardware utilization rate of the tasks in the task queue during the specified time period; hardware utilization includes the first utilization rate. For details, please refer to [link to relevant documentation]. Figure 10 Step S10032 of the illustrated embodiment will not be described again here.

[0126] Step S1104: Determine the second ratio of the sum of the hardware utilization rates of all tasks in the task queues executed on the hardware within the specified time period to the overall hardware utilization rate.

[0127] For current hardware, such as a current chip, this embodiment can calculate the sum of hardware utilization rates of all tasks in the task queues executed on the current chip, which can be expressed as: The overall hardware utilization rate can be obtained through methods such as chip monitoring, and can be expressed as Utilization. hw .

[0128] In this embodiment,

[0129] Step S1105: The second utilization rate of the target user's tasks on the hardware is calibrated using the second ratio. The second utilization rate is the sum of the hardware utilization rates of the tasks in the task queue corresponding to the target user within a specified time period. The calibrated second utilization rate is the product of the second utilization rate before calibration and the second ratio. The hardware utilization rate includes the calibrated second utilization rate.

[0130] Regarding the second utilization rate of hardware for the target user's task, the calibrated second utilization rate in this embodiment...

[0131] Among them, Utilization queue This indicates the second utilization rate before calibration.

[0132] This invention can accurately calculate the utilization rate of hardware resources for each user and effectively calibrate the utilization rate of hardware resources for each user based on the overall hardware utilization rate, thereby achieving accuracy down to how much hardware resources each user used within a certain period of time.

[0133] It should be understood that, Figure 11 One or more embodiments of the corresponding method for determining hardware utilization may be implemented in conjunction with, where feasible, a method for determining hardware utilization. Figure 10 , Figure 7 , Figure 6 , Figure 5 The invention may reasonably combine one or more embodiments of the method for determining hardware utilization corresponding to at least one of the accompanying drawings, and the combined embodiments and their reasonable modifications are also within the protection scope of the invention.

[0134] An embodiment of the present invention also provides an apparatus for determining hardware utilization, which is used to implement the above embodiments and preferred embodiments, and will not be repeated hereafter. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0135] This embodiment provides a device for determining hardware utilization, such as... Figure 12 As shown, it includes:

[0136] The statistics module 1201 is used to calculate the first total execution time and the second total execution time of the tasks executed on the hardware; wherein, the first total execution time is the sum of the execution times of the tasks executed on the hardware before a first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before a second specified time, the second specified time occurring before the first specified time.

[0137] The update module 1202 is used to update the first total execution time based on the first dynamic compensation time, and to update the second total execution time based on the second dynamic compensation time; the first dynamic compensation time and the second dynamic compensation time are used to compensate for the execution time of the tasks executed on the hardware.

[0138] The determination module 1203 is used to determine the hardware utilization rate within a specified time period by using the updated first total execution time, the updated second total execution time, and the specified time period; the specified time period is the time difference between the first specified time and the second specified time.

[0139] In some alternative implementations, the statistics module 1201 includes:

[0140] The first update unit is used to increase the first total execution time and the first dynamic compensation time by one sampling period each time a sampling period has elapsed and the task is being executed, based on the fact that the execution time of the task executed on the hardware before the first specified time is greater than the sampling period.

[0141] The second update unit is used to increase the second total execution time and the second dynamic compensation time by one sampling period each time a sampling period has elapsed and the task is being executed, based on the fact that the execution time of the task executed on the hardware before the second specified time is greater than the sampling period.

[0142] In some optional implementations, the statistics module 1201 further includes:

[0143] The third update unit is used to update the first total execution time based on the execution time of the task executed on the hardware before the first specified time, if the execution time is less than or equal to the sampling period. The updated first total execution time is equal to the sum of the first total execution time before the update and the execution time of the task.

[0144] The fourth update unit is used to update the second total execution time based on the execution time of the task executed on the hardware before the second specified time, if the execution time is less than or equal to the sampling period. The updated second total execution time is equal to the sum of the previous second total execution time and the task execution time.

[0145] In some alternative implementations, the update module 1202 includes:

[0146] The fifth update unit is used to update the first total execution time based on the first difference between the execution time of the task and the first dynamic compensation time when each task executed on the hardware is completed, if the execution time of the task is greater than the first dynamic compensation time, and the updated first total execution time is equal to the sum of the first total execution time before the update and the first difference; and to reset the first dynamic compensation time to a first preset value, the first preset value being the initial value of the first dynamic compensation time.

[0147] In some alternative implementations, the update module 1202 further includes:

[0148] The sixth update unit is used to update the first dynamic compensation time by using the second difference between the first dynamic compensation time and the task execution time when each task executed on the hardware is completed, if the execution time of the task is less than or equal to the first dynamic compensation time.

[0149] In some alternative implementations, the update module 1202 further includes:

[0150] The seventh update unit is used to update the second total execution time according to the third difference between the execution time of the task and the second dynamic compensation time when each task executed on the hardware is completed, if the execution time of the task is greater than the second dynamic compensation time, and the updated second total execution time is equal to the sum of the previous second total execution time and the third difference; and to reset the second dynamic compensation time to the second preset value, the second preset value being the initial value of the second dynamic compensation time.

[0151] In some alternative implementations, the update module 1202 further includes:

[0152] The eighth update unit is used to update the second dynamic compensation time by using the fourth difference between the second dynamic compensation time and the task execution time when each task executed on the hardware is completed, if the execution time of the task is less than or equal to the second dynamic compensation time.

[0153] In some alternative implementations, the task is a task in a task queue. The determining module 1203 includes:

[0154] The first determining unit is used to determine the fifth difference between the updated first total execution time and the updated second total execution time.

[0155] The second determining unit is used to determine the first ratio between the fifth difference and the specified time period, and to determine the first ratio as the first utilization rate of the hardware by the tasks in the task queue during the specified time period; the hardware utilization rate includes the first utilization rate.

[0156] In some alternative implementations, the apparatus for determining hardware utilization further includes a first computing module and a first calibration module.

[0157] The first calculation module is used to determine a second ratio of the sum of the hardware utilization rates of all tasks in the task queues executed on the hardware within a specified time period to the overall hardware utilization rate.

[0158] The first calibration module is used to calibrate the second utilization rate of the hardware by the target user's tasks using a second ratio. The second utilization rate is the sum of the hardware utilization rates of the tasks in the task queue corresponding to the target user within a specified time period. The calibrated second utilization rate is the product of the second utilization rate before calibration and the second ratio. The hardware utilization rate includes the calibrated second utilization rate.

[0159] In some alternative implementations, the apparatus for determining hardware utilization further includes a second computing module and a second calibration module.

[0160] The second calculation module is used to determine a third ratio of the number of threads used by the task to the total number of threads configured on the hardware during the process of updating the first total execution time and the second total execution time.

[0161] The second calibration module is used to calibrate the execution time and / or sampling period of the task using the third ratio. The execution time of the calibrated task is the product of the execution time of the task before calibration and the third ratio, and the sampling period after calibration is the product of the sampling period before calibration and the third ratio.

[0162] The further functional descriptions of each module and unit are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0163] In this embodiment, the device for determining hardware utilization is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0164] Embodiments of the present invention can also provide a computer device having the above-described features. Figure 12 The device shown is for determining hardware utilization.

[0165] Please see Figure 13 , Figure 13 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 13As shown, the computer device includes one or more processors 1301, memory 1302, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 13 Take the 1301 processor as an example.

[0166] Processor 1301 may be a central processing unit, a network processor, or a combination thereof. Processor 1301 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0167] The memory 1302 stores instructions executable by at least one processor 1301 to cause the at least one processor 1301 to perform the method shown in the above embodiments.

[0168] The memory 1302 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 1302 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 1302 may optionally include memory remotely located relative to the processor 1301, and these remote memories can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0169] The memory 1302 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 1302 may also include a combination of the above types of memory.

[0170] The computer device also includes a communication interface 1303 for communicating with other devices or communication networks.

[0171] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0172] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0173] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for determining hardware utilization, characterized in that, The method includes: The first total execution time and the second total execution time of the tasks executed on the hardware are calculated; wherein, the first total execution time is the sum of the execution times of the tasks executed on the hardware before a first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before a second specified time, wherein the second specified time occurs before the first specified time; The first total execution time is updated based on the first dynamic compensation time, and the second total execution time is updated based on the second dynamic compensation time; the first dynamic compensation time and the second dynamic compensation time are used to compensate for the execution time of the tasks executed on the hardware. The hardware utilization rate within the specified time period is determined by using the updated first total execution time, the updated second total execution time, and the specified time period; the specified time period is the time difference between the first specified time and the second specified time.

2. The method according to claim 1, characterized in that, The first total execution time and the second total execution time of the tasks executed on the statistical hardware include: If the execution time of the task executed on the hardware before the first specified time is greater than the sampling period, then every time the sampling period has elapsed and the task is being executed, the first total execution time and the first dynamic compensation time are both increased by one sampling period. If the execution time of the task executed on the hardware before the second specified time is greater than the sampling period, then every time the sampling period has elapsed and the task is being executed, the second total execution time and the second dynamic compensation time are both increased by one sampling period.

3. The method according to claim 2, characterized in that, The first total execution time and the second total execution time of the tasks executed on the statistical hardware also include: If the execution time of a task executed on the hardware before the first specified time is less than or equal to the sampling period, then the first total execution time is updated according to the execution time of the task. The updated first total execution time is equal to the sum of the first total execution time before the update and the execution time of the task. If the execution time of a task executed on the hardware before the second specified time is less than or equal to the sampling period, the second total execution time is updated according to the execution time of the task. The updated second total execution time is equal to the sum of the previous second total execution time and the execution time of the task.

4. The method according to any one of claims 1 to 3, characterized in that, The step of updating the first total execution time based on the first dynamic compensation time includes: When each task executed on the hardware is completed, if the execution time of the task is greater than the first dynamic compensation time, the first total execution time is updated according to the first difference between the execution time of the task and the first dynamic compensation time. The updated first total execution time is equal to the sum of the first total execution time before the update and the first difference. The first dynamic compensation time is reset to a first preset value, where the first preset value is the initial value of the first dynamic compensation time.

5. The method according to claim 4, characterized in that, The method further includes: When each task executed on the hardware is completed, if the execution time of the task is less than or equal to the first dynamic compensation time, the first dynamic compensation time is updated using the second difference between the first dynamic compensation time and the execution time of the task.

6. The method according to any one of claims 1 to 3, characterized in that, The step of updating the second total execution time based on the second dynamic compensation time includes: When each task executed on the hardware is completed, if the execution time of the task is greater than the second dynamic compensation time, the second total execution time is updated according to the third difference between the execution time of the task and the second dynamic compensation time. The updated second total execution time is equal to the sum of the previous second total execution time and the third difference. The second dynamic compensation time is reset to the second preset value, which is the initial value of the second dynamic compensation time.

7. The method according to claim 6, characterized in that, The method further includes: When each task executed on the hardware is completed, if the execution time of the task is less than or equal to the second dynamic compensation time, the second dynamic compensation time is updated using the fourth difference between the second dynamic compensation time and the execution time of the task.

8. The method according to any one of claims 1 to 3, characterized in that, The task refers to a task in a task queue. Determining the hardware utilization within the specified time period using the updated first total execution time, the updated second total execution time, and the specified time period includes: Determine the fifth difference between the updated first total execution time and the updated second total execution time; Determine the first ratio between the fifth difference and the specified time period, and determine the first ratio as the first utilization rate of the hardware by the task in the task queue during the specified time period; the hardware utilization rate includes the first utilization rate.

9. The method according to claim 8, characterized in that, The method further includes: Determine a second ratio of the sum of the utilization rates of the hardware by the tasks in all task queues executed on the hardware within the specified time period to the overall hardware utilization rate; The second ratio is used to calibrate the second utilization rate of the hardware for the target user's tasks. The second utilization rate is the sum of the utilization rates of the hardware for the tasks in the task queue corresponding to the target user within the specified time period. The calibrated second utilization rate is the product of the second utilization rate before calibration and the second ratio. The hardware utilization rate includes the calibrated second utilization rate.

10. The method according to claim 2 or 3, characterized in that, The method further includes: During the process of updating the first total execution time and the second total execution time, a third ratio is determined between the number of threads used by the task and the total number of threads configured on the hardware; The execution time of the task and / or the sampling period are calibrated using the third ratio. The calibrated execution time of the task is the product of the execution time of the task before calibration and the third ratio, and the calibrated sampling period is the product of the sampling period before calibration and the third ratio.

11. A device for determining hardware utilization, characterized in that, The device includes: The statistics module is used to calculate the first total execution time and the second total execution time of the tasks executed on the hardware; wherein, the first total execution time is the sum of the execution times of the tasks executed on the hardware before a first specified time, and the second total execution time is the sum of the execution times of the tasks executed on the hardware before a second specified time, wherein the second specified time occurs before the first specified time. An update module is used to update the first total execution time based on a first dynamic compensation time, and to update the second total execution time based on a second dynamic compensation time; the first dynamic compensation time and the second dynamic compensation time are used to compensate for the execution time of the tasks executed on the hardware; The determination module is used to determine the hardware utilization rate within the specified time period by using the updated first total execution time, the updated second total execution time, and the specified time period; the specified time period is the time difference between the first specified time and the second specified time.

12. A computer device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method for determining hardware utilization as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the method for determining hardware utilization as described in any one of claims 1 to 10.

14. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the method for determining hardware utilization as described in any one of claims 1 to 10.

Citation Information

Cited By

  • HLS-based two-dimensional array super-resolution direction-finding IP core design method

    CN121072414A