Resource processing methods, apparatus, equipment and storage media

CN114331804BActive Publication Date: 2026-08-14BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-23
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

相关技术中,随着多进程服务功能的引入,GPU资源的统计正确率降低

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114331804B_ABST
    Figure CN114331804B_ABST
Patent Text Reader

Abstract

This disclosure provides a resource processing method, apparatus, device, and storage medium, relating to the field of computer technology, and particularly to the field of cloud computing virtualization resource management. The specific implementation scheme is as follows: based on the M types of graphics processing unit (GPU) resources occupied by each process in M ​​shared memory locations of a graphics card, M first variable records and N second variable records are obtained; wherein, the different first variable records among the M first variable records represent the total usage of the M types of GPU resources by at least one process in different shared memory locations; the different second variable records among the N second variable records represent the usage of corresponding GPU resources by different processes in different shared memory locations; based on the M first variable records and N second variable records, the GPU resource usage value of the graphics card is obtained. According to the technical solution of this disclosure, the statistical accuracy of GPU resources can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more particularly to the field of cloud computing virtualization resource management technology. Background Technology

[0002] With the development of Graphics Processing Unit (GPU) virtualization technology, the statistical analysis of GPU resources has become increasingly important. However, with the introduction of multi-process service functionality, the accuracy of GPU resource statistics has decreased. Summary of the Invention

[0003] This disclosure provides a resource processing method, apparatus, device, storage medium, and computer program product.

[0004] According to a first aspect of this disclosure, a resource processing method is provided, comprising:

[0005] Based on the M shared memory locations of the graphics card, each process occupies M types of GPU resources, resulting in M ​​first variable records and N second variable records. Among the M first variable records, the different first variable records represent the total usage of M types of GPU resources by at least one process in different shared memory locations. The different second variable records in the N second variable records represent the usage of corresponding GPU resources by different processes in different shared memory locations. M is an integer greater than 1, and N is an integer greater than 1.

[0006] Based on M first variable records and N second variable records, the GPU resource usage value of the graphics card is obtained; where the GPU resource usage value includes the total usage of M types of GPU resources by all processes, and the usage of each process for each type of GPU resource.

[0007] According to a second aspect of this disclosure, a resource processing apparatus is provided, comprising:

[0008] The first recording unit is used to record the usage of M types of graphics processing unit (GPU) resources by each process in M ​​shared memory locations based on the graphics card, resulting in M ​​first variable records and N second variable records. Among the M first variable records, the different first variable records represent the total usage of M types of GPU resources by at least one process in different shared memory locations. Among the N second variable records, the different second variable records represent the usage of corresponding GPU resources by different processes in different shared memory locations. M is an integer greater than 1, and N is an integer greater than 1.

[0009] The statistics unit is used to obtain the GPU resource usage value of the graphics card based on M first variable records and N second variable records. The GPU resource usage value includes the total usage of M types of GPU resources by all processes, and the usage of each process for each type of GPU resource.

[0010] According to a third aspect of this disclosure, an electronic device is provided, comprising:

[0011] At least one processor; and

[0012] The memory is communicatively connected to the at least one processor; wherein,

[0013] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described above.

[0014] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing a computer to perform the methods described above.

[0015] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the methods described above.

[0016] According to the technical solution disclosed herein, the statistical accuracy of GPU resources can be improved.

[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0018] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0019] Figure 1 This is a schematic diagram of the implementation flow of the resource processing method according to the embodiments of this disclosure. Figure 1 ;

[0020] Figure 2 This is a schematic diagram of the implementation flow of the resource processing method according to the embodiments of this disclosure. Figure 2 ;

[0021] Figure 3 This is a schematic diagram illustrating the relationship between shared memory, GPU resources, and semaphores according to embodiments of this disclosure;

[0022] Figure 4 This is a schematic diagram of the implementation flow of the resource processing method according to the embodiments of this disclosure. Figure 3 ;

[0023] Figure 5 This is a schematic diagram of the implementation flow of the resource processing method according to the embodiments of this disclosure. Figure 4 ;

[0024] Figure 6 This is a schematic diagram of the structure of a resource processing apparatus according to an embodiment of the present disclosure. Figure 1 ;

[0025] Figure 7 This is a schematic diagram of the structure of a resource processing apparatus according to an embodiment of the present disclosure. Figure 2 ;

[0026] Figure 8 This is a block diagram of an electronic device used to implement the resource processing method of the embodiments of this disclosure. Detailed Implementation

[0027] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0028] The terms "first," "second," and "third," etc., used in the embodiments, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0029] This disclosure provides a resource processing method that can be applied to electronic devices, including but not limited to fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or regular servers. Mobile devices include, but are not limited to, one or more of mobile phones or tablets. Figure 1 As shown, the resource processing method includes:

[0030] S101: Based on the M shared memory locations of the graphics card, each process occupies M types of GPU resources, resulting in M ​​first variable records and N second variable records; where the different first variable records among the M first variable records represent the total usage of M types of GPU resources by at least one process in different shared memory locations; the different second variable records among the N second variable records represent the usage of corresponding GPU resources by different processes in different shared memory locations, where M is an integer greater than 1 and N is an integer greater than 1;

[0031] S102: Based on M first variable records and N second variable records, obtain the GPU resource usage value of the graphics card; where the GPU resource usage value includes the total usage of M types of GPU resources by all processes, and the usage of each process for each type of GPU resource.

[0032] The aforementioned graphics card is one capable of providing GPU resources. This graphics card can be a standalone unit or part of a multi-GPU system. This disclosure does not limit the type of graphics card.

[0033] The types of GPU resources are determined by the functions of the graphics card, and different graphics cards may have different numbers of GPU resource types.

[0034] Generally speaking, GPU resources can be divided into at least two types: video memory resources and computing power resources.

[0035] Here, the value of M is determined by the number of types of GPU resources corresponding to the graphics card. For example, if the graphics card has 2 types of GPU resources, then M = 2. Or, if the graphics card has 3 types of GPU resources, then M = 3.

[0036] Here, the value of N is determined by the number of processes using the GPU resources recorded in the shared memory. The value of N can be different for different shared memory types. For example, if the GPU resource corresponding to shared memory 1 is denoted as resource a, and the number of processes using resource a is n1, then N = n1; if the GPU resource corresponding to shared memory 2 is denoted as resource b, and the number of processes using resource b is n2, then N = n2, and n1 ≠ n2.

[0037] In some implementations, based on the M types of GPU resources used by each process within the M shared memory locations of the graphics card, M first variable records and N second variable records are obtained. This includes: recording the usage of each process's j-th type of GPU resource in the second variable record of the j-th shared memory location; obtaining the total usage of all processes' j-th type of GPU resource based on each process's usage; and recording the total usage of all processes' j-th type of GPU resource in the first variable record of the j-th shared memory location, where j is an integer greater than or equal to 1 and less than or equal to M. This facilitates rapid subsequent statistical analysis of GPU resource values.

[0038] In other implementations, based on the M types of GPU resources occupied by each process in the M shared memory locations of the graphics card, M first variable records and N second variable records are obtained. This includes: obtaining the total usage of the current j-th type of GPU resource and the usage of each process; recording the total usage of the j-th type of GPU resource in the first variable record of the j-th shared memory location; and recording the usage of the j-th type of GPU resource by each process in the second variable record of the j-th shared memory location, where j is an integer greater than or equal to 1 and less than or equal to M. This facilitates rapid subsequent statistical analysis of GPU resource values.

[0039] Here, the GPU resource usage values ​​of the graphics card include at least: the total usage of each type of GPU resource in the graphics card by all processes; and the usage of each type of GPU resource in the graphics card by each process.

[0040] In some implementations, the GPU resource usage value of the graphics card is obtained based on M first variable records and N second variable records, including: obtaining the total usage of M types of resources from the M first variable records; and calculating the usage of each process of the M types of GPU resources from the N second variables. In this way, the GPU resource usage value of the graphics card can be directly calculated from the first and second variable records, improving the efficiency of calculating the GPU resource usage value of the graphics card.

[0041] In other implementations, the GPU resource usage value of the graphics card is obtained based on M first variable records and N second variable records. This includes: calculating the usage of M types of GPU resources by each process from the N second variables; obtaining the total usage of the M types of GPU resources based on the usage of each process; comparing the total usage of the M types of GPU resources obtained from the usage of each process with the total usage of the M types of resources obtained from the M first variable records, and determining the total usage of the M types of resources based on the comparison result. For example, if the comparison result shows that the total usage of the M types of GPU resources obtained from the usage of each process is the same as the total usage of the M types of resources obtained from the M first variable records, then the total usage of the M types of resources will be obtained from the M first variable records and used as the final statistical result regarding the total usage of the M types of resources. For example, if the comparison result shows that the total usage of M types of GPU resources obtained from each process's usage of M types of GPU resources is different from the total usage of M types of resources obtained from the M first variable records, then new M first variable records and new N second variable records are obtained. The statistical steps are then repeated based on these new M first variable records and N second variable records until the total usage of M types of GPU resources obtained from each process's usage of M types of GPU resources matches the total usage of M types of resources obtained from the M first variable records. In this way, the second variable records can be fully utilized to verify the values ​​of the first variable records, thereby further ensuring the accuracy of the statistically obtained GPU resource usage values ​​for the graphics card.

[0042] The technical solution described in this disclosure obtains M first variable records and N second variable records by analyzing the M types of GPU resources occupied by each process in the M shared memory locations of the graphics card. Based on the M first variable records and N second variable records, the GPU resource usage value of the graphics card is statistically obtained. In this way, since the usage of each type of GPU resource by each process is recorded, the statistical results can be accurately obtained when statistically analyzing the GPU resource usage value of the graphics card, thus improving the accuracy of GPU resource statistics.

[0043] This disclosure provides a resource processing method that can be applied to electronic devices, including but not limited to fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or regular servers. Mobile devices include, but are not limited to, one or more of mobile phones or tablets. Figure 2 As shown, the resource processing method includes:

[0044] S201: Determine the M shared memory locations corresponding to the M types of GPU resources of the graphics card, with a one-to-one correspondence between the M types of GPU resources and the M shared memory locations;

[0045] S202: Determine the M semaphores corresponding to the M shared memory locations, with a one-to-one correspondence between the M shared memory locations and the M semaphores;

[0046] S203: Based on the situation of each process occupying M types of GPU resources in M ​​shared memory locations of the graphics card, obtain M first variable records and N second variable records; among them, the different first variable records in the M first variable records represent the total amount of M types of GPU resources used by at least one process in different shared memory locations; the different second variable records in the N second variable records represent the amount of GPU resources used by different processes in different shared memory locations.

[0047] S204: Based on M first variable records and N second variable records, obtain the GPU resource usage value of the graphics card; where the GPU resource usage value includes the total usage of M types of GPU resources by all processes, and the usage of each process for each type of GPU resource.

[0048] For each shared memory segment, a semaphore is maintained for each GPU resource. When a process needs to change the usage value of a certain GPU resource in the shared memory, it must first acquire the semaphore for that resource in the shared memory. Because there is only one semaphore for each shared memory segment on each resource, only one process can change the usage value of a certain GPU resource in that shared memory segment at a time.

[0049] In some implementations, determining the M shared memory locations corresponding to the M types of GPU resources on the graphics card includes: allocating M shared memory locations based on the M types of GPU resources on the graphics card, and establishing a one-to-one correspondence between the M types of GPU resources and the M shared memory locations. This allows for the allocation of M shared memory locations for the M types of GPU resources, resulting in a one-to-one correspondence between the M types of GPU resources and the M shared memory locations. This provides fundamental support for subsequently obtaining M first variable records and N second variable records regarding the M types of GPU resources occupied by each process within the M shared memory locations of the graphics card.

[0050] In other implementations, determining the M shared memory locations corresponding to the M types of GPU resources on the graphics card includes: selecting M shared memory locations from a pre-configured set of S shared memory locations; and, based on the size of the shared memory locations and the type of GPU resource, mapping the M shared memory locations to the M types of GPU resources one-to-one, where S is an integer greater than or equal to M, and at least one of the S shared memory locations has a different size than the others. This allows for the allocation of M shared memory locations appropriate to the M types of GPU resources, resulting in a one-to-one correspondence between the M types of GPU resources and the M shared memory locations. This provides a foundation for subsequently obtaining M first variable records and N second variable records regarding the M types of GPU resources occupied by each process within the M shared memory locations of the graphics card.

[0051] In some implementations, determining the M semaphores corresponding to the M shared memory locations includes: allocating M semaphores based on the M shared memory locations, thus establishing a one-to-one correspondence between the M shared memory locations and the M semaphores. This allows for the allocation of M semaphores to the M shared memory locations, resulting in a one-to-one correspondence between the M shared memory locations and the M semaphores. This provides ordered support for obtaining M first variable records and N second variable records for the M types of GPU resources occupied by each process within the M shared memory locations based on the graphics card.

[0052] In other implementations, determining the M semaphores corresponding to the M shared memory locations includes: based on the correspondence between M types of GPU resources and M semaphores, allocating M semaphores to the M shared memory locations corresponding to the M types of GPU resources, thus establishing a one-to-one correspondence between the M shared memory locations and the M semaphores. This allows for the allocation of M semaphores appropriate to the GPU resources stored in the M shared memory locations, resulting in a one-to-one correspondence between the M shared memory locations and the M semaphores. This provides ordering support for obtaining M first variable records and N second variable records for the M types of GPU resources occupied by each process within the M shared memory locations based on the graphics card.

[0053] In this embodiment of the disclosure, the relationship between the GPU resources of the graphics card, shared memory, and semaphores is illustrated as follows: Figure 3 As shown, from Figure 3 As can be seen, graphics card 1 has M types of GPU resources. Each type of GPU resource is allocated a shared memory, and each shared memory is allocated a semaphore. The M types of GPU resources correspond one-to-one with the M shared memory, and the M shared memory corresponds one-to-one with the M semaphores.

[0054] It should be understood that the above relationship diagram is merely exemplary and not restrictive, and it is scalable, which can include more graphics cards, and for each GPU resource of each graphics card, more shared memory and semaphores can be used to maintain, thereby enabling the simultaneous use of more graphics cards and more GPU resources to handle more process tasks.

[0055] For example, graphics card 1 has two types of GPU resources, denoted as resource a and resource b. Two shared memory locations are allocated to graphics card 1, denoted as shared memory 1 and shared memory 2. Shared memory 1 stores the usage information of resource a, and shared memory 2 stores the usage information of resource b. Furthermore, a semaphore, denoted as semaphore a', is allocated to shared memory 1, and a semaphore, denoted as semaphore b', is allocated to shared memory 2. For processes accessing shared memory 1, the process that acquires semaphore a' can prioritize recording the usage of resource a. For processes accessing shared memory 2, the process that acquires semaphore b' can prioritize recording the usage of resource b.

[0056] The technical solution described in this disclosure predetermines the M shared memory locations corresponding to the M types of GPU resources on the graphics card, and determines the M semaphores corresponding to the M shared memory locations. This provides fundamental support for the subsequent acquisition of M first variable records and N second variable records, helping to ensure the orderly progress of the recording work and thus improving the accuracy of GPU resource statistics. Furthermore, compared to multiple graphics cards sharing a single shared memory location and semaphore, it reduces synchronization blocking time and improves concurrency efficiency.

[0057] In some implementations, based on the M types of GPU resources used by each process in the M shared memory locations of the graphics card, M first variable records and N second variable records are obtained. This includes: when the i-th process acquires the semaphore of the j-th shared memory location among the M shared memory locations, recording the usage of the j-th GPU resource occupied by the i-th process in the j-th shared memory location, resulting in a first variable record and a second variable record for the j-th shared memory location, where i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1 and less than or equal to M. Thus, by maintaining a separate semaphore for each type of GPU resource, blocking caused by process synchronization during resource updates is reduced, improving concurrency efficiency.

[0058] For example, graphics card 1 includes two shared memory modules, denoted as shared memory 1 and shared memory 2. Shared memory 1 is used to record video memory resources, and its associated semaphore is denoted as semaphore 1. Shared memory 2 is used to record computing resources, and its associated semaphore is denoted as semaphore 2. Assume there are two processes, denoted as process 1 and process 2. When process 1 and process 2 access shared memory 1, they should first acquire semaphore 1. When process 1 and process 2 access shared memory 2, they should first acquire semaphore 2. If process 1 acquires semaphore 1 when accessing shared memory 1, process 2 cannot access shared memory 1 at that time. If process 1 acquires semaphore 1 and requests 100 MB of video memory, in shared memory 1, the first variable records x1 = 100 MB, where x1 represents the total video memory used; the second variable records y11 = 100 MB, where y11 represents the video memory used by process 1. If process 1 releases semaphore 1 and process 2 acquires semaphore 1 and requests 50 MB of video memory, in shared memory 1, the first variable records x1 = 150 MB, where x1 represents the total video memory used; the second variable records y11 = 100 MB, where y11 represents the video memory used by process 1; and y12 = 50 MB, where y12 represents the video memory used by process 2.

[0059] In other implementations, based on the M types of GPU resources used by each process in the M shared memory locations of the graphics card, M first variable records and N second variable records are obtained. This includes recording the usage of the j-th GPU resource by the i-th process in the j-th shared memory location according to the time sequence of the i-th process's access to the j-th shared memory location, resulting in a first variable record and a second variable record for the j-th shared memory location. Here, i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1 and less than or equal to M. This helps to sequentially write the information on the M types of GPU resources used by each process into the first and second variable records, thereby helping to prevent omissions.

[0060] The technical solution described in this disclosure uses shared memory and semaphores to record the GPU resource usage of each process in the graphics card, ensuring the accuracy of the GPU resource usage records for each process. This lays a correct data foundation for subsequent statistics on GPU resource usage, thereby improving the accuracy of GPU resource statistics.

[0061] This disclosure provides a resource processing method that can be applied to electronic devices, including but not limited to fixed devices and / or mobile devices. For example, fixed devices include, but are not limited to, servers, which can be cloud servers or regular servers. Mobile devices include, but are not limited to, one or more of mobile phones or tablets. Figure 4 As shown, the resource processing method includes:

[0062] S401: Based on the M types of GPU resources occupied by each process, determine the values ​​of M atomic variables in each process; based on the values ​​of the M atomic variables in each process, determine the M types of GPU resources occupied by each process.

[0063] S402: Based on the situation of each process occupying M types of GPU resources in M ​​shared memory locations of the graphics card, obtain M first variable records and N second variable records; among them, the different first variable records in the M first variable records represent the total amount of M types of GPU resources used by at least one process in different shared memory locations; the different second variable records in the N second variable records represent the amount of GPU resources used by different processes in different shared memory locations.

[0064] S403: Based on M first variable records and N second variable records, obtain the GPU resource usage value of the graphics card; where the GPU resource usage value includes the total usage of M types of GPU resources by all processes, and the usage of each process for each type of GPU resource.

[0065] Atomic variables are used to record the resource usage of each process. Each process maintains multiple atomic variables to record the resource usage of that process on different graphics cards.

[0066] It should be noted that the total number of atomic variables can differ across processes. The number of atomic variables for each process is equal to the total number of GPU resources currently being used by that process across all the graphics cards. For example, if the first process uses the GPU resources of two graphics cards, with two different types of GPU resources on the first card and three different types on the second card, then the first process has five atomic variables. If the first process uses the GPU resources of three graphics cards, with two different types of GPU resources on the first card and three different types on the second card, then the second process has nine atomic variables.

[0067] For example, multiple processes can run within each container, such as process A and process B. Each process can contain multiple threads, such as thread a1 and thread a2 in process A, and thread b1 and thread b2 in process B. Each thread can utilize GPU resources. GPU resources include video memory, computing power, and other resources, and these resources are distributed across different graphics cards, such as graphics card 1 and graphics card 2. For instance, process A maintains atomic variables A_1_1, A_1_2, A_2_1, and A_2_2, representing the video memory usage, computing power usage, video memory usage, and computing power usage of process A on graphics card 1 and graphics card 2, respectively.

[0068] In some implementations, the values ​​of M atomic variables in each process are determined based on the M types of GPU resources occupied by each process, including: determining the values ​​of M atomic variables in each process according to the priority order of the M types of GPU resources.

[0069] In other implementations, based on the M types of GPU resources occupied by each process, the values ​​of M atomic variables in each process are determined, including: determining the values ​​of M atomic variables in each process according to the time order in which each process occupies each type of GPU resource.

[0070] The technical solution described in this disclosure, by using atomic variables to back up the usage of each process for each type of GPU resource, can provide a data foundation for the first variable record and the second variable record in shared memory, ensuring the accuracy of the statistical usage of each process for each type of GPU resource, thereby further ensuring the correctness of the statistics on GPU resource usage.

[0071] based on Figure 4 In some embodiments of the technical solution shown, the resource processing method may further include:

[0072] In response to the update request from the i-th process for the GPU resource corresponding to the j-th shared memory, find the atomic variable corresponding to the GPU resource corresponding to the j-th shared memory in the i-th process; i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1 and less than or equal to M;

[0073] Based on the usage of GPU resources corresponding to the j-th shared memory requested in the update request, the value of the atomic variable corresponding to the j-th shared memory is updated.

[0074] Here, update requests apply to scenarios where the GPU resources used by a process change, including but not limited to requests to request GPU resources and requests to release GPU resources.

[0075] For example, when thread a1 of process A requests GPU resources, it will find the atomic variable in process A that records the type of resource based on the graphics card and resource type corresponding to the requested GPU resources. For example, A_1_2 represents the computing power resource usage of process A on graphics card 1. The value of the atomic variable A_1_2 will be changed, specifically by adding the requested computing power resource usage to the original value.

[0076] For example, when thread b2 of process B releases GPU resources, it will find the atomic variable in process B that records the type of resource based on the graphics card and resource type corresponding to the released GPU resources. For example, B_1_1 represents the amount of video memory used by process B on graphics card 1. The value of the atomic variable B_1_1 will be changed, specifically by subtracting the amount of video memory resources released from the original value.

[0077] It should be noted that atomic variables ensure that only one thread in a process is updating the value of an atomic variable at a time. When one thread changes the value of an atomic variable, other threads in the same process that are also requesting or releasing similar GPU resources will not be able to operate on the value of that atomic variable and will have to wait for the previous thread to complete updating the value of the atomic variable.

[0078] The technical solution described in this disclosure, by orderly changing the values ​​of atomic variables in a process, can ensure the accuracy of the usage of each GPU resource by each process recorded by the atomic variables, thereby further ensuring the correctness of the statistical GPU resource usage.

[0079] based on Figure 4 The resource processing method shown in the technical solution may further include: adjusting the first variable record and the second variable record in the j-th shared memory based on the updated value of the atomic variable corresponding to the GPU resource corresponding to the j-th shared memory in the i-th process.

[0080] In some implementations, adjusting the first and second variable records in the j-th shared memory based on the updated values ​​of the atomic variables corresponding to the GPU resources corresponding to the j-th shared memory in the i-th process includes: updating the total usage of the j-th type of GPU resources by all processes in the first variable record in the j-th shared memory based on the updated values ​​of the atomic variables corresponding to the GPU resources corresponding to the j-th shared memory in the i-th process; and updating the usage of the j-th type of GPU resources by the i-th process in the second variable record in the j-th shared memory based on the updated values ​​of the atomic variables corresponding to the GPU resources corresponding to the j-th shared memory in the i-th process.

[0081] For example, graphics card 1 includes two shared memory modules, denoted as Shared Memory 1 and Shared Memory 2. Shared Memory 1 records video memory resources, and its associated semaphore is denoted as Semaphore 1. Shared Memory 2 records computing resources, and its associated semaphore is denoted as Semaphore 2. Assume there are two processes, denoted as Process 1 and Process 2. In Shared Memory 1, the first variable records x1 = 150 MB, where x1 represents the total video memory usage; the second variable records y11 = 100 MB, where y11 represents the video memory usage of Process 1; and y12 = 50 MB, where y12 represents the video memory usage of Process 2. When Process 1 and Process 2 access Shared Memory 1, they must first acquire semaphore 1. If Process 1 acquires semaphore 1 when accessing Shared Memory 1, Process 2 cannot access Shared Memory 1 at that time. In shared memory 2, the first variable record includes x2 = 200 MB, where x2 represents the total computing resource usage; the second variable record includes y21 = 50 MB, where y21 represents the computing resource usage of process 1; and y22 = 150 MB, where y22 represents the computing resource usage of process 2. Both process 1 and process 2 must acquire semaphore 2 before accessing shared memory 2. If process 2 acquires semaphore 2 when accessing shared memory 2, process 1 cannot access shared memory 2 at this time. If process 1 requests to release 100 MB of video memory resources, then the value of the atomic variable used to record video memory resources in process 1 is updated from 100 MB to 0, which in turn updates x1 in the first variable record of shared memory 1 from 150 MB to 50 MB, and updates y11 in the second variable record of shared memory 1 from 100 MB to 0.

[0082] In other implementations, the first and second variable records in the j-th shared memory are adjusted based on the updated values ​​of the atomic variables corresponding to the GPU resources corresponding to the j-th shared memory in the i-th process. This includes: updating the usage of the j-th type of GPU resources by the i-th process in the second variable record in the j-th shared memory based on the updated values ​​of the atomic variables corresponding to the GPU resources corresponding to the j-th shared memory in the i-th process; and updating the total usage of the j-th type of GPU resources by all processes in the first variable record in the j-th shared memory based on the updated usage of the j-th type of GPU resources by the i-th process in the second variable record.

[0083] It should be noted that this disclosure does not limit the order of adjustment of the first variable record and the second variable record. They can be updated simultaneously, or the first variable record can be updated first, or the second variable record can be updated first.

[0084] The technical solution described in this disclosure uses atomic variables in processes, shared memory and semaphores in graphics cards, and allocates independent shared memory and semaphores for each type of GPU resource. After the multi-process service is started, the usage of GPU resources can be accurately counted. Compared with multiple graphics cards sharing a single shared memory and semaphore, the blocking time of synchronization is reduced and the concurrency efficiency is improved.

[0085] When a process exits normally, its resource usage is actively deducted from the corresponding shared memory. If the process terminates abnormally, although it is no longer using GPU resources, the resource usage value in the corresponding shared memory may not have been updated in time due to the abnormal termination, leading to errors in the GPU resource usage statistics. To avoid this error, embodiments of this disclosure provide a resource processing method, such as... Figure 5 As shown, the resource processing method includes:

[0086] S501: Based on the situation of each process occupying M types of GPU resources in M ​​shared memory locations of the graphics card, we obtain M first variable records and N second variable records; among them, the different first variable records in the M first variable records represent the total amount of M types of GPU resources used by at least one process in different shared memory locations; the different second variable records in the N second variable records represent the amount of GPU resources used by different processes in different shared memory locations.

[0087] S502: In the case of a process that terminates abnormally in M ​​shared memory locations, determine the usage of M types of GPU resources corresponding to the abnormally terminated process, remove the usage of the M types of GPU resources corresponding to the abnormally terminated process from the M shared memory locations, and obtain M new first variable records and N new second variable records.

[0088] S503: Based on the new M first variable records and the new N second variable records, obtain the GPU resource usage value of the graphics card; where the GPU resource usage value includes the total usage of M types of GPU resources by all processes, and the usage of each process for each type of GPU resource.

[0089] Here, abnormal termination is relative to normal exit. Therefore, a process that exits abnormally can be understood as a process that terminates abnormally.

[0090] It should be noted that S502 occurs during the entire resource statistics process, but can also be executed when the inspection conditions are met; this disclosure does not impose any restrictions on this.

[0091] In one specific implementation, the check condition is triggered in real time. In response to triggering the check condition, that is, in real time, it checks whether there are any abnormally terminated processes in the M shared memory. If there are abnormally terminated processes, the amount of resources that the abnormally terminated processes should release is determined, and the amount of resources that the abnormally terminated processes should release is removed from the shared memory.

[0092] In another specific implementation, the check condition is triggered periodically. In response to the triggering of the check condition, that is, when the check period is reached, it is checked whether there are any abnormal terminations of certain processes in the M shared memory. If there are abnormally terminated processes, the amount of resources that the abnormally terminated processes should release is determined, and the amount of resources that the abnormally terminated processes should release is removed from the shared memory.

[0093] In another specific implementation, the check condition is triggered when a new process starts. In response to triggering the check condition, that is, when a new process starts, it is checked whether there are any abnormal processes in the M shared memory. If there are abnormally terminated processes, the amount of resources that the abnormally terminated processes should release is determined, and then the amount of resources that the abnormally terminated processes should release is removed from the shared memory.

[0094] For example, graphics card 1 includes two shared memory modules, denoted as Shared Memory 1 and Shared Memory 2. Shared Memory 1 records video memory resources, and its associated semaphore is denoted as Semaphore 1. Shared Memory 2 records computing resources, and its associated semaphore is denoted as Semaphore 2. Assume there are two processes, denoted as Process 1 and Process 2. In Shared Memory 1, the first variable records x1 = 150 MB, where x1 represents the total video memory usage; the second variable records y11 = 100 MB, where y11 represents the video memory usage of Process 1; and y12 = 50 MB, where y12 represents the video memory usage of Process 2. When Process 1 and Process 2 access Shared Memory 1, they must first acquire semaphore 1. If Process 1 acquires semaphore 1 when accessing Shared Memory 1, Process 2 cannot access Shared Memory 1 at that time. In shared memory 2, the first variable records x2 = 200 MB, where x2 represents the total computing resources used; the second variable records y21 = 50 MB, where y21 represents the computing resources used by process 1; and y22 = 150 MB, where y22 represents the computing resources used by process 2. When process 1 and process 2 access shared memory 2, they must first acquire semaphore 2. If process 2 acquires semaphore 2 when accessing shared memory 2, process 1 cannot access shared memory 2 at that time. If process 3 accesses shared memory 1 for the first time, since shared memory 1 records the usage of video memory resources by processes 1 and 2, it checks whether processes 1 and 2 have terminated abnormally. If neither has terminated abnormally, assuming process 3 requests 30 MB of video memory resources, then in the first variable record, x1 is changed from 150 MB to 180 MB; the second variable record, after the change, includes y11 = 100 MB, where y11 represents the video memory resource usage of process 1; y12 = 50 MB, where y12 represents the video memory resource usage of process 2; and y13 = 30 MB, where y13 represents the video memory resource usage of process 3. If process 1 terminates abnormally, first change x1 in the first variable record from 150 MB to 50 MB; change y11 in the second variable record from 100 MB to 0, where y11 represents the video memory usage of process 1; y12 = 50 MB, where y12 represents the video memory usage of process 2; assuming process 3 requests 30 MB of video memory, then change the first variable record to x1 = 80 MB; change the second variable record to y11 = 0, where y11 represents the video memory usage of process 1; y12 = 50 MB, where y12 represents the video memory usage of process 2; and y13 = 30 MB, where y13 represents the video memory usage of process 3.

[0095] The technical solution described in this disclosure, through a process detection mechanism for abnormally terminated processes, ensures the accuracy of the statistical GPU resource usage even in the event of abnormal process exit.

[0096] The resource processing method provided in this disclosure can be used in projects such as resource statistics or resource allocation. For example, the executing entity of the method can be an electronic device or a server. The resource processing method provided in this disclosure can be applied to multi-process service scenarios. For example, when a business process in a container needs to use virtualization services, it can directly link to a GPU virtualization dynamic library that can be mounted to different container environments. The above resource processing method can be applied to this GPU virtualization dynamic library, which handles the process's resource allocation or release requests, ensuring the management of GPU resource usage in the container.

[0097] This disclosure also provides a resource processing apparatus, such as... Figure 6 As shown, the resource processing device includes:

[0098] The first recording unit 610 is used to record the information of each process occupying M types of graphics processing unit (GPU) resources in M ​​shared memory locations based on the graphics card, obtaining M first variable records and N second variable records; wherein, the different first variable records among the M first variable records represent the total usage of M types of GPU resources occupied by at least one process in different shared memory locations; the different second variable records among the N second variable records represent the usage of corresponding GPU resources occupied by different processes in different shared memory locations, where M is an integer greater than 1 and N is an integer greater than 1;

[0099] The statistics unit 620 is used to obtain the GPU resource usage value of the graphics card based on M first variable records and N second variable records. The GPU resource usage value includes the total usage of M types of GPU resources by all processes and the usage of each process for each type of GPU resource.

[0100] In some implementations, such as Figure 7 As shown, the device may further include:

[0101] The determination unit 630 is used to determine the M shared memory locations corresponding to the M types of GPU resources of the graphics card, with each of the M types of GPU resources corresponding to one of the M shared memory locations; and to determine the M semaphores corresponding to the M shared memory locations, with each of the M shared memory locations corresponding to one of the M semaphores.

[0102] In some implementations, the first recording unit 610 is specifically used to record the usage of the j-th GPU resource occupied by the i-th process in the j-th shared memory when the i-th process obtains the semaphore of the j-th shared memory in the M shared memory, and to obtain the first variable record and the second variable record of the j-th shared memory, where i is an integer greater than or equal to 1 and j is an integer greater than or equal to 1 and less than or equal to M.

[0103] In some implementations, such as Figure 7As shown, the device may further include:

[0104] The second recording unit 640 is used to determine the values ​​of M atomic variables in each process based on the M types of GPU resources occupied by each process.

[0105] In some implementations, such as Figure 7 As shown, the device may further include: a lookup unit 650; wherein the lookup unit 650 is used to, in response to the update request of the i-th process for the GPU resource corresponding to the j-th shared memory, search for the atomic variable corresponding to the GPU resource corresponding to the j-th shared memory in the i-th process; i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1 and less than or equal to M;

[0106] The second recording unit 640 is also used to update the value of the atomic variable corresponding to the GPU resource corresponding to the j-th shared memory based on the usage of the GPU resource corresponding to the j-th shared memory requested by the update request.

[0107] In some implementations, such as Figure 7 As shown, the device may further include:

[0108] The adjustment unit 660 is used to adjust the first variable record and the second variable record in the j-th shared memory based on the updated value of the atomic variable corresponding to the GPU resource corresponding to the j-th shared memory in the i-th process.

[0109] In some implementations, such as Figure 7 As shown, the device may further include:

[0110] The correction unit 670 is used to determine the usage of M types of GPU resources corresponding to the abnormally terminated process when there is an abnormally terminated process in M ​​shared memory, and remove the usage of M types of GPU resources corresponding to the abnormally terminated process from the M shared memory.

[0111] Those skilled in the art should understand that the functions of each processing module in the resource processing apparatus of this disclosure embodiment can be understood with reference to the relevant description of the foregoing resource processing method. Each processing module in the resource processing apparatus of this disclosure embodiment can be implemented by an analog circuit that implements the functions described in the embodiments of this disclosure, or it can be implemented by running software that performs the functions described in the embodiments of this disclosure on an electronic device.

[0112] The resource processing apparatus of this disclosure can improve the statistical accuracy of GPU resources.

[0113] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0114] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0115] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0116] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0117] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0118] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as resource processing methods. For example, in some embodiments, the resource processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the resource processing method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform resource processing methods by any other suitable means (e.g., by means of firmware).

[0119] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0120] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (EPROM), flash memory, optical fiber, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0122] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0123] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0124] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0125] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0126] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A resource processing method, comprising: Based on the M types of graphics processing unit (GPU) resources used by each process, determine the values ​​of M atomic variables for each process; Based on the M types of GPU resources used by each process in M ​​shared memory locations of a graphics card, M first variable records and N second variable records are obtained; wherein, the different first variable records among the M first variable records represent the total amount of M types of GPU resources used by at least one process in different shared memory locations; the different second variable records among the N second variable records represent the amount of GPU resources used by different processes in different shared memory locations, where M is an integer greater than 1 and N is an integer greater than 1; Based on the M first variable records and the N second variable records, the GPU resource usage value of the graphics card is obtained; wherein, the GPU resource usage value includes the total usage of the M types of GPU resources by all processes, and the usage of each process for each type of GPU resource; In response to the update request from the i-th process for the GPU resource corresponding to the j-th shared memory, the atomic variable corresponding to the GPU resource corresponding to the j-th shared memory is searched in the i-th process; i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1 and less than or equal to M; Based on the usage of GPU resources corresponding to the j-th shared memory requested in the update request, the value of the atomic variable corresponding to the GPU resource corresponding to the j-th shared memory is updated. Based on the updated values ​​of the atomic variables corresponding to the GPU resources of the j-th shared memory in the i-th process, adjust the first variable record and the second variable record in the j-th shared memory.

2. The method according to claim 1, further comprising: The M types of GPU resources of the graphics card are identified as corresponding to the M shared memory locations, and the M types of GPU resources correspond one-to-one with the M shared memory locations; Determine the M semaphores corresponding to the M shared memory locations, with each of the M shared memory locations corresponding to one of the M semaphores.

3. The method according to claim 2, wherein, The M types of GPU resources used by each process in the M shared memory locations based on the graphics card are used to obtain M first variable records and N second variable records, including: When the i-th process obtains the semaphore of the j-th shared memory among the M shared memory, the usage of the j-th GPU resource occupied by the i-th process in the j-th shared memory is recorded, and the first variable record and the second variable record of the j-th shared memory are obtained, where i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1 and less than or equal to M.

4. The method according to claim 1, further comprising: If there is an abnormally terminated process in the M shared memory locations, determine the usage of the M types of GPU resources corresponding to the abnormally terminated process, and remove the usage of the M types of GPU resources corresponding to the abnormally terminated process from the M shared memory locations.

5. A resource processing apparatus, comprising: The second recording unit is used to determine the values ​​of M atomic variables for each process based on the M types of graphics processing unit (GPU) resources occupied by each process. The first recording unit is used to record the usage of M types of GPU resources by each process in M ​​shared memory locations based on the graphics card, resulting in M ​​first variable records and N second variable records. The different first variable records among the M first variable records represent the total usage of M types of GPU resources by at least one process in different shared memory locations. The different second variable records among the N second variable records represent the usage of corresponding GPU resources by different processes in different shared memory locations. M is an integer greater than 1, and N is an integer greater than 1. The statistics unit is used to obtain the GPU resource usage value of the graphics card based on the M first variable records and the N second variable records; wherein, the GPU resource usage value includes the total usage of the M types of GPU resources by all processes, and the usage of each process for each type of GPU resource; The lookup unit is used to respond to the update request of the i-th process for the GPU resource corresponding to the j-th shared memory, and to look up the atomic variable corresponding to the GPU resource corresponding to the j-th shared memory in the i-th process; i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1 and less than or equal to M; The second recording unit is further configured to update the value of the atomic variable corresponding to the GPU resource corresponding to the j-th shared memory based on the usage of the GPU resource corresponding to the j-th shared memory requested by the update request; The adjustment unit is used to adjust the first variable record and the second variable record in the j-th shared memory based on the updated values ​​of the atomic variables corresponding to the GPU resources corresponding to the j-th shared memory in the i-th process.

6. The apparatus according to claim 5, further comprising: The determining unit is used to determine the M shared memory locations corresponding to the M types of GPU resources of the graphics card, wherein the M types of GPU resources correspond one-to-one with the M shared memory locations; and to determine the M semaphores corresponding to the M shared memory locations, wherein the M shared memory locations correspond one-to-one with the M semaphores.

7. The apparatus according to claim 6, wherein, The first recording unit is used for: When the i-th process obtains the semaphore of the j-th shared memory among the M shared memory, the usage of the j-th GPU resource occupied by the i-th process in the j-th shared memory is recorded, and the first variable record and the second variable record of the j-th shared memory are obtained, where i is an integer greater than or equal to 1, and j is an integer greater than or equal to 1 and less than or equal to M.

8. The apparatus according to claim 5, further comprising: The correction unit is configured to, in the case of an abnormally terminated process existing in the M shared memory locations, determine the usage of the M types of GPU resources corresponding to the abnormally terminated process, and remove the usage of the M types of GPU resources corresponding to the abnormally terminated process from the M shared memory locations.

9. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4.

10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-4.

11. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • GPU video memory management control method and related device

    CN110930291A