Resource scheduling method and electronic equipment
By adding a PCIe memory management unit to the PCIe protocol layer, direct mapping of video memory and hard disk and parallel transmission of multi-channels is realized, the task interruption caused by insufficient computing resources is solved, data transmission efficiency and task execution effect are improved, and cost is reduced.
Patent Information
- Application Number
- CN202510542654.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-22
AI Technical Summary
Insufficient computing resource capacity leads to task interruption or affects task execution effectiveness, and the existing technical solutions have problems of data transmission bottlenecks and high costs.
By adding a PCIe memory management unit to the PCIe protocol layer, the hardware-level mapping of the video memory physical address and the hard disk logical block address is realized, data is directly transmitted, memory redirection is avoided, multi-channel parallel transmission and enhanced CRC verification are adopted to ensure the accuracy and efficiency of data transmission.
It realizes the expansion of computing resources and the improvement of data transmission efficiency, reduces delay, improves task execution effect and efficiency, and reduces costs.
Smart Images

Figure CN120353600A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data storage, and in particular, to a resource scheduling method and an electronic device. Background Art
[0002] During the execution of different tasks (e.g., model training), the demand for computing resources shows an exponential growth. If the computing resource capacity is insufficient, it is easy to cause task interruption or affect the execution effect of the task. Summary of the Invention
[0003] In view of this, the present disclosure provides a resource scheduling method and an electronic device.
[0004] According to a first aspect of the present disclosure, a resource scheduling method is provided, including: obtaining candidate data of a target task, where the candidate data represents data in a first computing resource that does not meet the requirements for executing the target task; in the case of an address mapping relationship, according to the address mapping relationship, transmitting the candidate data from the first computing resource to a second computing resource, where the address mapping relationship represents the corresponding relationship between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource; in the case of no address mapping relationship, interrupting the data transmission between the first computing resource and the second computing resource.
[0005] According to an embodiment of the present disclosure, the method further includes: determining the virtual address of the candidate data in the second computing resource according to the physical address of the candidate data in the first computing resource, and establishing an address mapping relationship.
[0006] According to an embodiment of the present disclosure, transmitting the candidate data from the first computing resource to the second computing resource according to the address mapping relationship includes: encapsulating the candidate data into multiple data packets, where the data volume of the data packets is the same as the capacity of the read / write unit of the second computing resource; transmitting the multiple data packets to the second computing resource in parallel through multiple channels.
[0007] According to an embodiment of the present disclosure, the method further includes: determining a recombination table according to the multiple data packets, where the recombination table includes the address information of the multiple data packets in the first computing resource; storing the multiple data packets and the recombination table corresponding to the data packets in the second computing resource.
[0008] According to an embodiment of the present disclosure, the data packet includes first payload information, and the method further includes: in response to the second computing resource receiving the data packet, calculating the second payload information of the data packet; in the case where the second payload information is consistent with the first payload information, storing the data packet in the second computing resource; in the case where the second payload information is inconsistent with the first payload information, generating a control signal, where the control signal is used to indicate retransmission of the data packet.
[0009] According to an embodiment of the present disclosure, the method further includes: in response to a target task's invocation of candidate data, transferring the candidate data from a second computing resource to a first computing resource.
[0010] According to an embodiment of the present disclosure, the method further includes: in response to a target task's access to a target physical address of a first computing resource, converting the target physical address into a target virtual address of a second computing resource according to an address mapping relationship, where the target virtual address is used to store candidate data; and loading the candidate data to the target physical address of the first computing resource.
[0011] According to an embodiment of the present disclosure, the method further includes: in response to the completion of the transfer of candidate data, deleting the candidate data in the first computing resource.
[0012] According to an embodiment of the present disclosure, the method further includes: splitting the candidate data to obtain a plurality of slice data, each slice data containing a complete semantic unit; and encapsulating the slice data into a plurality of data packets.
[0013] A second aspect of the present disclosure provides a resource scheduling apparatus, including: an acquisition module, configured to acquire candidate data of a target task, where the candidate data represents data in a first computing resource that does not meet the requirements for executing the target task; a transmission module, configured to, in the presence of an address mapping relationship, transfer the candidate data from the first computing resource to a second computing resource according to the address mapping relationship, where the address mapping relationship represents the corresponding relationship between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource; and in the absence of an address mapping relationship, interrupting the data transmission between the first computing resource and the second computing resource.
[0014] A third aspect of the present disclosure provides an electronic device, including: an electronic device, including:
[0015] A first computing resource; a second computing resource, used as an extended resource of the first computing resource; a processor, configured to acquire candidate data of a target task, where the candidate data represents data in the first computing resource that does not meet the requirements for executing the target task; in the presence of an address mapping relationship, transfer the candidate data from the first computing resource to the second computing resource according to the address mapping relationship, where the address mapping relationship represents the corresponding relationship between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource; and in the absence of an address mapping relationship, interrupting the data transmission between the first computing resource and the second computing resource.
[0016] A fourth aspect of the present disclosure further provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is caused to execute the above resource scheduling method.
[0017] The fifth aspect of the present disclosure also provides a computer program product, including a computer program which, when executed by a processor, implements the above-mentioned resource scheduling method.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become clearer. In the drawings:
[0020] Figure 1 Schematically shows a flowchart of the resource scheduling method according to an embodiment of the present disclosure;
[0021] Figure 2 Schematically shows a schematic diagram of resource scheduling between video memory and a hard disk according to an embodiment of the present disclosure;
[0022] Figure 3 Schematically shows a schematic diagram of the processing principle of the address mapping relationship between video memory and a hard disk in the resource scheduling method according to an embodiment of the present disclosure;
[0023] Figure 4 Schematically shows an architecture diagram of the resource scheduling method according to an embodiment of the present disclosure;
[0024] Figure 5A Schematically shows one of the schematic diagrams of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example;
[0025] Figure 5B Schematically shows another schematic diagram of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example;
[0026] Figure 5C Schematically shows a third schematic diagram of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example;
[0027] Figure 5D Schematically shows a third schematic diagram of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example;
[0028] Figure 5E Schematically shows a fourth schematic diagram of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example;
[0029] Figure 6 Schematically shows a structural block diagram of the resource scheduling device according to an embodiment of the present disclosure;
[0030] Figure 7 FIG. 0 schematically shows a block diagram of an electronic device suitable for implementing a resource scheduling method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0031] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0032] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0033] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0034] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C).
[0035] Embodiments of the present disclosure provide a resource scheduling method and an electronic device. Before introducing the technical solutions provided by the embodiments of the present disclosure, the related technologies involved in the present disclosure will be described first.
[0036] During the execution of different tasks (such as model training), the demand for computing resources shows an exponential growth. If the computing resource capacity is insufficient, it is easy to cause task interruption or affect the execution effect of the task.
[0037] Exemplarily, the demand for video memory in large model training shows an exponential growth. Insufficient single-card video memory capacity is likely to cause training interruption or a decline in the performance of model training.
[0038] To address the problem of insufficient single - card video memory capacity, on the one hand, a memory swapping solution can be adopted: video memory (GPU) → system memory (RAM) → hard disk (SSD). That is, the data not in use in the video memory is first transferred to the memory (RAM) via the PCIe bus, and then temporarily stored or transferred to the SSD by the memory (RAM) via the PCIe bus. However, in this solution, the memory bandwidth (about 50GB / s) is lower than the bandwidth of PCIe 4.0×16 (about 63GB / s bidirectional), becoming a data transfer bottleneck. And it requires the CPU to schedule data transfer, introducing an additional 30% latency, resulting in a slower model training speed. On the other hand, the video memory of multiple graphics cards can be pooled through NVLink (a high - speed GPU interconnection technology) to expand the total capacity. However, this solution only supports high - end graphics cards of the same brand (such as NVIDIA A100), with extremely high costs (3 - 5 times increase). And it requires restructuring the distributed training code, with high development complexity. On yet another hand, the video memory data can be stored in the SSD after real - time compression to save video memory occupancy. However, this solution introduces a 10 - 15% loss of GPU computing power during compression / decompression, and in some scenarios, it is also likely to cause a decline in model accuracy.
[0039] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure are explained, and the nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations.
[0040] PCIe (PCI Express) is a high - speed bus for communication between the CPU (memory) and devices such as GPU (video memory) and SSD (hard disk).
[0041] CRC (Cyclic Redundancy Check) is a data checksum used to detect errors during transmission. Traditional CRC: The PCIe standard packet header comes with a CRC (LCRC), which only checks the header and does not check the data part.
[0042] The ioctl (Input / Output Control) command processes user - space requests. For example, it sends NVMe I / O commands (such as reading and writing to the SSD).
[0043] NVMe instructions are a standardized command set defined in the NVM Express (NVMe) protocol for operations such as data transfer, device management, and status control between the host (CPU) and NVMe SSD.
[0044] The following will be through Figures 1 to 4 Describe in detail the resource scheduling method of the embodiments of the present disclosure.
[0045] Figure 1 Schematically shows a flowchart of the resource scheduling method according to the embodiments of the present disclosure.
[0046] As shown Figure 1 in the figure, the resource scheduling method of this embodiment includes operations S210 to S230.
[0047] In operation S210, candidate data for the target task is obtained, and the candidate data represents data in the first computing resource that does not meet the requirements for executing the target task.
[0048] In operation S220, in the case where there is an address mapping relationship, according to the address mapping relationship, the candidate data is transmitted from the first computing resource to the second computing resource. The address mapping relationship represents the corresponding relationship between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource.
[0049] In operation S230, in the case where there is no address mapping relationship, the data transmission between the first computing resource and the second computing resource is interrupted.
[0050] Exemplarily, the target task can be an operation that requires computing resources during execution. For example, the target task can be model training, data analysis, image processing, etc.
[0051] The computing resource can be the hardware or software resources required to execute the target task. For example, the computing resource can be memory (RAM), solid-state drive (SSD), bandwidth, etc. The types of the first computing resource and the second computing resource are different, and data cannot be directly transmitted between the first computing resource and the second computing resource. It is necessary to rely on other computing resources for transfer to achieve data transfer and storage. Taking model training as an example of the target task, the first computing resource can be video memory (GPU), and the second computing resource can be a hard disk (SSD). In the case where there is no address mapping relationship between the physical address of the video memory and the virtual address of the hard disk, the data that is temporarily not used in the video memory (candidate data) needs to be transferred to the hard disk. First, the candidate data is transferred and stored in the memory (RAM) through the PCIe bus, and then transferred from the memory (RAM) to the hard disk through the PCIe bus to achieve the transfer and storage of the candidate data between the video memory and the hard disk.
[0052] The candidate data can be redundant data during the execution of the target task in the first computing resource. That is, during the execution of the target task, the activity level of the candidate data is less than the first threshold. In other words, the candidate data can be that during the execution of the target task, the number of accesses to the candidate data by the system is less than the second threshold. The candidate data may not be directly useful in the target task but may be data required for a future task or calculation step. For example, in model training, the data can be divided into a training set and a test set. The data required for the training task is the test set, and the data required for the test task is the test set. During the training task, the data in the test set is the candidate data.
[0053] The address mapping relationship can be the correspondence between the physical address of the first computing resource and the virtual address of the second computing resource. The physical address represents the actual storage location of the candidate data in the first computing resource, and the virtual address represents the logical address used to access the candidate data in the second computing resource. For example, when the first computing resource is the video memory and the second computing resource is the hard disk, the hard disk can be regarded as an extended space of the video memory, and the storage space of the hard disk is regarded as "virtual video memory". A PCIe Memory Management Unit (PCIe-MMU) can be added to the PCIe protocol layer hardware to be responsible for the direct mapping between the physical address of the video memory and the logical block address (LBA) of the hard disk, realizing the direct connection at the physical layer between the video memory and the hard disk. When data is transferred between the video memory and the hard disk, data can be directly transferred through PCIe, bypassing the traditional memory exchange path (video memory → memory → SSD), without going through the memory for transfer, eliminating the CPU scheduling overhead and the memory bandwidth bottleneck.
[0054] It can be understood that by establishing the mapping relationship between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource, the direct exchange of data between the first computing resource and the second computing resource can be realized without going through the transfer of other computing resources, so as to expand the space of the first computing resource. Moreover, data can be directly transferred between the first computing resource and the second computing resource, reducing latency and improving the effect and efficiency of task execution.
[0055] As described above, for the address mapping relationship, it can be obtained through the following operations: according to the physical address of the candidate data in the first computing resource, determine the virtual address of the candidate data in the second computing resource, and establish the address mapping relationship.
[0056] In one example, continuing with the above example where the first computing resource is the video memory (GPU) and the second computing resource is the hard disk (SSD) for illustration. The hardware-level mapping from the physical address of the video memory to the logical block address (LBA) of the SSD can be realized through the PCIe Memory Management Unit (PCIe MMU), enabling the GPU to directly address the SSD. Record the correspondence between the virtual address range in the video memory and the logical block address (LBA) of the SSD to obtain the address mapping relationship.
[0057] When the GPU attempts to access the video memory address 0xA0000000 (which is actually mapped to the SSD), query the address mapping relationship, and obtain that the address 0xA0000000 in the video memory corresponds to the address 1024 in the hard disk. Generate a PCIe read request and send it to the SSD. The SSD returns the data corresponding to the address 1024 and writes it to the address 0xA0000000 of the GPU video memory.
[0058] As described above, in operation S220, according to the address mapping relationship, the candidate data is transmitted from the first computing resource to the second computing resource. In one implementable manner, this operation may further include operations S221 to S222.
[0059] In operation S321, the candidate data is encapsulated into multiple data packets, and the data volume of the data packets is the same as the capacity of the read / write unit of the second computing resource.
[0060] In operation S322, the multiple data packets are transmitted to the second computing resource in parallel through multiple channels.
[0061] Exemplarily, the minimum unit of PCIe data transmission is a Transaction Layer Packet (TLP). For example, the transmission granularity of a traditional TLP is relatively large (such as 128B - 2KB), while the read / write unit of a hard disk SSD is usually a 4KB block. If an 8KB video memory block is directly transmitted, the SSD needs to be internally split, which may increase the latency.
[0062] In one example, a large data request (such as 16KB) of the GPU video memory is split into multiple 4KB TLP packets to match the read / write granularity of the SSD. Multiple channels in the PCIe link can transmit the TLP packets in parallel, allowing the later - sent packets to arrive first, avoiding blocking the entire link due to the latency of a certain packet. The data packets returned by the SSD are re - combined into continuous data required by the video memory at the PCIe layer.
[0063] It can be understood that by adapting the transmitted data packets to the physical characteristics of the second computing resource (aligning the read / write unit), redundant operations are reduced, thereby reducing the additional operation latency of the second computing resource. At the same time, the split data packets can be scheduled more flexibly, avoiding blocking the link due to large packets.
[0064] To facilitate the understanding of the direct transmission of data between the video memory and the hard disk, the following will be further described in combination with Figure 2 for further illustration.
[0065] Figure 2 Schematically shows a schematic diagram of resource scheduling between the video memory and the hard disk according to an embodiment of the present disclosure.
[0066] Refer to Figure 2, in one example, the high - bandwidth memory of the GPU (such as HBM2e) is used to store data such as model parameters and gradients during training. However, the video memory capacity is limited, and when training large models, data needs to be frequently swapped to external storage. A silver - disk SSD can be used as an extended storage for the video memory, providing high - speed read and write through the NVMe protocol. The PCIe bus, which is the data transfer path between the GPU and the SSD, is located on the same chip (PCIe Root Complex). A hardware - level mapping from the physical address of the video memory to the logical block address (LBA) of the SSD can be achieved through the PCIe memory management unit, enabling the GPU to directly address the SSD. The data packet is split into data packets with a granularity consistent with the hard - disk read - write unit size (matching the SSD block size) to achieve atomic transfer and improve transfer efficiency.
[0067] As described above, the resource scheduling method of the embodiments of the present disclosure may further include operations: determining a recombination table according to multiple data packets, where the recombination table includes address information of multiple data packets in the first computing resource; storing the multiple data packets and the recombination table corresponding to the data packets in the second computing resource.
[0068] Exemplarily, the recombination table may be used to record the logical order of data blocks corresponding to multiple data packets. For example, data packet 1 is in the front and data packet 2 is in the back. The recombination table is stored in both the first computing resource and the second computing resource.
[0069] In one example, multiple data packets are sequentially stored in the second computing resource in the order they arrive at the second computing resource.
[0070] If the SSD receives out - of - order TLP packets (for example, first receives data packet 2 and then data packet 1), then first stores TLP 2 in the first position in the SSD and stores TLP 1 in the second position in the SSD.
[0071] It should be noted here that in the case where the system needs to call the candidate data corresponding to TLP 1 and TLP 2 stored in the SSD, TLP 1 and TLP 2 need to be recombined according to the logical order recorded in the recombination table and then loaded into the video memory.
[0072] By querying the recombination table, it is obtained that the recombination needs to be in the order of TLP 1→TLP 2. After all TLP packets arrive, they are stored in the SSD in order.
[0073] In another example, multiple data packets are stored in the second computing resource in the logical order represented by the address information.
[0074] The SSD receives out-of-order arrived TLP packets (e.g., first receives packet 2 and then receives packet 1). By querying the reorganization table, it is obtained that the TLP packets need to be reorganized in the order of TLP 1 → TLP 2. After all TLP packets arrive, they are stored in the SSD in order.
[0075] It can be understood that by storing multiple data packets in the second computing resource in order, the number of read and write operations on the second computing resource can be reduced, and the lifespan of the hardware of the second computing resource can be improved.
[0076] As described above, the data packet includes the first payload information.
[0077] Exemplarily, the first payload information may include the packet header and data payload of the data packet. Through the enhanced CPC (ECRC): additionally verify the entire TLP packet (including the header + data payload) to improve the fault tolerance of data transmission.
[0078] When the candidate data is transmitted in the form of TLP (Transaction Layer Packet) packets through PCIe, it is vulnerable to electromagnetic noise, resulting in transmission errors. To avoid transmission errors. The resource scheduling method of the embodiments of the present disclosure may further include operations: in response to the second computing resource receiving a data packet, calculating the second payload information of the data packet; in the case where the second payload information is consistent with the first payload information, storing the data packet in the second computing resource; in the case where the second payload information is inconsistent with the first payload information, generating a control signal, and the control signal is used to indicate retransmission of the data packet.
[0079] In one example, when the video memory sends a data packet, it calculates the entire packet header and data payload of the data packet and attaches them to the end of the packet. After the SSD receives the data packet, it recalculates the ECRC of the data packet and compares it with the recorded value at the end of the received data packet. In the case where the information comparison is inconsistent, it is determined that a transmission error has occurred, and a control signal (e.g., Negative Acknowledge, NAK signal) is immediately sent to the video memory to request retransmission of the data packet. In the case where the information comparison is consistent, it is determined that the transmission is correct, and the data packet is stored in the SSD.
[0080] It should be noted here that if the data packet transmission fails (i.e., the first payload information and the second payload information are inconsistent), the data packet will not be stored in the SSD but will be completely discarded to ensure data consistency.
[0081] As described above, the resource scheduling method of the embodiments of the present disclosure may further include operations: in response to a target task calling the candidate data, transmitting the candidate data from the second computing resource to the first computing resource.
[0082] In one example, during the execution of a target task by the system, for data in the video memory that is not required for the execution of the target task (candidate data), according to the address mapping relationship, the data can be directly transferred through PCIe and the candidate data can be transferred and stored to the hard disk, so that only the data related to the current execution of the target task remains in the video memory. During the execution of a task related to the candidate data by the system, the candidate data previously stored in the hard disk needs to be loaded from the hard disk into the video memory. At the same time, the data of the current task in the video memory is transferred and stored to the hard disk, and the transfer and call of the data are realized in this cycle.
[0083] As described above, for the above operation: in response to the target task's call for candidate data, the candidate data is transferred from the second computing resource to the first computing resource. This operation can further include the operations of: in response to the target task's access to the target physical address of the first computing resource, according to the address mapping relationship, converting the target physical address into the target virtual address of the second computing resource, where the target virtual address is used to store the candidate data; and loading the candidate data into the target physical address of the first computing resource.
[0084] In one example, when the GPU attempts to access candidate data corresponding to the video memory address 0xA0000000 (which is actually mapped to the SSD), the address mapping relationship is queried, and it is obtained that the address 0xA0000000 in the video memory corresponds to the address 1024 in the hard disk. A PCIe read request is generated and sent to the SSD. The SSD returns the data corresponding to the address 1024 and writes it to the address 0xA0000000 of the GPU video memory. Multiple data packets of the candidate data in the hard disk can be transmitted in parallel through multiple channels. After reaching the video memory in sequence, the video memory reorganizes the multiple data packets according to the logical order of each data packet in the reorganization table and stores the reorganized multiple data packets in sequence in the video memory.
[0085] As described above, the resource scheduling method of the embodiments of the present disclosure can further include the operation of: deleting the candidate data in the first computing resource in response to the completion of the transmission of the candidate data.
[0086] In one example, after all the candidate data in the video memory is transferred and stored to the hard disk, the storage space occupied by the candidate data in the video memory is released to improve the utilization rate of the resources in the video memory.
[0087] As described above, the resource scheduling method of the embodiments of the present disclosure can further include the operations of: splitting the candidate data to obtain multiple slice data, where each slice data contains a complete semantic unit; and encapsulating the slice data into multiple data packets.
[0088] Exemplarily, the slice data can be obtained by decomposing ultra-large-scale data into slices that can be processed in batches. The slice data contains a complete semantic unit. For example, it can be the weight matrix of a certain layer of model training.
[0089] For example, a single graphics card's video memory (such as 24GB) cannot hold the parameter data of a large language model (such as a model with 175 billion parameters). The model parameters of the large language model can be sliced into multiple slices (Slices) by layer or tensor. For example, each slice is 4MB. Only the slices required for the current calculation are retained in the video memory, and the rest are temporarily stored in the SSD. When it is necessary to transfer the slice data from the video memory to the SSD, the slice data is split into multiple data packets of the same size as the SSD read / write unit, and according to the address relationship mapping table between the video memory and the SSD, the multiple data packets are sent to the SSD through the PCIe link, and the SSD controller writes to the physical flash memory according to the LBA.
[0090] To facilitate the understanding of the resource scheduling method of the embodiments of the present disclosure, the following will be further described in conjunction with Figure 3 and Figure 4 for further illustration.
[0091] Figure 3 Schematically shows the processing principle diagram of the address mapping relationship between the video memory and the hard disk in the resource scheduling method according to the embodiments of the present disclosure.
[0092] Referring to Figure 3 , continuing with the first computing resource being the video memory and the second computing resource being the hard disk as an example.
[0093] In the application layer, Python code (such as PyTorch / TensorFlow) calls an interface similar to tensor_exit to mark the video memory data (i.e., candidate data, such as inactive tensors) that needs to be swapped out to the SSD. Obtain the unique device address (Unique Device Address, UIDA) of the data to be swapped out in the GPU video memory for subsequent hardware addressing. Allocate logical block address (LBA) space on the SSD and establish a mapping relationship with the video memory address.
[0094] In the kernel driver layer (CPU), the CPU, through the PCIe memory management unit (PCIe MMU), performs a hardware-level mapping between the physical address of the GPU video memory and the LBA of the SSD, enabling the PCIe DMA engine to directly address. Write the GPU address register: specify the source data address in the video memory. Write the SSD LBA register: specify the target storage location in the SSD. Start the PCIe controller to initiate direct data transfer between the video memory and the SSD (without the CPU participating in data copying).
[0095] The hardware layer is implemented in the PCIe bus controller and is responsible for executing the physical data transfer between the video memory and the SSD, using the PCIe bandwidth (such as PCIe 4.0×16 bidirectional 63GB / s) to transfer data. After the transfer is completed, the hardware generates a message signal interrupt (MSI-X) to notify the CPU. After receiving the interrupt signal, the CPU calls the interrupt handler or file system callback of the kernel to update the data status in the SSD. The user-space application (such as Python) is notified that the data eviction operation has been completed and can continue to execute subsequent computing tasks.
[0096] Figure 4 Schematically shows an architecture diagram of a resource scheduling method according to an embodiment of the present disclosure.
[0097] Refer to Figure 4 , in the hardware layer, a PCIe memory management unit is added to the CPU: used to manage the address mapping relationship between the GPU video memory and the SSD. The PCIe controller is used to control the data transfer from the video memory to the SSD, and the packet processor is used to split / recombine PCIe packets. The hard disk SSD is directly connected to the GPU through the PCIe interface, and the NAND array stores the evicted video memory data (that is, the data evicted from the GPU video memory is temporarily stored using the NAND flash array of the SSD). Therefore, the physical address direct mapping between the GPU video memory and the SSD can be realized without memory transfer.
[0098] The kernel-mode driver notifies the CPU through an interrupt signal (MSI-X) after the packet transfer is completed, and subsequent computing tasks can continue to be executed. At the same time, the status of the packet is verified: the first payload information at the end of the packet is compared with the second payload information of the newly calculated packet of the SSD. The controller is used to manage the descriptor queue, for example, batch aggregation requests. It is also responsible for implementing the address conversion between the video memory and the hard disk. MMU configuration (address mapping relationship table in the video memory and the hard disk): maintains the mapping table between the GPU video memory and the SSD LBA (logical block address). PRP list configuration (address mapping relationship table in the hard disk and the video memory): the physical region page table in the NVMe protocol of the hard disk, describing the distribution of data in the SSD. The NVMe instruction synchronizer is used to manage the submission queue (SQ) and completion queue (CQ) of NVMe; at the same time, it is also used to ensure the synchronization of NVMe instructions (such as read / write) initiated by the GPU with the SSD.
[0099] User-mode middleware integrates a deep learning framework and automatically captures tensors (candidate data) to be swapped out in the video memory through the framework's extension interface (such as register_post_hook in PyTorch). TensorFlow operators cooperate with the PCIe bus (DirectSwap API) to be responsible for the exchange of candidate data between the GPU memory and the NVMe SSD. Through the PCIe bus, tensors can be migrated from the video memory to the SSD or loaded from the SSD back into the video memory. The candidate data manager is used to track pinned memory (preventing it from being swapped out by the system) and maintain metadata (such as the location, size, and access frequency of candidate data).
[0100] Through TLP batch aggregation: Combine multiple small requests to reduce PCIe protocol overhead. Interrupt mechanism: Reduce the impact on the CPU from high-frequency interrupts. Lossless descriptor ring: Efficiently manage DMA descriptors to avoid memory copying.
[0101] In one example, PyTorch training triggers swap_out(tensor), and the middleware marks the tensor as swappable. The candidate data manager locks the video memory page and generates metadata (such as LBA mapping). The kernel driver configures the DMA descriptor and initiates a PCIe atomic TLP transfer (4KB×N). The SSD receives the data and writes it to the NAND, and notifies completion through an MSI-X interrupt. The GPU video memory releases space, and the tensor metadata is updated to "located on the SSD".
[0102] To facilitate the understanding of the resource scheduling method in the embodiments of the present disclosure, the target task is taken as an example of model training for further illustration.
[0103] Figure 5A Schematically shows one of the schematic diagrams of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example; Figure 5B Schematically shows another schematic diagram of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example; Figure 5C Schematically shows a third schematic diagram of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example; Figure 5D Schematically shows a third schematic diagram of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example; Figure 5E Schematically shows a fourth schematic diagram of the resource scheduling method according to an embodiment of the present disclosure, taking model training as an example.
[0104] The model training process includes: forward propagation, backward propagation, and gradient descent. The model parameters are sliced to obtain multiple slice data (slice 1, slice 2,..., slice n).
[0105] The data exchange between the video memory and the hard disk can be managed by an intelligent scheduling module. Here, the intelligent scheduling module is similar to the management function of the above user-mode middleware.
[0106] Referring to Figure 5A , it schematically shows the process of data exchange of model parameters between the video memory and the hard disk during the forward propagation of model training. In the video memory, there are the layer parameters currently being calculated (such as slice 1) and the activation values of the current layer. In the hard disk, there are inactive parameters (such as slice 2 to slice n) and historical activation values. Slice 1 and the activation values of the current layer are retained in the video memory, and the rest of the data (such as slice 2) has not been loaded. During forward propagation, slice 2 is loaded from the SSD to the video memory, replacing slice 1 that is no longer needed. The slice data flows between the video memory and the SSD as needed. After multiple iterations, until slice n is loaded from the SSD to the video memory.
[0107] Referring to Figure 5B , it schematically shows the process of data exchange of model parameters between the video memory and the hard disk during the backward propagation of model training. The video memory includes the slice data of the current model parameters, the activation values, gradient values, and loss values of the current layer. The hard disk includes other slice data of the model parameters, historical activation values, and gradient values. The video memory only retains the slice (such as slice n) required for the current calculation, and the rest is temporarily stored in the SSD. After calculating the gradient of slice n during backward propagation, slice n - 1 is loaded into the video memory. After multiple iterations, until slice 1 is loaded from the SSD to the video memory.
[0108] Referring to Figure 5C , it schematically shows the process of data exchange of model parameters between the video memory and the hard disk during the gradient descent of model training. The video memory includes the slice data of the current model parameters, the gradient values of the current layer, and the state of the current optimizer. The model parameters corresponding to the slices in the optimizer are updated using the gradient descent algorithm. The hard disk includes other slice data of the model parameters, gradient values, and the state of the historical optimizer. The video memory only retains the slice (such as slice n) required for the current calculation, and the rest is temporarily stored in the SSD. The video memory loads slice n and the optimizer state. After calculating the gradient of slice n during backward propagation, the parameters and optimizer state of slice n are updated according to the gradient. The optimizer state is updated synchronously with the parameter slices, and the optimizer state of slice n in the hard disk is also updated. Then slice n - 1 is loaded into the video memory for further gradient calculation and update. After multiple iterations, until slice 1 is loaded from the SSD to the video memory, completing the optimization of the model parameters.
[0109] Referring to Figure 5D, schematically shows the process of data exchange and update of model parameters in model training between video memory and hard disk. The video memory includes slice data of the current model parameters, gradient values of the current layer, and the state of the current optimizer. The model parameters corresponding to the slices in the optimizer are updated using the gradient descent algorithm. The hard disk includes other slice data of the model parameters, gradient values, and the state of the historical optimizer. The video memory only retains the slice (such as slice n) required for the current calculation, and the rest is temporarily stored in the SSD. The video memory loads slice n and the optimizer state from the hard disk. After backpropagating to calculate the gradient of slice n, the parameters and optimizer state of slice n are updated according to the gradient. The optimizer state is updated synchronously with the parameter slice, and the optimizer state of slice n in the hard disk is updated. Then, the updated slice n is transferred to the hard disk. Then slice n - 1 is loaded into the video memory for further gradient calculation and update. After multiple iterations, until slice 1 is loaded from the SSD into the video memory to complete the optimization of the model parameters, the data of the updated slices n - 1 are transferred to the hard disk.
[0110] Referring to Figure 5E , schematically shows the processing method when the model parameters in model training are completed. After the model parameters are trained, the slice data of the updated model parameters are transferred to the hard disk, and the slice data in the video memory are cleared.
[0111] Based on the above resource scheduling method, the present disclosure also provides a resource scheduling device. The following will be combined with Figure 6 to describe the device in detail.
[0112] Figure 6 Schematically shows a structural block diagram of a resource scheduling device according to an embodiment of the present disclosure.
[0113] As Figure 6 shown, the acquisition device 300 of this embodiment includes an acquisition module 310 and a transmission module 320.
[0114] The acquisition module 310 is used to acquire candidate data for a target task, and the candidate data represents data in the first computing resource that does not meet the requirements for executing the target task. In one embodiment, the acquisition module 310 can be used to perform the operation S210 described above, which will not be elaborated here.
[0115] The transmission module 320 is used to, in the case of an address mapping relationship, transfer the candidate data from the first computing resource to the second computing resource according to the address mapping relationship, where the address mapping relationship represents the corresponding relationship between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource; in the case of no address mapping relationship, interrupt the data transmission between the first computing resource and the second computing resource. In one embodiment, the transmission module 320 can be used to perform the operation S220 described above, which will not be elaborated here.
[0116] According to an embodiment of the present disclosure, any plurality of modules among the acquisition module 310 and the transmission module 320 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the acquisition module 310 and the transmission module 320 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system in package, an application specific integrated circuit (ASIC), or any other reasonable manner that can integrate or package circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in any appropriate combination of several of them. Alternatively, at least one of the acquisition module 710 and the transmission module 720 may be at least partially implemented as a computer program module, and when the computer program module is run, it can execute corresponding functions.
[0117] The electronic device according to an embodiment of the present disclosure includes: a first computing resource; a second computing resource for serving as an extended resource of the first computing resource.
[0118] A processor is configured to acquire candidate data for a target task, where the candidate data represents data in the first computing resource that does not meet the requirements for executing the target task; in the case of an address mapping relationship, according to the address mapping relationship, transmit the candidate data from the first computing resource to the second computing resource, where the address mapping relationship represents the correspondence between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource; in the case of no address mapping relationship, interrupt the data transmission between the first computing resource and the second computing resource.
[0119] Exemplarily, a PCIe memory management unit and an intelligent scheduling unit are integrated in the processor. The PCIe memory management unit is configured to establish an address mapping relationship between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource. The intelligent scheduling unit is configured to split all the data of the target task into slice data, and transmit the slice data that is not required at the first moment for executing the target task (candidate data), by splitting it into data packets, from the first computing resource to the second computing resource. Load the slice data required at the second moment for executing the target task from the second computing resource to the first computing resource. And execute the target task in this cycle.
[0120] Figure 7 A block diagram of an electronic device suitable for implementing a resource scheduling method according to an embodiment of the present disclosure is schematically shown.
[0121] As Figure 7As shown, an electronic device 400 according to an embodiment of the present disclosure includes a processor 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage section 408 into a random access memory (RAM) 403. The processor 401 can include, for example, a general-purpose microprocessor (e.g., CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 401 can also include on-board memory for caching purposes. The processor 401 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0122] In the RAM 403, various programs and data required for the operation of the electronic device 400 are stored. The processor 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. The processor 401 performs various operations of the method flow according to an embodiment of the present disclosure by executing programs in the ROM 402 and / or the RAM 403. It should be noted that the program can also be stored in one or more memories other than the ROM 402 and the RAM 403. The processor 401 can also perform various operations of the method flow according to an embodiment of the present disclosure by executing programs stored in one or more memories.
[0123] According to an embodiment of the present disclosure, the electronic device 400 may further include an input / output (I / O) interface 405, and the input / output (I / O) interface 405 is also connected to the bus 404. The electronic device 400 may further include one or more of the following components connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read from it can be installed into the storage section 408 as needed.
[0124] The present disclosure also provides a computer-readable storage medium, which may be included in the device / device / system described in the above embodiment; or may exist separately and not be assembled into the device / device / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to an embodiment of the present disclosure is implemented.
[0125] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, which may include, for example, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program may be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the above-described ROM 402 and / or RAM 403 and / or one or more memories other than ROM 402 and RAM 403.
[0126] An embodiment of the present disclosure further includes a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the resource scheduling method provided by the embodiment of the present disclosure.
[0127] When the computer program is executed by the processor 401, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0128] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part 409, and / or be installed from the removable medium 411. The program code contained in the computer program may be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0129] In such an embodiment, the computer program may be downloaded and installed from the network through the communication part 409, and / or be installed from the removable medium 411. When the computer program is executed by the processor 401, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. may be implemented by computer program modules.
[0130] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0132] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0133] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present disclosure.
Claims
1. A resource scheduling method, comprising: Obtaining candidate data of a target task, where the candidate data represents data in a first computing resource that does not meet the requirements for executing the target task; In the case of an address mapping relationship, according to the address mapping relationship, transmitting the candidate data from the first computing resource to a second computing resource, where the address mapping relationship represents the correspondence between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource; In the case where the address mapping relationship does not exist, interrupting the data transmission between the first computing resource and the second computing resource.
2. The method according to claim 1, the method further comprising: Determining the virtual address of the candidate data in the second computing resource according to the physical address of the candidate data in the first computing resource, and establishing the address mapping relationship.
3. The method according to claim 1, transmitting the candidate data from the first computing resource to a second computing resource according to the address mapping relationship, comprising: Encapsulating the candidate data into a plurality of data packets, where the data volume of the data packet is the same as the capacity of the read / write unit of the second computing resource; Transmitting the plurality of data packets to the second computing resource in parallel through a plurality of channels.
4. The method according to claim 3, the method further comprising: Determining a recombination table according to the plurality of data packets, where the recombination table includes the address information of the plurality of data packets in the first computing resource; Storing the plurality of data packets and the recombination table corresponding to the data packets in the second computing resource.
5. The method according to claim 3, the data packet includes first payload information, the method further comprising: In response to the second computing resource receiving the data packet, calculating second payload information of the data packet; In the case where the second payload information is consistent with the first payload information, storing the data packet in the second computing resource; In the case where the second payload information is inconsistent with the first payload information, generating a control signal, where the control signal is used to indicate retransmitting the data packet.
6. The method according to claim 1, the method further comprising: In response to the target task calling the candidate data, transmitting the candidate data from the second computing resource to the first computing resource.
7. The method according to claim 6, the method further comprising: In response to the target task accessing the target physical address of the first computing resource, according to the address mapping relationship, converting the target physical address into the target virtual address of the second computing resource, where the target virtual address is used to store the candidate data; Loading the candidate data to the target physical address of the first computing resource.
8. The method according to claim 1, the method further comprising: In response to the completion of the transmission of the candidate data, deleting the candidate data in the first computing resource.
9. The method according to claim 3, the method further comprising: Slice the candidate data to obtain multiple slice data, each slice data containing a complete semantic unit; Encapsulate the slice data into multiple data packets.
10. An electronic device, comprising: A first computing resource; A second computing resource for serving as an extended resource of the first computing resource; A processor for obtaining candidate data of a target task, the candidate data representing data in the first computing resource that does not meet the requirements for executing the target task; in the case of an address mapping relationship, according to the address mapping relationship, transmit the candidate data from the first computing resource to the second computing resource, the address mapping relationship representing the correspondence between the physical address of the candidate data with respect to the first computing resource and the virtual address of the second computing resource; In the case of the absence of the address mapping relationship, interrupt the data transmission between the first computing resource and the second computing resource.