Memory resource scheduling method and electronic device
Patent Information
- Application Number
- CN202512058743.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-12-31
AI Technical Summary
这导致部分内存处于高负荷的繁忙状态或内存不足时,部分内存却因任务需求未达上限而处于闲置状态,造成各加速器间的内存负载不均
[0008] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
Smart Images

Figure CN121433922B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multi-accelerator system technology, and more specifically to a memory resource scheduling method and electronic device. Background Technology
[0002] Accelerators are dedicated processors used for high-performance computing, and their computations rely on their own onboard memory. Currently, in multi-accelerator server deployments, the memory of each accelerator is typically used independently by its own running tasks. This results in some memory being under heavy load or insufficient, while other memory remains idle due to unmet task demands, causing uneven memory load among the accelerators. When an accelerator's memory is insufficient, the system needs to exchange data with the Central Processing Unit (CPU)'s memory. However, exchanging data with the CPU's memory via the high-speed Peripheral Component Interconnect Express (PCIe bus) is often limited by the bandwidth and latency of buses such as PCIe, affecting the overall execution efficiency of the computing tasks. Summary of the Invention
[0003] In view of the above problems, this application provides a memory resource scheduling method and an electronic device.
[0004] The first aspect of this application provides a memory resource scheduling method, the method comprising: for at least one accelerator in a multi-accelerator system, when the memory occupancy rate of at least one accelerator is detected to reach a preset threshold, determining at least one idle accelerator from the other accelerators in the multi-accelerator system, excluding at least one accelerator, based on the memory status information of each accelerator in the multi-accelerator system, wherein the idle accelerator is an accelerator with idle memory; determining a target accelerator for at least one accelerator from the idle accelerators based on the idle memory size of the idle accelerator and the link distance of the idle accelerator relative to at least one accelerator; and migrating the data to be swapped out in at least one accelerator to the memory space of the target accelerator.
[0005] A second aspect of this application provides a memory resource scheduling apparatus, comprising: a first determining module, configured to, for at least one accelerator in a multi-accelerator system, determine at least one idle accelerator from the other accelerators in the multi-accelerator system, excluding the at least one accelerator, based on memory status information of each accelerator in the multi-accelerator system, when the memory occupancy rate of the at least one accelerator is detected to reach a preset threshold, wherein the idle accelerator is an accelerator with idle memory; a second determining module, configured to determine a target accelerator for the at least one accelerator from the idle accelerators, based on the idle memory size of the idle accelerator and the link distance of the idle accelerator relative to the at least one accelerator; and a migration module, configured to migrate data to be swapped out from the at least one accelerator to the memory space of the target accelerator.
[0006] A third aspect of this application provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.
[0007] A fourth aspect of this application also provides a computer-readable storage medium having a computer program or instructions stored thereon, which, when executed by a processor, implement the steps of the above-described method.
[0008] The fifth aspect of this application also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.
[0009] By monitoring when the memory usage of any accelerator reaches a threshold, idle accelerators with available memory are automatically selected from other accelerators. This allows for the timely discovery and utilization of distributed memory resources, alleviating memory pressure on a single accelerator. Instead of swapping data out to processor memory, idle memory from other accelerators within the same server is utilized, enabling data migration via high-speed inter-accelerator interconnects and avoiding bottlenecks in PCIe bus bandwidth and latency. By comprehensively evaluating the idle memory size of idle accelerators and their link distance to the requesting accelerator, the optimal target accelerator in terms of capacity and access performance can be selected for efficient data migration. Leveraging the high-speed interconnect capabilities of hardware, previously idle memory resources are transformed into a usable high-performance storage pool, thereby alleviating memory pressure while maintaining the overall execution efficiency of computing tasks.
[0010] This method avoids swapping data out to higher-latency processor memory, reducing performance degradation caused by remote access, while improving the overall utilization of memory resources within the server and enhancing the computational efficiency of multiple accelerators when handling unbalanced memory loads. Furthermore, by introducing a hierarchical processing mechanism, it prioritizes the use of the high-performance memory extension layer formed by the high-speed interconnects between accelerators, only degrading to the central processing unit memory layer when no resources are available within the accelerator cluster. This fully leverages the performance of the high-speed hardware interconnects while ensuring reliability in extreme scenarios, thereby improving the memory management efficiency and resource utilization of the multi-accelerator cluster. Attached Figure Description
[0011] The above-mentioned contents, other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0012] Figure 1 The illustration shows application scenarios of a memory resource scheduling method, apparatus, electronic device, storage medium, and program product according to embodiments of this application.
[0013] Figure 2 A flowchart of a memory resource scheduling method according to an embodiment of this application is shown.
[0014] Figure 3A A schematic diagram of a traditional memory resource scheduling method is shown.
[0015] Figure 3B A schematic diagram of a memory resource scheduling method according to an embodiment of this application is shown.
[0016] Figure 4 A schematic diagram of a queue to be processed according to an embodiment of this application is shown.
[0017] Figure 5 A schematic diagram illustrating the processing of page faults according to an embodiment of this application is shown.
[0018] Figure 6A A schematic diagram of parallel prefetching of data to be prefetched according to an embodiment of this application is shown.
[0019] Figure 6B A schematic diagram of parallel prefetching of data to be prefetched is shown according to another embodiment of this application.
[0020] Figure 7 A structural block diagram of a memory resource scheduling apparatus according to an embodiment of this application is shown.
[0021] Figure 8 A block diagram of an electronic device suitable for implementing a memory resource scheduling method according to an embodiment of this application is shown. Detailed Implementation
[0022] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0023] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0025] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0026] Current servers typically integrate multiple accelerators to provide aggregated computing capabilities, and these accelerators are interconnected via high-speed interconnect technology. When running computing tasks, each accelerator needs to load data and model parameters into its own onboard memory. The widely used memory management approach treats the memory of each accelerator as an independent resource pool, allocating it only to local tasks. This static or isolated allocation model limits the overall memory resource utilization within the server to the task demands of a single accelerator. When a task's memory requirements exceed the local capacity of its accelerator, the overflowing data can only be exchanged via system buses such as PCIe to the CPU-managed main memory (e.g., dynamic random access memory) for storage. However, the bandwidth of such cross-device data exchange is low, failing to fully utilize the hardware interconnect features within the server.
[0027] In view of the above-mentioned technical problems, embodiments of this application provide a memory resource scheduling method, including: for at least one accelerator in a multi-accelerator system, when the memory occupancy rate of at least one accelerator is detected to reach a preset threshold, determining at least one idle accelerator from the other accelerators in the multi-accelerator system other than at least one accelerator based on the memory status information of each accelerator in the multi-accelerator system, wherein the idle accelerator is an accelerator with idle memory; determining a target accelerator for at least one accelerator from the idle accelerators based on the idle memory size of the idle accelerator and the link distance of the idle accelerator relative to at least one accelerator; and migrating the data to be swapped out in at least one accelerator to the memory space of the target accelerator.
[0028] Figure 1 The illustration shows application scenarios of a memory resource scheduling method, apparatus, electronic device, storage medium, and program product according to embodiments of this application.
[0029] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a request accelerator 110, a target accelerator 120, and a central processing unit 130.
[0030] The requesting accelerator 110 is a computing accelerator (such as a graphics processing unit or tensor processor) that requires memory expansion. It is used to perform computing tasks and trigger data swapping and prefetching operations when its local memory is insufficient. It is connected to the target accelerator 120 via a high-speed interconnect protocol and to the central processing unit 130 via a PCIe bus. The high-speed interconnect protocol refers to the interconnect protocol between accelerators, such as Ultra Accelerator Link (UALink) or Omni-directional Intelligent Sensing Architecture (OISA).
[0031] The target accelerator 120 is an idle accelerator that provides idle memory resources for temporarily storing data swapped out from the request accelerator 110 or loading data that may be accessed in the future in response to prefetch instructions. It is directly interconnected with the request accelerator 110 via a high-speed interconnect protocol to achieve high-speed data transmission, and is connected to the central processing unit 130 via an independent PCIe link to support cooperative operations involving system memory.
[0032] The central processing unit 130 can be used to store spare data pages and back up page data swapped out from the requesting accelerator 110 and / or the target accelerator 120. When the requesting accelerator 110 triggers a memory usage rate exceeding a preset threshold and there is an idle accelerator, the central processing unit 130 can receive and store the data swapped out from the target accelerator 120; when the requesting accelerator 110 triggers a memory usage rate exceeding a preset threshold and there is no idle accelerator, the central processing unit 130 can receive and store the data swapped out from the requesting accelerator 110; at the same time, when the target accelerator 120 needs to reclaim its temporary data memory space, the central processing unit 130 can also receive the data copy backed up from the target accelerator 120, thereby realizing the dynamic release and recycling of memory resources between accelerators.
[0033] It should be understood that Figure 1 The number of request accelerators and target accelerators shown is merely illustrative. Depending on implementation needs, there can be any number of request accelerators and target accelerators.
[0034] The following will be based on Figure 1 The described scene, through Figures 2-6B The memory resource scheduling method of the embodiments of this application will be described in detail.
[0035] Figure 2 A flowchart of a memory resource scheduling method according to an embodiment of this application is shown; Figure 3A A schematic diagram of a traditional memory resource scheduling method is shown; Figure 3B A schematic diagram of a memory resource scheduling method according to an embodiment of this application is shown. The following is in conjunction with... Figure 2 , Figure 3A and Figure 3B Please provide a detailed explanation.
[0036] like Figure 2 As shown, the memory resource scheduling method 200 of this embodiment includes operations S210 to S230, which are executed by the memory manager.
[0037] In operation S210, for at least one accelerator in the multi-accelerator system, when the memory occupancy rate of at least one accelerator is detected to reach a preset threshold, at least one idle accelerator is determined from the other accelerators in the multi-accelerator system other than at least one accelerator, based on the memory status information of each accelerator in the multi-accelerator system. The idle accelerator is an accelerator with idle memory.
[0038] When the memory utilization rate of at least one accelerator reaches a preset threshold, it indicates that the local memory utilization of that accelerator is nearing saturation and a capacity bottleneck is about to occur, which may trigger memory expansion or data swapping. Accelerators whose memory utilization rate reaches the preset threshold can be identified as requesting accelerators, indicating that the accelerator needs to request available memory space from other idle accelerators. For example... Figure 3A and Figure 3B The request accelerator shown is shown.
[0039] Memory status information refers to the current memory usage data of each accelerator, including total memory capacity, used memory capacity, and idle memory capacity. This information is used to filter accelerators with available (idle) memory capacity and to preliminarily determine other accelerators with available memory. For example, if the memory usage rate of a requesting accelerator reaches a preset threshold, based on the memory status information of each accelerator (e.g., including the first accelerator, second accelerator, third accelerator, etc.), the first and second accelerators are identified as idle accelerators with available memory.
[0040] Memory status information can be stored in the memory of the accelerator or memory scheduling device as a linked list. Each accelerator's corresponding linked list maintains information about its idle memory blocks and periodically synchronizes its status with the global memory management module via a high-speed interconnect protocol. The linked list may contain fields such as accelerator number, idle memory size, idle memory location, link information between accelerators, PCIe link information between the accelerator and the central processing unit, topology information, and redundancy information, as shown in Table 1.
[0041] Table 1
[0042]
[0043] The accelerator ID identifies the accelerator; the idle memory size records the available memory capacity; the idle memory location records the physical address range or virtual address mapping; the link information between accelerators indicates the connection status and available bandwidth; the PCIe link information indicates the link status and bandwidth capacity; the topology information indicates the relative position of the accelerator in the fully interconnected network; and the redundancy information identifies the available space for data backup. When the memory occupancy rate of any accelerator reaches a preset threshold, the idle memory location and link distance information of candidate accelerators are read from the linked list to determine the available accelerator.
[0044] For example, the physical memory of an accelerator is divided into 4MB memory blocks for management. Each memory block contains metadata describing its status and physical address. A metadata linked list of free 4MB memory blocks is maintained for each accelerator. By querying the free linked list of each accelerator, the available idle memory of each accelerator can be quickly retrieved, thus obtaining the idle accelerators.
[0045] In operation S220, a target accelerator for at least one accelerator is determined from the idle accelerators based on the idle memory size of the idle accelerators and the link distance of the idle accelerators relative to at least one accelerator.
[0046] Idle memory size refers to the current unused memory capacity of multiple candidate idle accelerators. Link distance characterizes the actual physical distance between the requesting accelerator and candidate idle accelerators (e.g., the first and second accelerators) for communication via physical interconnection, and is used to characterize or indicate the potential latency and bandwidth performance of data transmission between the two. Typically, link distance is directly proportional to latency and inversely proportional to effective bandwidth. The target accelerator is the accelerator (e.g., the first and second accelerators) ultimately selected from the idle accelerators (e.g., the first accelerator) as the remote memory extension area for the requesting accelerator to receive the swapped-out data.
[0047] In operation S230, the data to be swapped out in at least one accelerator is migrated to the memory space of the target accelerator.
[0048] Data to be swapped out refers to the portion of data selected from the requesting accelerator's local memory that needs to be moved out. This is typically data with low access frequency or low priority, to free up local memory space for new computational data. Migration refers to the process of transferring the data to be swapped out from the requesting accelerator to the target accelerator (first accelerator) via an interconnect network, achieving a physical transfer of the data's location. The target accelerator's memory space refers to the portion of idle memory on the target accelerator that is allocated to receive and store the migrated data; this serves as the requesting accelerator's effective memory.
[0049] When the requesting accelerator requests memory from the target accelerator, the memory manager marks the 4MB memory block in the metadata as originating from the target accelerator.
[0050] After operation S230, the memory resource scheduling method may further include: in the absence of a target accelerator, migrating the data to be swapped out in at least one accelerator to the memory space corresponding to the CPU.
[0051] When all other accelerators lack sufficient free memory, data to be swapped out from any accelerator can be migrated to the corresponding memory space on the CPU, preventing memory expansion failure due to a lack of a target accelerator. Although accessing CPU memory is less efficient than accessing the memory of a neighboring accelerator, it maintains computational tasks within a manageable scope compared to task stalls or interruptions caused by memory exhaustion.
[0052] By monitoring when the memory usage of any accelerator reaches a threshold, idle accelerators with available memory are automatically selected from other accelerators. This allows for the timely discovery and utilization of distributed memory resources, alleviating memory pressure on a single accelerator. Instead of swapping data out to processor memory, idle memory from other accelerators within the same server is utilized, enabling data migration via high-speed inter-accelerator interconnects and avoiding bottlenecks in PCIe bus bandwidth and latency. By comprehensively evaluating the idle memory size of idle accelerators and their link distance to the requesting accelerator, the optimal target accelerator in terms of capacity and access performance can be selected for efficient data migration. Leveraging the high-speed interconnect capabilities of hardware, previously idle memory resources are transformed into a usable high-performance storage pool, thereby alleviating memory pressure while maintaining the overall execution efficiency of computing tasks.
[0053] This method avoids swapping data out to higher-latency processor memory, reducing performance degradation caused by remote access, while improving the overall utilization of memory resources within the server and enhancing the computational efficiency of multiple accelerators when handling unbalanced memory loads. Furthermore, by introducing a hierarchical processing mechanism, it prioritizes the use of the high-performance memory extension layer formed by the high-speed interconnects between accelerators, only degrading to the central processing unit memory layer when no resources are available within the accelerator cluster. This fully leverages the performance of the high-speed hardware interconnects while ensuring reliability in extreme scenarios, thereby improving the memory management efficiency and resource utilization of the multi-accelerator cluster.
[0054] When the memory usage of any accelerator reaches a preset threshold, some data needs to be swapped out from the accelerator's local memory to free up space for the current computing task to continue. To determine the specific data to be swapped out from local memory, the Least Recently Used (LRU) strategy is used to filter the pages to be swapped out.
[0055] Specifically, migrating the data to be swapped out in any accelerator to the memory space of the target accelerator includes: pre-allocating a memory space of a predetermined memory capacity in the memory space of the target accelerator; and migrating the data to be swapped out in any accelerator to the memory space of the target accelerator according to the memory space of the predetermined memory capacity.
[0056] Pre-allocating large pages of predetermined memory capacity (e.g., 4MB) allows for the transfer of 4MB pages at a time during each data swap between the accelerator and the CPU or space accelerator, thus merging what would otherwise require 512 4KB small page transfers into a single large page operation. This strategy effectively matches the circular memory access characteristics with long reuse distances prevalent in workloads such as graph computing and deep neural networks. The large page swapping mechanism reduces page table entry processing and transfer scheduling overhead, and its general applicability can be extended to all application scenarios involving accelerator and CPU memory data exchange, thereby improving the overall management throughput and resource scheduling efficiency of heterogeneous memory systems.
[0057] When performing a swap-out operation, the memory manager first selects the least accessed memory block as a candidate swap-out object based on the local LRU list of each accelerator. The selection of the swap-out path is based on whether there are available free accelerator memory resources in the current system: if other accelerator memory space is detected to be idle, the background thread will directly migrate the selected memory block to the free accelerator memory; if no such resource is available, the memory block will be moved to CPU-managed memory.
[0058] Indicatively, such as Figure 3A As shown, in traditional memory resource scheduling methods, the request accelerator needs to read pages A and B, but it detects that the memory usage of the request accelerator has reached a preset threshold. At this point, some pages currently in memory are swapped out to free up space. For example... Figure 3A As shown, the request accelerator evicts pages X and Y, which were originally residing in its memory, to CPU memory, and reads pages A and B from CPU memory to its local memory.
[0059] Indicatively, such as Figure 3B As shown in the memory resource scheduling method of this application embodiment, the request accelerator needs to read pages A and B, but it is detected that the memory occupancy rate of the request accelerator has reached a preset threshold. At this time, the idle memory of other accelerators in the system is preferentially used as an extended resource. If it is detected that the first accelerator has available memory space, the first accelerator is selected as the target accelerator, and the pages to be swapped out in the request accelerator (such as pages X and Y) are pre-evicted to the local memory of the target accelerator, instead of being directly swapped out to the CPU memory.
[0060] Indicatively, such as Figure 3BAs shown, the requesting accelerator needs to evict pages X and Y into the target accelerator's memory and retrieve the required pages A and B from the CPU memory. Based on this, the requesting accelerator evicts pages X and Y into the target accelerator's memory. Simultaneously, the target accelerator prefetches pages X and Y from the CPU and stores them in its own memory. Then, the requesting accelerator can directly read pages A and B from the target accelerator without accessing the CPU's memory. Furthermore, the target accelerator can back up the data of pages X and Y to the CPU's memory so that when it needs memory, it can restore its own memory space by releasing the marked pages X and Y, thus ensuring the data security of pages X and Y.
[0061] Once a memory block is successfully swapped out to free accelerator memory, its corresponding metadata management structure will be removed from the original LRU list and registered in a swapped-out list specifically used to track swapped-out memory blocks. Since the memory block currently only exists within the free accelerator memory as temporary storage, to ensure data access consistency and reliability, the memory manager prohibits the free accelerator from unilaterally reclaiming this memory space at this stage. This ensures that the original requesting accelerator can still correctly access this swapped-out data through the memory management mechanism.
[0062] For example, when a memory block is successfully swapped out from the request accelerator to the first accelerator's memory, its metadata in the request accelerator's LRU list is removed and added to a swapped-out list specifically used to track secondary storage data. Since this data currently only exists in the first accelerator's memory, to ensure that the request accelerator can still access this data correctly, the memory manager will temporarily prevent the first accelerator from reclaiming or overwriting this memory area, even if the first accelerator itself subsequently experiences memory pressure.
[0063] By migrating the metadata of swapped-out memory blocks from the local LRU list to a globally managed swapped-out list, cross-device memory address mapping tracking capabilities are built, enabling requesting accelerators to still access data migrated to the target accelerator's (e.g., the first accelerator) memory through a unified virtual address space. By temporarily prohibiting the target accelerator (the first accelerator) from unilaterally reclaiming borrowed memory space, data still relied upon by the original accelerator is prevented from being accidentally overwritten or released due to the target accelerator's own memory needs, ensuring data persistence and reliability in cross-device memory sharing scenarios. Decoupling the lifecycle state of memory blocks from their physical storage location establishes a controllable state transition foundation for subsequent multi-level data write-back (e.g., from the first accelerator to CPU memory) and memory space reclamation, thereby supporting more complex dynamic memory resource scheduling strategies.
[0064] After the initial swapping out of the memory to the free accelerator memory is completed, the memory manager will further start an asynchronous write-back thread to gradually copy the memory block copies temporarily stored in the free accelerator memory to the CPU system memory. After the copying is completed, the corresponding area in the free accelerator memory is marked as removable and its metadata is transferred to the removable list.
[0065] According to an embodiment of this application, the data to be swapped out is backed up to the memory space of the central processing unit and marked as removable; in response to a memory release instruction for the memory space of the target accelerator, the memory space occupied by the data marked as removable is released in the memory space of the target accelerator.
[0066] When the requesting accelerator needs memory space, the data to be swapped out is migrated to the target accelerator's memory. When the target accelerator subsequently needs to reclaim memory, to avoid directly discarding this swapped-out data, the data to be swapped out is backed up to the CPU's memory space and marked as removable. Simultaneously, within the target accelerator, the memory region corresponding to this data is marked as removable. Thus, while the data swapped out from the requesting accelerator is being migrated to the target accelerator, a complete backup is also performed in the CPU's memory.
[0067] When the target accelerator itself needs more available memory, in response to a memory release command for the target accelerator, the memory space occupied by data marked as removable in the target accelerator is released. Since the data is already backed up in CPU memory, the reclamation operation will not result in data loss.
[0068] By introducing CPU memory as a secondary backup storage, when an idle accelerator needs to reclaim memory resources, it can prioritize releasing these memory regions marked as removable, thus promptly returning the memory space to its original accelerator for its own tasks. This achieves dynamic and secure reclamation of memory resources among accelerators. Data backup ensures the security and recoverability of data swapped out by requesting accelerators, thereby improving the liquidity and overall utilization of memory resources in multi-accelerator systems while ensuring data integrity, and enhancing the practicality of memory management methods.
[0069] It's important to note that before the write-back thread completes migrating data from idle accelerator memory to CPU memory, the occupied idle accelerator memory space cannot be reclaimed to serve memory expansion requests from other accelerators. Therefore, the data throughput efficiency of the write-back thread becomes a key factor affecting overall memory resource turnover. To improve write-back performance, the memory manager employs a big-page transfer mode in 4MB units to increase the data size of a single swap-out operation, thereby reducing the number of transfers and protocol overhead, speeding up the execution of the write-back thread, and ultimately improving the resource scheduling efficiency and task responsiveness of the entire hierarchical memory management system.
[0070] According to an embodiment of this application, when there are multiple idle accelerators, determining a target accelerator for at least one accelerator from among the idle accelerators includes: sorting the multiple idle accelerators according to the idle memory size of the idle accelerators to obtain a first sorting result; sorting the multiple idle accelerators according to the link distance of the idle accelerators relative to at least one accelerator to obtain a second sorting result; weighting the first sorting result and the second sorting result according to a preset memory weight and a preset distance weight to obtain a comprehensive sorting result; and determining the target accelerator from among the multiple idle accelerators according to the comprehensive sorting result.
[0071] Idle memory size refers to the maximum contiguous or aggregated memory capacity currently available for each candidate accelerator. The first ranking result is a sequence of candidate accelerators arranged from largest to smallest or smallest to largest in terms of idle memory capacity. After the first ranking, accelerators with larger idle memory space can be prioritized to ensure that the data swapping requirements of the requested accelerator can be met at once, reducing the overhead of multiple migrations due to insufficient capacity. The second ranking result is a sequence of candidate accelerators arranged from nearest to farthest or farthest to nearest in terms of link distance. After the second ranking, candidate accelerators with shorter communication paths and lower latency can be prioritized to minimize data migration time and subsequent remote access latency. The preset memory weight refers to the priority coefficient assigned to the dimension of memory capacity. The preset distance weight refers to the priority coefficient assigned to the dimension of link distance. The weighted processing calculates the linear combination of the position or score of each candidate accelerator in the two rankings according to the corresponding weights, referring to formula (1).
[0072] m = a × memory accelerator sequence + b × distance accelerator sequence (1);
[0073] Where m is the overall ranking result of any idle accelerator, a is the preset memory weight, and b is the preset distance weight.
[0074] The overall ranking result reflects the combined advantages and disadvantages of each candidate accelerator in terms of capacity provision and communication performance. Configurable weighting coefficients allow for flexible adjustments to decisions based on different application scenarios. For example, different weighting coefficients can be selected to calculate the overall ranking result for memory-intensive or latency-sensitive tasks. The target accelerator is typically the idle accelerator that ranks highest or has the highest overall score in the overall ranking result. Among multiple feasible options, the accelerator that achieves the best balance between memory capacity and communication performance is selected as the remote memory expansion target, thereby optimizing the overall performance of data migration and subsequent access. When multiple accelerators are requesting memory, there may be situations where multiple accelerators simultaneously request idle memory corresponding to the same target accelerator. In this case, memory is allocated to each accelerator sequentially according to the order of their requests.
[0075] When all other accelerators within the server have no available free memory, data is swapped out to the local CPU memory directly connected to the requesting accelerator via the PCIe bus, serving as the final memory extension layer. The data transfer rate is limited by the bandwidth limit of the PCIe bus, but by confining the swapping to local CPU memory, the higher latency and protocol overhead associated with accessing remote CPU memory in a more complex, inconsistent memory access architecture are avoided. This ensures functionality even in extreme cases while keeping performance degradation within a known and relatively controllable range.
[0076] By simultaneously considering the idle memory capacity of candidate accelerators and their link distance to the requesting accelerator, and comprehensively ranking them according to configurable weighting coefficients, the target node that achieves the optimal balance between memory supply capacity and communication performance can be selected from multiple available idle accelerators. This avoids the high latency access problem that may result from selecting solely based on capacity, and also prevents the risk of insufficient memory capacity caused by selecting solely based on distance. This enables more refined resource matching and more efficient data migration, which is beneficial to improving the overall memory utilization and task execution performance of multi-accelerator clusters.
[0077] According to an embodiment of this application, when at least one accelerator triggers a page fault, the page fault is recorded in a processing queue; when the processing queue meets the page fault handling conditions, the page faults in the processing queue are processed in batches.
[0078] A page fault is triggered when an accelerator (such as a request accelerator) encounters a problem accessing a virtual address and finds that the required data is not in its local or mapped physical memory. The system does not immediately process the error but instead adds its relevant information (such as the fault address, process ID, etc.) to a processing queue for temporary storage. Specific triggering conditions are set, such as the queue length reaching a threshold, a fixed time interval elapsed, or an idle period. When the conditions are met, multiple page faults accumulated in the queue are processed all at once. Batch processing typically involves operations such as allocating physical memory, loading the required data from CPU memory or other accelerator memory, and updating page tables.
[0079] For example, a multi-graphics processing unit (GPU) server is running a deep learning training task, responsible for processing several layers of a large model.
[0080] During training iterations, the GPU needs to access model parameter blocks P1 and P2, but neither of these blocks is currently loaded into its memory. The GPU triggers two consecutive page faults (corresponding to P1 and P2 respectively). Therefore, these two error messages are recorded sequentially in the GPU's corresponding processing queue, which has a length of 2. The system's batch processing condition is a queue length ≥ 3 or processing once every 5 milliseconds. When the GPU subsequently triggers a third page fault due to accessing parameter P3, the queue length reaches 3, satisfying the immediate processing condition. Alternatively, if the length threshold is not reached before the 5-millisecond timer expires, the process is triggered based on the time condition. The three page faults (P1, P2, P3) in the queue are processed at once. The processing may include: reading the three parameter blocks from CPU memory (or other GPU memory storing P1, P2, and P3) into the GPU's local memory through one or a few large-capacity transfers; allocating physical pages for these three memory blocks and updating the GPU's page table; and clearing the processing queue after processing.
[0081] By temporarily storing triggered page faults in a processing queue and batch processing them when conditions are met, the overhead of multiple discrete page fault responses can be merged into a single centralized processing flow, thereby reducing the cumulative cost of system scheduling and management.
[0082] Figure 4 A schematic diagram of a queue to be processed according to an embodiment of this application is shown.
[0083] like Figure 4 As shown, batch processing of page faults in the queue to be processed includes: controlling at least one accelerator and a target accelerator to process page faults in parallel, starting from the head and tail of the queue to be processed, respectively.
[0084] The request accelerator, which issues a memory access request and triggers a page fault, retrieves and processes page faults sequentially from the head of the processing queue. Simultaneously, the target accelerator, which provides free memory resources, retrieves and processes page faults in reverse order from the tail of the processing queue. Thus, the request accelerator and the target accelerator process their respective page fault entries independently, achieving parallel processing.
[0085] In parallel processing, mutex locks are used to synchronize access to and protect the state of the currently processed page, preventing data races, inconsistencies, or memory access conflicts caused by multiple processing threads simultaneously modifying the metadata or content of the same page. Specifically, when any accelerator's processing thread begins loading or modifying a particular page, it first acquires the mutex lock associated with that page. This ensures that other threads cannot perform write or inconsistent read operations on that page until the thread completes its operation (such as data transfer or state update) and releases the lock. This guarantees the correctness, order, and integrity of memory operations in a parallel environment, preventing data corruption or system errors caused by concurrent access, while minimizing the impact on parallel performance through fine-grained lock design.
[0086] For example, such as Figure 4 As shown, the request accelerator triggers multiple page faults consecutively, which are recorded sequentially in a processing queue. The queue order is: [10, 11, 12, 13, 14, 15, 16, ..., N] (10 is the head, N is the tail). The request accelerator is controlled to process from the head, with its task sequence being: 10, 11, 12... Simultaneously, the first accelerator is controlled to process from the tail, with its task sequence being: N, ..., 16.
[0087] By enabling the request accelerator and the target accelerator to process page faults in parallel from the head and tail of the processing queue, respectively, batch processing tasks can be split into bidirectional parallel task flows. This fully utilizes the computation and input / output (I / O) capabilities of multiple accelerators, improving the throughput of page fault resolution and data loading. At the same time, the request accelerator and the target accelerator can simultaneously call their respective computing units, distributing the I / O pressure originally concentrated on one accelerator to two physical devices, thus achieving load balancing between the two accelerators.
[0088] Figure 5 A schematic diagram illustrating the processing of page faults according to an embodiment of this application is shown.
[0089] Parallel processing of page faults includes: controlling the processing thread of at least one accelerator to retrieve page faults from the head of the queue to be processed and loading the data corresponding to the page faults into the memory space of at least one accelerator; and controlling the processing thread of the target accelerator to retrieve page faults from the tail of the queue to be processed and loading the data corresponding to the page faults into the memory space of the target accelerator.
[0090] like Figure 5 As shown, when the request accelerator triggers page faults for pages A and B, the request accelerator calls its processing thread to sequentially retrieve page faults for page A from the head of the processing queue. Simultaneously, the target accelerator calls its processing thread to sequentially retrieve page faults for page B from the tail of the processing queue; that is, the request accelerator and the target accelerator process page faults for pages A and B in parallel. If both pages A and B are stored in CPU memory, the request accelerator retrieves page A from CPU memory and moves it to its own memory, while the target accelerator retrieves page B from CPU memory and moves it to its own memory. Specifically, the request accelerator can directly retrieve page A from CPU memory, and simultaneously, it can directly retrieve page B from CPU memory. At the same time, the request accelerator can swap out its infrequently used data pages X and Y from its memory to the target accelerator to occupy free memory in the target accelerator. The target accelerator then synchronously backs up pages X and Y to CPU memory.
[0091] Ultimately, after processing, the requesting accelerator only needs the data of the currently used page A. The target accelerator stores its own stored pages E, X, and Y, as well as the retrieved page B. The CPU memory stores the originally stored pages C and D, and the backed-up pages X and Y. Thus, the missing data in the first half of the processing queue (such as page A) is already in the requesting accelerator's memory and can be used immediately. The data in the second half of the processing queue (such as page B) is located in the target accelerator's (first accelerator's) memory, so that when the requesting accelerator needs to prefetch pages B, X, and Y, it can directly retrieve them from the target accelerator's (such as the first accelerator's) local memory via a high-speed interconnect protocol, without having to access the CPU memory through the lower bandwidth and higher latency PCIe bus. This effectively improves the response speed of data access and the overall I / O efficiency of the system, while reducing the pressure on CPU memory bandwidth, and realizing distributed optimization and load balancing of memory resources in a multi-accelerator environment.
[0092] By enabling the request accelerator and the target accelerator to perform data loading in parallel from the head and tail of the processing queue, respectively, the single sequential data transfer task is split into two concurrent sub-task streams. This fully utilizes the I / O capabilities and bus bandwidth of the dual accelerators, reducing the total time spent on batch page fault processing. At the same time, by temporarily storing some data locally on the target accelerator, the request accelerator can subsequently access it remotely via a high-speed interconnect protocol. This avoids PCIe bandwidth contention and reduces access latency, while improving response efficiency in multi-accelerator page fault scenarios.
[0093] According to an embodiment of this application, after page fault processing in the queue to be processed is completed, for at least one page fault, data to be prefetched corresponding to at least one page fault is determined; at least one accelerator and a target accelerator are controlled to prefetch the data to be prefetched in parallel.
[0094] After all accumulated and triggered page faults have been processed, the recently processed page fault is analyzed, and related data that may be accessed subsequently is inferred based on predefined rules. These predefined rules can be sequential prefetching or access pattern prediction based on machine learning models. From this, adjacent data blocks that may be accessed next are inferred. The inferred data is the data to be prefetched. This data may be stored in CPU memory, in the free manager's memory, or a combination thereof.
[0095] The request accelerator can prefetch some or all of the data to be prefetched from the memory corresponding to the target accelerator, or it can prefetch some or all of the data to be prefetched from the memory corresponding to the CPU. In the case where the request accelerator and the target accelerator prefetch data in parallel, the request accelerator can prefetch the data stored in the target accelerator's memory, while the target accelerator can prefetch the data stored in the CPU's memory. By initiating parallel data prefetching after completing the current page fault handling, potentially accessed data can be loaded into the local memory of both the request accelerator and the target accelerator in advance, thereby mitigating potential page fault issues and improving memory access latency in subsequent computations.
[0096] According to an embodiment of this application, determining the data to be prefetched corresponding to at least one page fault includes: determining the data access direction based on the analysis results obtained by analyzing the address access sequence before triggering at least one page fault; selecting a data block of a preset data size along the data access direction starting from the address of at least one page fault, and determining the data corresponding to the data block as the data to be prefetched.
[0097] An address access sequence refers to a historical record of a series of memory addresses accessed by the accelerator within a certain period prior to triggering the current page fault. The data access direction is a signed offset used to indicate the primary direction and step size of the address changes. For example, analysis might reveal an address sequence following a pattern of +4096 bytes, +4096 bytes, +4096 bytes… in which case the access direction is determined to be forward, with a step size of +4096 bytes. Alternatively, a pattern of -8192 bytes, -8192 bytes… might be found, in which case the direction is reversed, with a step size of -8192 bytes.
[0098] The starting point refers to the memory address where a page fault was triggered and has been handled. Starting from this starting point, one or more addresses that may be accessed in the future are calculated according to the access direction and step size determined in the previous step. The preset data size determines how many units of data are prefetched. For example, if the prefetch depth is set to 2, then starting from the starting point, the next two data blocks are prefetched along the direction. The data blocks corresponding to the calculated future addresses are marked as objects that need to be prefetched, i.e., the data to be prefetched.
[0099] By analyzing the address access sequence before a page fault is triggered, the data access direction and step size are determined. Based on this, subsequent data blocks are prefetched from the page fault address along the predicted direction. This allows the prefetching behavior to closely match the actual memory access pattern of the application during runtime, thereby improving the accuracy and timeliness of the prefetched data.
[0100] Figure 6A A schematic diagram of parallel prefetching of data to be prefetched according to an embodiment of this application is shown; Figure 6B A schematic diagram of parallel prefetching of data to be prefetched is shown according to another embodiment of this application.
[0101] According to an embodiment of this application, parallel prefetching of data to be prefetched includes: controlling a prefetching thread of at least one accelerator to prefetch a first sub-data located in the memory space of the target accelerator from the data to be prefetched to the memory space of the at least one accelerator; and controlling a prefetching thread of the target accelerator to prefetch a second sub-data located in the memory space of the central processing unit from the data to be prefetched to the memory space of the target accelerator.
[0102] The first sub-data refers to the data to be prefetched that already exists in the memory of the target accelerator (such as the first accelerator), for example... Figure 6A Page E is shown in the image. The first sub-data may be temporarily stored in the target accelerator's memory during a previous swap-out operation, or it may be data loaded into the target accelerator's memory during a previous prefetching phase. The second sub-data refers to the data to be prefetched that currently only exists in CPU memory, for example... Figure 6A Pages C and D are shown in the image.
[0103] like Figure 6AAs shown, the prefetch thread of the requesting accelerator is responsible for prefetching page E from the target accelerator's memory to its local memory. The prefetch thread of the target accelerator (such as the first accelerator) is responsible for prefetching pages C and D from the CPU memory to the target accelerator's local memory.
[0104] The prefetch threads of both accelerators start simultaneously, each processing its assigned portion of data. For the requesting accelerator, accessing the target accelerator's memory typically occurs via a high-speed interconnect protocol, a path with higher bandwidth and lower latency than the PCIe path used to access CPU memory. Therefore, the requesting accelerator experiences faster data transfer speeds when fetching the first piece of data from the target accelerator. For the target accelerator, loading the second piece of data from CPU memory uses the PCIe path. This effectively creates a high-speed cache for the requesting accelerator; if it needs this data later, it can directly retrieve it from the target accelerator's memory via the high-speed interconnect protocol, bypassing the slower CPU memory access. After prefetching, the requesting accelerator locally possesses the first piece of data; while the second piece of data is cached in the target accelerator's memory, thus creating a data distribution that provides a backup for the requesting accelerator.
[0105] Furthermore, in Figure 6A Based on this, the requesting accelerator obtains pages A, B, and E through swapping and prefetching steps, while the target accelerator obtains pages X, Y, C, and D. Figure 6B Furthermore, if the requesting accelerator needs to prefetch pages C and D, it will directly prefetch pages C and D from the target accelerator. At the same time, the target accelerator can also prefetch page F from CPU memory into the target accelerator's memory in parallel for use by the requesting accelerator.
[0106] Based on the above memory resource scheduling method, this application also provides a memory resource scheduling device. The following will be combined with... Figure 7 The device is described in detail.
[0107] Figure 7 A structural block diagram of a memory resource scheduling apparatus according to an embodiment of this application is shown.
[0108] like Figure 7 As shown, the memory resource scheduling device 700 of this embodiment includes a first determining module 710, a second determining module 720, and a migration module 730.
[0109] The first determining module 710 is configured to, for at least one accelerator in a multi-accelerator system, determine at least one idle accelerator from the other accelerators in the multi-accelerator system, based on the memory status information of each accelerator in the multi-accelerator system when the memory occupancy rate of at least one accelerator is detected to reach a preset threshold. The idle accelerator is defined as an accelerator with available idle memory. In one embodiment, the first determining module 710 may be used to execute the operation S210 described above, which will not be repeated here.
[0110] The second determining module 720 is configured to determine a target accelerator for at least one accelerator from among the idle accelerators based on the idle memory size of the idle accelerators and the link distance of the idle accelerators relative to at least one accelerator. In one embodiment, the second determining module 720 may be used to perform the operation S220 described above, which will not be repeated here.
[0111] The migration module 730 is used to migrate data to be swapped out from at least one accelerator to the memory space of a target accelerator. In one embodiment, the migration module 730 can be used to perform the operation S230 described above, which will not be repeated here.
[0112] According to an embodiment of this application, when there are multiple idle accelerators, the second determination module 720 includes: a first sorting submodule, a second sorting submodule, a weighted processing submodule, and a first determination submodule.
[0113] The first sorting submodule is used to sort multiple idle accelerators according to the idle memory size of the idle accelerators to obtain a first sorting result; the second sorting submodule is used to sort multiple idle accelerators according to the link distance of the idle accelerators relative to at least one accelerator to obtain a second sorting result; the weighted processing submodule is used to perform weighted processing on the first sorting result and the second sorting result according to preset memory weight and preset distance weight to obtain a comprehensive sorting result; the first determination submodule is used to determine the target accelerator from the multiple idle accelerators according to the comprehensive sorting result.
[0114] According to an embodiment of this application, the memory resource scheduling method further includes: a backup module and a release module.
[0115] The backup module is used to back up the data to be swapped out to the memory space of the central processing unit and mark the data to be swapped out as removable; the release module is used to release the memory space occupied by the data marked as removable in the memory space of the target accelerator in response to the memory release command for the memory space of the target accelerator.
[0116] According to embodiments of this application, the memory resource scheduling method further includes a recording module and a batch processing module.
[0117] The recording module is used to record page faults to the processing queue when at least one accelerometer triggers a page fault. The batch processing module is used to process page faults in the processing queue in batches when the processing queue meets the page fault handling conditions.
[0118] According to an embodiment of this application, the batch processing module includes: a first control submodule.
[0119] The first control submodule is used to control at least one accelerator and a target accelerator to process page faults in parallel, starting from the head and tail of the queue to be processed, respectively.
[0120] According to an embodiment of this application, the first control submodule includes: a first control unit and a second control unit.
[0121] The first control unit is used to control the processing thread of at least one accelerator to obtain page faults from the head of the queue to be processed and load the data corresponding to the page faults into the memory space of at least one accelerator; the second control unit is used to control the processing thread of the target accelerator to obtain page faults from the tail of the queue to be processed and load the data corresponding to the page faults into the memory space of the target accelerator.
[0122] According to an embodiment of this application, the memory resource scheduling method further includes: a third determining module and a control module.
[0123] The third determination module is used to determine the data to be prefetched corresponding to at least one page fault, given that page fault handling in the queue has been completed; the control module is used to control at least one accelerator and the target accelerator to prefetch the data to be prefetched in parallel.
[0124] According to an embodiment of this application, the third determining module includes: a second determining submodule and a selecting submodule.
[0125] The second determining submodule is used to determine the data access direction based on the analysis results obtained from analyzing the address access sequence before triggering at least one page fault; the selecting submodule is used to select a data block of a preset amount of data along the data access direction, starting from the address of at least one page fault, and determine the data corresponding to the data block as the data to be prefetched.
[0126] According to an embodiment of this application, the control module includes a second control submodule and a third control submodule.
[0127] The second control submodule is used to control the prefetching thread of at least one accelerator to prefetch the first sub-data located in the memory space of the target accelerator from the data to be prefetched to the memory space of at least one accelerator; the third control submodule is used to control the prefetching thread of the target accelerator to prefetch the second sub-data located in the memory space of the central processing unit from the data to be prefetched to the memory space of the target accelerator.
[0128] According to embodiments of this application, any plurality of modules among the first determining module 710, the second determining module 720, and the migration module 730 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules can be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this application, at least one of the first determining module 710, the second determining module 720, and the migration module 730 can be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented by any other reasonable means of integrating or packaging the circuit, or implemented in any one of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, at least one of the first determining module 710, the second determining module 720, and the migration module 730 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0129] Figure 8 A block diagram of an electronic device suitable for implementing a memory resource scheduling method according to an embodiment of this application is shown.
[0130] like Figure 8 As shown, an electronic device 800 according to an embodiment of this application includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage portion 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of this application.
[0131] RAM 803 stores various programs and data required for the operation of electronic device 800. Processor 801, ROM 802, and RAM 803 are interconnected via bus 804. Processor 801 executes various operations of the method flow according to embodiments of this application by executing programs in ROM 802 and / or RAM 803. It should be noted that the programs may also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 may also execute various operations of the method flow according to embodiments of this application by executing programs stored in said one or more memories.
[0132] According to embodiments of this application, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to a bus 804. The electronic device 800 may also include one or more of the following components connected to the input / output (I / O) interface 805: an input section 806 including a keyboard, mouse, etc.; an output section 807 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 808 including a hard disk, etc.; and a communication section 809 including a network interface card such as a LAN card, modem, etc. The communication section 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the input / output (I / O) interface 805 as needed. A removable medium 811, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 810 as needed so that computer programs read from it can be installed into the storage section 808 as needed.
[0133] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0134] According to embodiments of this application, the computer-readable storage medium can be a non-volatile computer-readable storage medium, such as including but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this application, the computer-readable storage medium may include ROM 802 and / or RAM 803 and / or one or more memories other than ROM 802 and RAM 803 described above.
[0135] Embodiments of this application also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code is used to cause the computer system to implement the methods provided in the embodiments of this application.
[0136] When the computer program is executed by the processor 801, it performs the functions defined in the system / apparatus of this application embodiment. According to the embodiments of this application, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0137] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 809, and / or installed from a removable medium 811. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0138] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, it performs the functions defined in the system of this application embodiment. According to the embodiments of this application, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0139] According to embodiments of this application, program code for executing the computer programs provided in the embodiments of this application can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0140] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0141] Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application.
[0142] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. A memory resource scheduling method applied to a multi-accelerator system, characterized in that, The method includes: For at least one accelerator in the multi-accelerator system, when the memory occupancy rate of the at least one accelerator is detected to reach a preset threshold, at least one idle accelerator is determined from the other accelerators in the multi-accelerator system other than the at least one accelerator, based on the memory status information of each accelerator in the multi-accelerator system. The idle accelerator is an accelerator with idle memory. Based on the idle memory size of the idle accelerators and the link distance of the idle accelerators relative to the at least one accelerator, a target accelerator for the at least one accelerator is determined from the idle accelerators; and Migrate the data to be swapped out from the at least one accelerator to the memory space of the target accelerator; In the absence of a target accelerator, migrate the data to be swapped out from at least one accelerator to the memory space of the central processing unit. The method further includes: If a page fault is detected by the at least one accelerator, the page fault is recorded in the processing queue. If the processing queue meets the page fault handling conditions, the processing thread of the at least one accelerator is controlled to retrieve the page fault from the head of the processing queue and load the data corresponding to the page fault into the memory space of the at least one accelerator for immediate use by the at least one accelerator. The processing thread of the target accelerator is controlled to obtain page faults from the tail of the queue to be processed, and load the data corresponding to the page faults into the memory space of the target accelerator so that the at least one accelerator can directly obtain the data from the local memory of the target accelerator through a high-speed interconnect protocol. Control the prefetching thread of the at least one accelerator to prefetch the first sub-data located in the memory space of the target accelerator from the data to be prefetched into the memory space of the at least one accelerator; Control the prefetch thread of the target accelerator to prefetch the second sub-data located in the memory space of the central processing unit from the data to be prefetched to the memory space of the target accelerator, so that when the at least one accelerator needs the second sub-data, it can be obtained directly from the memory space of the target accelerator through a high-speed interconnect protocol; The data to be swapped out is backed up to the memory space of the central processing unit, and the data to be swapped out is marked as removable. In response to a memory release command for the target accelerator, the memory space occupied by the data marked as removable in the memory space of the target accelerator is released.
2. The method according to claim 1, characterized in that, In the case where there are multiple idle accelerators, determining the target accelerator for the at least one accelerator from the idle accelerators includes: Based on the idle memory size of the idle accelerators, the multiple idle accelerators are sorted to obtain a first sorting result; Based on the link distance of the idle accelerator relative to the at least one accelerator, the multiple idle accelerators are sorted to obtain a second sorting result; The first and second sorting results are weighted according to the preset memory weight and preset distance weight to obtain the comprehensive sorting result. Based on the comprehensive ranking results, the target accelerator is determined from the plurality of idle accelerators.
3. The method according to claim 1, characterized in that, Also includes: Once the page fault handling in the queue is completed, for at least one page fault, determine the data to be prefetched corresponding to the at least one page fault; Control the at least one accelerator and the target accelerator to prefetch the data to be prefetched in parallel.
4. The method according to claim 3, characterized in that, The determination of the prefetch data corresponding to the at least one page fault includes: Based on the analysis results obtained from analyzing the address access sequence prior to triggering the at least one page fault, the data access direction is determined; Starting from the address of the at least one page fault, a data block of a preset amount of data is selected along the data access direction, and the data corresponding to the data block is determined as the data to be prefetched.
5. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Accelerator, memory management method for accelerator and data processing system
CN106959893A
Data processing system, method and controller
CN115576661A
Memory management method and device and computing equipment
CN119537267A