Dynamic management method for device side memory and host memory
Patent Information
- Application Number
- CN202510468453.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
Smart Images

Figure CN120336023A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of memory allocation management, and more specifically, to a method for dynamically managing device-side memory and host memory. Background Art
[0002] Currently, with the rise of artificial intelligence technology, the computation of a huge amount of data has imposed certain pressure on the size and bandwidth of device-side memory; at the same time, as more and more AI chips integrate hybrid IP modules, such as multiple computing units like RVV, GPU, NPU, VPU, etc., software calls multiple IPs to run in parallel, involving a large amount of globally shared data. As the time scope of the data extends, the long-term occupation of memory space will cause the phenomenon of insufficient memory in the NPU during peak computing. Especially with the current popularity of large models, huge numbers of parameters participate in the computation, which puts forward higher requirements for the size and bandwidth of on-chip memory.
[0003] Generally, the general processing method is to simply use a pcie module in the NPU acceleration card, connect to the host through pcie, map the address bus of the chip device to the IO address range of the host, and the host driver manages the device-side memory through the fixed mapping method of ioremap, without the characteristics of dynamic management.
[0004] Device-side memory and host memory are two different media. For the NPU computing unit, there are obvious differences in the data reading and writing speeds. Generally, device-side memory mostly uses GDDR, LPDDR5, HBM or SEDRAM, with higher access bandwidth. How to adopt a simple, efficient, general, and easily dynamically extensible method to optimize the management of these two memory media is a problem to be solved. Summary of the Invention
[0005] Aiming at the technical problem of limited device-side memory in the prior art, the present invention provides a method for dynamically managing device-side memory and host memory, which can solve the problem of insufficient device-side memory when facing big data processing.
[0006] The present invention provides a method for dynamically managing device-side memory and host memory, including:
[0007] Maintain an LRU double-linked hash list in the host program, and maintain corresponding configuration information for each node in the LRU double-linked hash list. The configuration information includes the key value of the node, the valid flag, and the physical memory space of the device-side memory mapped by the node data. The valid flag is used to mark that the node data is located in the device-side memory or in the host memory;
[0008] When there is insufficient device - side memory space for a new application, traverse forward from the tail end of the LRU doubly - linked hash list in sequence, and move the node data located in the device - side memory to the host memory until there is enough device - side memory space to meet the size requirement of the device - side memory space to be applied for.
[0009] Based on the above - mentioned technical solution, the present invention can also be improved as follows.
[0010] Optionally, in the configuration information, the flag valid = 1 indicates that the node data is located in the device - side memory, and valid = 0 indicates that the node data is located in the host memory. The physical memory space of the device - side memory mapped by the node data includes the starting physical address and data size of the node data.
[0011] Optionally, when there is insufficient device - side memory space for a new application, traverse from the tail end of the LRU doubly - linked hash list in sequence, and move the node data located in the device - side memory to the host memory until there is enough device - side memory space to meet the size requirement of the device - side memory space to be applied for, including:
[0012] When the host program receives a request for device - side memory space, it determines whether the device - side memory space is sufficient. If it is sufficient, map a section of device - side memory space, create a node, maintain the configuration information of the node, and then place the created node at the head of the LRU doubly - linked hash list;
[0013] If the device - side memory space is insufficient, through the LRU mechanism, traverse forward in sequence from the tail end of the LRU doubly - linked hash list for nodes with valid = 1;
[0014] Allocate the same - sized host memory, move the device - side memory data of this node to the host memory through the DMA function, and set the valid value of this node to 0;
[0015] Refresh the memory space of this node until enough device - side memory space is reserved to meet the size requirement of the device - side memory space to be applied for;
[0016] Map the corresponding device - side memory, set the memory space parameters for this node, set a key value for this node and set the valid of this node to 1, and move this node to the head of the LRU doubly - linked hash list.
[0017] Optionally, it further includes:
[0018] When the data that has been moved to the host memory is required by the computing unit again, determine whether the device - side memory space is sufficient;
[0019] If the device-side memory space is insufficient, the LRU bidirectional hash linked list is traversed from the tail end to the front, and the node data in the device-side memory is moved to the host memory until sufficient device-side memory space is reserved to meet the space size requirement for copying.
[0020] Optionally, traverse the LRU bidirectional hash linked list from the tail end to the front, and move the node data in the device-side memory to the host memory until enough device-side memory space is reserved to meet the space size requirement to be copied, including:
[0021] According to the LRU mechanism, the valid of the node with a valid value of 1 is set to 0 from the end of the LRU bidirectional hash linked list in turn, and the device-side memory space data of the node is transferred to the host memory through the DMA function, and the memory space parameter corresponding to the node is replaced until enough device-side memory space is reserved to meet the space size requirement to be copied;
[0022] Map the corresponding device-side memory, copy the data required by the computing unit to the device-side memory, refresh the memory space parameters corresponding to the node, reset the valid of the node to 1, and place the node at the head of the LRU bidirectional hash linked list.
[0023] Optionally, if the device-side memory space is sufficient, the device-side memory of the required size is mapped out, a node is maintained, and the newly required host memory data is moved to the device-side memory space of the node through the DMA function, the host memory is released, and the valid value of the node is set to 1;
[0024] The node is mentioned at the head of the LRU bidirectional hash linked list, and the starting physical address of the node data is sent to the device end, so that the device end processes the node data.
[0025] Optionally, also include:
[0026] When the applied memory space is used up, the device-side memory is unmapped and the corresponding node on the LRU bidirectional hash linked list is deleted to release the device-side memory space and the host memory space.
[0027] A dynamic management method for device - side memory and host memory provided by the present invention maintains an LRU doubly - linked hash list in the host program, and maintains corresponding configuration information for each node. The configuration information includes the key value of the node, the valid flag, and the physical memory space of the device - side memory mapped by the node data. The valid flag is used to mark whether the node data is located in the device - side memory or in the host memory. When there is insufficient space for a newly applied device - side memory, traverse from the end to the front of the LRU doubly - linked hash list in turn, and move the node data located in the device - side memory to the host memory until there is enough device - side memory space to meet the size requirement of the newly applied device - side memory space. When the device - side memory is insufficient, the present invention opens up the host memory as a backup storage and proposes an LRU management mechanism to dynamically manage the allocation of the device - side memory and the host memory, solving the problem of limited device - side memory resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 FIG. is a flowchart of a dynamic management method for device - side memory and host memory provided by an embodiment of the present invention;
[0029] Figure 2 FIG. is a schematic diagram of the configuration information of each node of the LRU doubly - linked hash list according to an embodiment of the present invention;
[0030] Figure 3 FIG. is a schematic diagram of the LRU doubly - linked hash list according to an embodiment of the present invention;
[0031] Figure 4 FIG. is a flowchart of applying for a new device - side memory space according to an embodiment of the present invention;
[0032] Figure 5 FIG. is a flowchart of the processing when the data that has been moved to the host memory is required by the computing unit again according to an embodiment of the present invention;
[0033] Figure 6 FIG. is a flowchart of destroying memory data according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention. In addition, the technical features in various embodiments or individual embodiments provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. Such combination is not restricted by the order of steps and / or the pattern of structural composition, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be achieved, it should be considered that such combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0035] In many existing application scenarios, when the large model is reasoning, the cache of KV values causes data inflation after long-term operation, which will inevitably lead to insufficient memory on the device side. In view of the limited memory resources on the device side in the prior art, the present invention allocates host memory as backup storage. The device-side memory belongs to the near memory of the computing unit and has fast access and calculation speeds; the host memory is regarded as the backup storage of the computing unit. When the memory resources on the device side are insufficient, the data that has not been used for a long time is transferred to the host memory with a larger capacity and lower cost, so as to free up enough device-side memory space to give priority to high-performance computing.
[0036] The present invention proposes an LRU management mechanism to dynamically manage the device-side memory and the host memory. Among them, the core of using the LRU algorithm is based on the assumption that the data recently used is more likely to be used again in the future, while the data that has not been used for a long time is less likely to be used in the future. The data that is frequently used is called hot data, and the data that is not used for a long time is called cold data.
[0037] Figure 1 The flowchart of a method for dynamically managing the device-side memory and the host memory provided by the present invention is as Figure 1 shown, and the method includes:
[0038] Step 1, maintain an LRU doubly linked hash list in the host program, and maintain corresponding configuration information for each node in the LRU doubly linked hash list. The configuration information includes the key value of the node, the valid flag, and the physical memory space of the device-side memory mapped by the node data. The valid flag is used to mark whether the node data is located in the device-side memory or in the host memory.
[0039] It is understandable that, in order to improve the utilization efficiency of memory space and address the issue of insufficient memory on the device side, a unified dynamic memory management software is developed to achieve dynamic management of host memory and device memory. The memory management strategy of the present invention includes the following strategies:
[0040] The application allocates device-side memory through the inbound of PCIE, maintains an LRU doubly linked hash list in the host program, and maintains corresponding configuration information for each node in the list. The configuration information includes the node key value, the node valid flag (valid = 1 indicates that the node data is located in the device-side memory, valid = 0 indicates that the node data is in the host memory), and the physical memory space of the device-side memory mapped by the node data (including the starting physical address of the node data and the size of the data). The data structure of each node is as Figure 2 shown.
[0041] In the configuration information of the node, the corresponding storage space is found by binding the key value, rather than directly using the virtual address. Since the key value and the node are bound one-to-one and can ensure simultaneous creation and destruction, the phenomenon of access mismatch caused by data storage space transfer is avoided.
[0042] Among them, based on the LRU management mechanism, the newly inserted node or the node reused again is placed at the head of the list. Figure 3 It is a schematic diagram of the structure of the maintained LRU doubly linked hash list.
[0043] Step 2, when the newly applied device-side memory space is insufficient, traverse from the tail end to the front of the LRU doubly linked hash list in turn, and move the node data located in the device-side memory to the host memory until there is enough device-side memory space to meet the size requirement of the device-side memory space to be applied.
[0044] See Figure 4 , in an embodiment of the present invention, when the newly applied device-side memory space is insufficient, traverse from the tail end of the LRU doubly linked hash list in turn, and move the node data located in the device-side memory to the host memory until there is enough device-side memory space to meet the size requirement of the device-side memory space to be applied, including:
[0045] When the host program receives a new application request for device-side memory space, it judges whether the device-side memory space is sufficient. If it is sufficient, map a section of device-side memory space, create a node, and after maintaining the configuration information of the node, place the created node at the head of the LRU doubly linked hash list.
[0046] If the memory space on the device side is insufficient, through the LRU mechanism, traverse the nodes with valid value of 1 from the end to the front of the LRU doubly linked hash list in turn. Allocate the same size of host memory, transfer the data in the device-side memory space of this node to the host memory through the DMA function, and set the valid value of this node to 0.
[0047] Refresh the memory space address of this node until enough device-side memory space is reserved to meet the size requirement of the device-side memory space to be applied for. Map the corresponding device-side memory, set the memory space parameters to this node, set the key value for this node, set the valid of this node to 1, and move this node to the head of the LRU doubly linked hash list.
[0048] The present invention innovatively introduces the LRU mechanism management into the device-side memory management, so that the device-side memory that has not been used for a long time and the memory cache that cannot be immediately released are cached in the host memory.
[0049] In an embodiment of the present invention, when the data that has been transferred to the host memory is required by the computing unit again, it is judged whether the device-side memory space is sufficient. If the device-side memory space is not enough, traverse from the end to the front of the LRU doubly linked hash list in turn, and transfer the node data in the device-side memory to the host memory until enough device-side memory space is reserved to meet the size requirement of the space to be copied; copy the host memory data required by the computing unit to the reserved device-side memory space.
[0050] Specifically, refer to Figure 5 , when the data that has been transferred to the host memory is required by the computing unit again and the device-side memory is insufficient, according to the LRU mechanism, set the valid of the nodes with valid value of 1 to 0 in turn from the end of the LRU doubly linked hash list, transfer the data in the device-side memory space corresponding to this node to the host memory through the DMA function, replace the memory space parameters corresponding to this node until enough device-side memory space is reserved to meet the size requirement of the space to be copied; map the corresponding device-side memory, copy the data required by the computing unit again to the device-side memory, refresh the memory space parameters corresponding to this node, reset the valid of this node to 1, and place this node at the head of the LRU doubly linked hash list.
[0051] If the device-side memory is sufficient, map the device-side memory with the required space size, maintain a node, transfer the host memory data that is needed again to the device-side memory space of this node through the DMA function, release the host memory, and set the valid value of this node to 1. Move this node to the head of the LRU doubly linked hash list, and send the starting physical address of the node data to the device side so that the device side can process the docking point data.
[0052] In one embodiment of the present invention, when the device - side memory space of the application is used up, the mapping of the device - side memory is released, and the corresponding node is deleted at the same time.
[0053] Among them, the dynamic memory management and allocation method of the present invention is applicable to the following scenarios: when large - model inference is carried out, the cache of KV values causes data expansion after long - term operation, which will inevitably lead to insufficient device - side memory. Through the solution of this patent, the KV cache is moved to a larger and cheaper main memory, and they are only retrieved back to the GPU or NPU when needed, thereby improving the context of large - model inference.
[0054] Specifically, refer to Figure 6 , when a data destruction request is sent to the application, the application indexes to the corresponding node according to the key value of the data and obtains the valid flag of the node. When the valid flag is 1, the device - side memory mapping is released. When the valid value is 0, the host memory space is released. After releasing the memory space, the corresponding node on the LRU double - linked hash list is deleted.
[0055] A dynamic management method for device - side memory and host memory provided by an embodiment of the present invention aims at the problem that when managing the NPU device - side memory in a static fixed - mapping manner currently, when the device - side memory is insufficient, only the main memory can be used to participate in data storage and calculation. Since high - performance computing has high requirements for storage media, placing data in the main memory will affect the access bandwidth and thus affect the computing performance. A set of driver software for unified memory management for the NPU chip is designed. When the device - side memory resources are insufficient, the data that has not been used for a long time is transferred to a larger and cheaper main memory, thereby freeing up space to give priority to high - performance computing and solving the problem of limited device - side memory resources.
[0056] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0057] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer - usable storage media (including but not limited to disk storage, CD - ROM, optical storage, etc.) containing computer - usable program code.
[0058] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in the Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0059] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in the Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0060] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0061] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0062] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for dynamically managing device - side memory and host memory, characterized in that, Including: Maintain an LRU doubly linked hash list in the host program, and maintain corresponding configuration information for each node in the LRU doubly linked hash list. The configuration information includes the key value of the node, the valid flag, and the physical memory space of the device-side memory mapped by the node data. The valid flag is used to mark whether the node data is located in the device-side memory or in the host memory; When the available new device-side memory space is insufficient, traverse forward from the tail end of the LRU doubly linked hash list in sequence, and move the node data located in the device-side memory to the host memory until there is enough device-side memory space to meet the size requirement of the device-side memory space to be applied for.
2. The dynamic management method for device - side memory and host memory according to claim 1, wherein In the configuration information, the flag valid = 1 indicates that the node data is located in the device-side memory, and valid = 0 indicates that the node data is located in the host memory. The physical memory space of the device-side memory mapped by the node data includes the starting physical address and the data size of the node data.
3. The dynamic management method of device - side memory and host memory according to claim 2, wherein, When the available new device-side memory space is insufficient, traverse from the tail end of the LRU doubly linked hash list in sequence, and move the node data located in the device-side memory to the host memory until there is enough device-side memory space to meet the size requirement of the device-side memory space to be applied for, including: When the host program receives a request for applying for device-side memory space, judge whether the device-side memory space is sufficient. If it is sufficient, map a section of device-side memory space, create a node, maintain the configuration information of the node, and then place the created node at the head of the LRU doubly linked hash list; If the device-side memory space is insufficient, through the LRU mechanism, traverse forward from the tail end of the LRU doubly linked hash list in sequence for the nodes with valid being 1; Allocate host memory of the same size, move the device-side memory data of this node to the host memory through the DMA function, and set the valid value of this node to 0; Refresh the memory space of this node until enough device-side memory space is reserved to meet the size requirement of the device-side memory space to be applied for; Map the corresponding device-side memory, set the memory space parameters to this node, set the key value for this node and set the valid of this node to 1, and move this node to the head of the LRU doubly linked hash list.
4. The dynamic management method for device - side memory and host memory according to claim 2, characterized in that, Also including: When the data that has been moved to the host memory is required by the computing unit again, judge whether the device-side memory space is sufficient; If the device-side memory space is not enough, traverse forward from the tail end of the LRU doubly linked hash list in sequence, and move the node data located in the device-side memory to the host memory until enough device-side memory space is reserved to meet the size requirement of the space to be copied.
5. The dynamic management method for device - side memory and host memory according to claim 4, wherein, Traverse forward from the tail end of the LRU doubly linked hash list in sequence, and move the node data located in the device-side memory to the host memory until enough device-side memory space is reserved to meet the size requirement of the space to be copied, including: According to the LRU mechanism, the valid of the node with valid value of 1 is set to 0 in turn from the tail end of the LRU double - linked hash list, and the data in the device - side memory space of this node is transferred to the host memory through the DMA function, replacing the memory space parameters corresponding to this node until enough device - side memory space is reserved to meet the space size requirement to be copied; Map the corresponding device - side memory, copy the data required by the computing unit again to the device - side memory, refresh the memory space parameters corresponding to this node, reset the valid of this node to 1, and place this node at the head of the LRU double - linked hash list.
6. The dynamic management method of device - side memory and host memory according to claim 4, characterized in that, If the device - side memory space is sufficient, map the device - side memory of the required space size, maintain a node, transfer the host memory data required again to the device - side memory space of this node through the DMA function, release the host memory, and set the valid value of this node to 1; Move this node to the head of the LRU double - linked hash list, and send the starting physical address of the node data to the device - side so that the device - side processes the node data.
7. The dynamic management method for device - side memory and host memory according to claim 1, characterized in that, It further includes: After the applied memory space is used up, unmap the device - side memory, and at the same time delete the corresponding node on the LRU double - linked hash list to release the device - side memory space and the host memory space.