A memory management method, device, equipment and medium

By slicing and unified addressing the memory of the heterogeneous acceleration computing system, the problem of memory isolation between AI heterogeneous accelerator devices is solved, and more efficient computing resources and memory resources are achieved, reducing operation and maintenance costs.

CN114020454BActive Publication Date: 2025-05-13LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111257276.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2025-05-13
Estimated Expiration
2041-10-27

AI Technical Summary

Technical Problem

Due to the physical isolation between different AI heterogeneous accelerator devices, memory sharing between AI heterogeneous computing architectures cannot be achieved in different architectures, resulting in the deployment of artificial intelligence algorithm models limited by the onboard memory space size of a single AI accelerator device, and there is a waste of computing resources when the model is executed in parallel.

Method used

By slicing the host memory of the heterogeneous acceleration computing system and the onboard memory of each AI accelerator device separately, the common memory slice space is determined, and the unified address space is addressed, so that each processor can access the public memory slice space to complete the AI ​​algorithm computing task.

Benefits of technology

It breaks through the physical memory isolation limitation between AI heterogeneous acceleration devices, improves the computing resources and memory resource utilization efficiency of heterogeneous acceleration computing systems, and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114020454B_ABST
    Figure CN114020454B_ABST
Patent Text Reader

Abstract

The present application discloses a memory management method, device, equipment and medium, including: slicing the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device to obtain corresponding memory slice space; determining a common memory slice space from all the memory slice spaces; addressing all the common memory slice spaces with a unified address space to obtain corresponding address space; when executing an artificial intelligence algorithm calculation task, deploying the artificial intelligence algorithm model in the address space so that each processor can access the corresponding common memory slice space in the address space to complete the artificial intelligence algorithm calculation task. It can break through the physical isolation limitation of memory between AI heterogeneous accelerated computing devices and improve the utilization efficiency of computing resources and memory resources of heterogeneous accelerated computing systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage technology, and in particular to a memory management method, device, equipment and medium. Background Art

[0002] As the size of data sets increases and models become more complex, the computational cost of AI (Artificial Intelligence) network models is getting higher and higher. As a platform and foundation for carrying AI applications, computing power has promoted the progress and rapid evolution of the entire AI system and is one of the most core elements of AI. The new hybrid heterogeneous computing architecture that solves AI application computing systems has become a hot topic for competition in industry and academia at home and abroad. Heterogeneous accelerators for accelerating AI algorithm applications are emerging in an endless stream, such as GPU (graphics processing unit), FPGA (Field Programmable Gate Array), TPU (tensor processing unit) and various customized AI accelerators. The development of heterogeneous devices provides diverse underlying hardware support for AI scenario applications.

[0003] As the complexity of AI application scenarios increases, and customized AI heterogeneous accelerators are often adapted and optimized for specific computing scenarios, AI applications have put forward requirements for hybrid heterogeneous computing systems. In order to improve the energy efficiency of AI computing systems, AI heterogeneous acceleration devices with different computing characteristics are integrated in a single server system at the same time. Different AI accelerators are used to be responsible for different computing tasks in the same complex AI application scenario. By building an efficient super-heterogeneous computing system, efficient collaboration between different AI heterogeneous acceleration devices in the same AI computing system can be achieved. Although different AI accelerators provide high-capacity onboard memory, due to the physical isolation between different AI heterogeneous accelerator devices, memory sharing between different AI heterogeneous computing architectures cannot be achieved, resulting in the deployment of AI algorithm models often being limited by the size of the onboard memory space of a single AI accelerator device. The current method mainly adopts the model parallel method to split the large-scale AI algorithm model in the network structure, and realizes the training and reasoning of the model by deploying different modules of the model on the onboard memory of different AI accelerators. Then, the model parallel pipeline execution method is used to improve the operating efficiency of the system. However, due to the inability to fully eliminate the bubbling between pipeline stages, it will still cause a certain degree of waste of computing resources. Summary of the invention

[0004] In view of this, the purpose of this application is to provide a memory management method, device, equipment and medium that can break through the physical isolation limitations of memory between AI heterogeneous acceleration devices and improve the utilization efficiency of computing resources and memory resources of heterogeneous accelerated computing systems.

[0005] In a first aspect, the present application discloses a memory management method, comprising:

[0006] The host memory of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device are sliced ​​separately to obtain corresponding memory slice space;

[0007] Determine a common memory slice space from all the memory slice spaces;

[0008] Performing unified address space addressing on all the public memory slice spaces to obtain corresponding addressing spaces;

[0009] When executing an artificial intelligence algorithm computing task, the artificial intelligence algorithm model is deployed in the addressing space so that each processor accesses the corresponding public memory slice space in the addressing space to complete the artificial intelligence algorithm computing task.

[0010] Optionally, the determining a common memory slice space from all the memory slice spaces includes:

[0011] Respectively determine the designated memory slice space among all the memory slice spaces corresponding to the host end and all the memory slice spaces corresponding to the onboard memory of each of the AI ​​accelerator devices as the private memory slice space;

[0012] All non-specified memory slice spaces are determined as common memory slice spaces.

[0013] Optionally, the onboard memory of each AI accelerator device in the heterogeneous accelerated computing system is sliced ​​separately, including:

[0014] All AI accelerator devices in the heterogeneous accelerated computing system are traversed, and the onboard memory of each traversed AI accelerator device is sliced ​​respectively.

[0015] Optionally, in the heterogeneous accelerated computing system, all the AI ​​accelerator devices are mounted to the host end based on a PCIe interface.

[0016] Optionally, the heterogeneous accelerated computing system includes AI accelerator devices with different architectures.

[0017] Optionally, the slicing of the host-side memory of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device to obtain corresponding memory slice space includes:

[0018] According to the preset slice space size, the host-side memory of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device are sliced ​​separately to obtain the corresponding memory slice space.

[0019] Optionally, also include:

[0020] When it is detected that a new AI accelerator device is mounted in the heterogeneous accelerated computing system, the onboard memory of the new AI accelerator device is sliced ​​to obtain a newly added memory slice space;

[0021] Determine a newly added public memory slice space from the newly added memory slice space memory;

[0022] The addressing space is updated based on the newly added public memory slice space to obtain a new addressing space.

[0023] In a second aspect, the present application discloses a memory management device, comprising:

[0024] The memory slice processing module is used to slice the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device to obtain the corresponding memory slice space;

[0025] A public memory slice space determination module, used to determine a public memory slice space from all the memory slice spaces;

[0026] A unified addressing module, used for performing unified address space addressing on all the public memory slice spaces to obtain corresponding addressing spaces;

[0027] The intelligent algorithm computing task execution module is used to deploy the artificial intelligence algorithm model in the addressing space when executing the artificial intelligence algorithm computing task, so that each processor can access the corresponding public memory slice space in the addressing space to complete the artificial intelligence algorithm computing task.

[0028] In a third aspect, the present application discloses an electronic device, comprising:

[0029] Memory, used to store computer programs;

[0030] The processor is used to execute the computer program to implement the aforementioned memory management method.

[0031] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned memory management method when executed by a processor.

[0032] It can be seen that the present application first slices the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device respectively to obtain corresponding memory slice spaces, then determines the common memory slice space from all the memory slice spaces, and then addresses all the common memory slice spaces with a unified address space to obtain the corresponding address space. When executing the artificial intelligence algorithm calculation task, the artificial intelligence algorithm model is deployed in the address space so that each processor can access the corresponding common memory slice space in the address space to complete the artificial intelligence algorithm calculation task. That is, the present application first slices the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device separately, and then determines a common memory slice space therefrom, uniformly addresses the common memory slice space, obtains a unified addressing space, deploys the artificial intelligence algorithm model in the addressing space, and each processor accesses the corresponding common memory slice space in the addressing space to complete the computing task. In this way, the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device are uniformly addressed and managed, realizing the memory resource pooling of the heterogeneous accelerated computing system, breaking through the physical isolation limitation of memory between AI heterogeneous acceleration devices, and improving the computing resource and memory resource utilization efficiency of the heterogeneous accelerated computing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0034] Figure 1 A flow chart of a memory management method disclosed in this application;

[0035] Figure 2 An architectural diagram of a heterogeneous accelerated computing system disclosed in this application;

[0036] Figure 3 A specific memory slice schematic diagram disclosed in this application;

[0037] Figure 4 A schematic diagram of the structure of a memory management device disclosed in this application;

[0038] Figure 5 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION

[0039] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0040] In existing artificial intelligence computing systems, due to the physical isolation between different AI heterogeneous accelerator devices, it is impossible to achieve memory sharing between different AI heterogeneous computing architectures, resulting in artificial intelligence algorithm models being often limited by the size of the onboard memory space of a single AI accelerator device during deployment. The traditional model parallel pipeline method cannot fully improve the overall efficiency of the system due to natural bubbling. To this end, the present application provides a memory management solution that can break through the physical isolation limitations of memory between AI heterogeneous acceleration devices and improve the computing resource and memory resource utilization efficiency of heterogeneous accelerated computing systems.

[0041] See also Figure 1 As shown, the embodiment of the present application discloses a memory management method, including:

[0042] Step S11: Slice the host memory of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device to obtain corresponding memory slice space.

[0043] In a specific implementation, all AI accelerator devices in a heterogeneous accelerated computing system can be traversed, and the onboard memory of each traversed AI accelerator device can be sliced ​​and processed respectively.

[0044] Among them, the heterogeneous accelerated computing system includes AI accelerator devices with different architectures.

[0045] Moreover, in the heterogeneous accelerated computing system, all the AI ​​accelerator devices are mounted to the host end based on a PCIe (peripheral component interconnect express, a high-speed serial computer expansion bus standard) interface.

[0046] That is, in this embodiment, all AI accelerator devices are mounted on the host through the PCIe interface, realizing a unified standard interconnection.

[0047] Step S12: determining a common memory slice space from all the memory slice spaces.

[0048] In a specific implementation, designated memory slice spaces among all memory slice spaces corresponding to the host end and all memory slice spaces corresponding to the onboard memory of each AI accelerator device can be respectively determined as private memory slice spaces; and all non-designated memory slice spaces can be determined as public memory slice spaces.

[0049] Among them, the private memory slice space can only be accessed by the local processor, that is, the private memory slice space of the onboard memory on any AI accelerator device can only be accessed by the AI ​​chip on the AI ​​accelerator device, and the private memory slice space on the host side can only be accessed by the CPU (central processing unit) on the host side. The public memory slice space slice can be accessed by all processors in the system.

[0050] In a specific implementation, the host-side memory of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device can be sliced ​​separately according to the preset slice space size to obtain corresponding memory slice space.

[0051] The preset slice space size can be determined according to the actual scenario and is not limited here.

[0052] Step S13: performing unified address space addressing on all the public memory slice spaces to obtain corresponding addressing spaces.

[0053] Step S14: When executing the artificial intelligence algorithm calculation task, the artificial intelligence algorithm model is deployed in the addressing space so that each processor accesses the corresponding public memory slice space in the addressing space to complete the artificial intelligence algorithm calculation task.

[0054] The memory management method provided in this application can be applied to the host side of a heterogeneous accelerated computing system.

[0055] Further, when it is detected that a new AI accelerator device is mounted in the heterogeneous accelerated computing system, the onboard memory of the new AI accelerator device is sliced ​​to obtain a newly added memory slice space;

[0056] Determine a newly added public memory slice space from the newly added memory slice space memory;

[0057] The addressing space is updated based on the newly added public memory slice space to obtain a new addressing space.

[0058] It can be seen that the present application first slices the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device respectively to obtain corresponding memory slice spaces, then determines the common memory slice space from all the memory slice spaces, and then addresses all the common memory slice spaces with a unified address space to obtain the corresponding address space. When executing the artificial intelligence algorithm calculation task, the artificial intelligence algorithm model is deployed in the address space so that each processor can access the corresponding common memory slice space in the address space to complete the artificial intelligence algorithm calculation task. That is, the present application first slices the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device separately, and then determines a common memory slice space therefrom, uniformly addresses the common memory slice space, obtains a unified addressing space, deploys the artificial intelligence algorithm model in the addressing space, and each processor accesses the corresponding common memory slice space in the addressing space to complete the computing task. In this way, the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device are uniformly addressed and managed, realizing the memory resource pooling of the heterogeneous accelerated computing system, breaking through the physical isolation limitation of memory between AI heterogeneous acceleration devices, and improving the computing resource and memory resource utilization efficiency of the heterogeneous accelerated computing system.

[0059] See also Figure 2 shown, this Figure 2 This is an architecture diagram of a heterogeneous accelerated computing system disclosed in an embodiment of the present application. All AI heterogeneous accelerator devices are mounted in the server using a PCIe interface, that is, mounted on the host side. In the AI ​​accelerator list Device {1, 2, 3, ..., n} of the server system, Device n is represented by Device#. AI accelerator devices of different architectures are supported at the same time. The host-side CPU device can use the PCIe bus interface to access the onboard memory of the AI ​​accelerator device.

[0060] Accordingly, a specific memory management method disclosed in an embodiment of the present application includes:

[0061] Step 01: Slice the memory of the Host (i.e., the host side) and divide it into m memory slice spaces: {H1, H2, H3, ... H#}, where H1 slice is determined as a private memory slice space, and the rest are public memory slice spaces.

[0062] Step 02: Traverse all AI accelerator devices in the host system and slice the onboard memory of all AI accelerator devices {D1, D2, D3, …, D#} into k memory slice spaces: {M1, M2, M3, …, M#}. In any AI accelerator device, the M1 slice is a private memory slice space, and the rest are public memory slice spaces.

[0063] For example, see Figure 3 As shown, Figure 3 A specific memory slice diagram disclosed in an embodiment of the present application. Secondly, the onboard memory of each AI accelerator and the memory of the host side are sliced, and the sliced ​​memory space includes private memory slice space and public memory slice space. The memory slice space can only be accessed by the local processor. For example, the M1 onboard memory slice on the Device1 (ie D1) device can only be accessed by the AI ​​chip on the Device1 accelerator, and the H1 memory slice on the host side can only be accessed by the CPU on the host side. The public memory slice space can be accessed by all processors in the system.

[0064] Step 03: Address the public memory slice space in a unified address space, and the address space is Mem_Addr{H2,…,H#,D1:M2,D1:M3,…,D1:M#,D2:M2,D2:M3,…,D2:M#,…,D#:M#}.

[0065] Step 04: Deploy the artificial intelligence algorithm model in the unified addressing space of the public memory slice space. During the actual computing process, the host CPU and various AI processors access the data in the memory slice on demand.

[0066] That is, the embodiment of the present application uniformly addresses the public memory slice space on the host side and all AI accelerator devices in the heterogeneous accelerated computing system, maintains the unified addressing space of the system memory slice on the host side, and based on the unified addressing space, realizes the memory resource pooling of the super-heterogeneous computing system that supports hybrid heterogeneous acceleration. When facing the computing tasks of artificial intelligence algorithm applications, the AI ​​algorithm model is uniformly deployed in the pooled memory address space, and all AI heterogeneous accelerator devices and host-side CPUs access memory slice data at different addresses as required when performing calculations. In this way, all memories in the heterogeneous accelerated computing system are uniformly managed, the onboard memory space and the host memory space are sliced, and the resource pooling of the onboard memory of the AI ​​heterogeneous accelerator device is realized, and memory access across AI computing accelerator devices is realized, breaking through the limitation of the AI ​​onboard memory capacity on the size of the AI ​​algorithm model, improving the computing resource and memory resource utilization efficiency of the super-heterogeneous computing system, and reducing the operation and maintenance cost of the heterogeneous accelerated computing system.

[0067] See also Figure 4 As shown, the embodiment of the present application discloses a memory management device, including:

[0068] The memory slice processing module 11 is used to slice the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device to obtain corresponding memory slice space;

[0069] A public memory slice space determining module 12, configured to determine a public memory slice space from all the memory slice spaces;

[0070] A unified addressing module 13, used to perform unified address space addressing on all the public memory slice spaces to obtain corresponding addressing spaces;

[0071] The intelligent algorithm computing task execution module 14 is used to deploy the artificial intelligence algorithm model in the addressing space when executing the artificial intelligence algorithm computing task, so that each processor can access the corresponding public memory slice space in the addressing space to complete the artificial intelligence algorithm computing task.

[0072] It can be seen that the present application first slices the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device respectively to obtain corresponding memory slice spaces, then determines the common memory slice space from all the memory slice spaces, and then addresses all the common memory slice spaces with a unified address space to obtain the corresponding address space. When executing the artificial intelligence algorithm calculation task, the artificial intelligence algorithm model is deployed in the address space so that each processor can access the corresponding common memory slice space in the address space to complete the artificial intelligence algorithm calculation task. That is, the present application first slices the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device separately, and then determines a common memory slice space therefrom, uniformly addresses the common memory slice space, obtains a unified addressing space, deploys the artificial intelligence algorithm model in the addressing space, and each processor accesses the corresponding common memory slice space in the addressing space to complete the computing task. In this way, the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device are uniformly addressed and managed, realizing the memory resource pooling of the heterogeneous accelerated computing system, breaking through the physical isolation limitation of memory between AI heterogeneous acceleration devices, and improving the computing resource and memory resource utilization efficiency of the heterogeneous accelerated computing system.

[0073] Among them, the public memory slice space determination module 12 is specifically used to determine the designated memory slice space among all the memory slice spaces corresponding to the host end and all the memory slice spaces corresponding to the onboard memory of each of the AI ​​accelerator devices as private memory slice space; and determine all non-designated memory slice space as public memory slice space.

[0074] In addition, the memory slice processing module 11 is specifically used to traverse all AI accelerator devices in the heterogeneous accelerated computing system, and slice the onboard memory of each traversed AI accelerator device respectively.

[0075] In the heterogeneous accelerated computing system, all the AI ​​accelerator devices are mounted to the host end based on the PCIe interface.

[0076] Furthermore, the heterogeneous accelerated computing system includes AI accelerator devices with different architectures.

[0077] The memory slice processing module 11 is specifically used to slice the host-side memory of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device according to a preset slice space size to obtain corresponding memory slice space.

[0078] Furthermore, the device also includes: a newly added mounting detection module, which is used to detect whether a new AI accelerator device is mounted in the heterogeneous accelerated computing system.

[0079] Correspondingly, the memory slice processing module 11 is used to slice the onboard memory of the new AI accelerator device to obtain a new memory slice space when the newly added mounting detection module detects that a new AI accelerator device is mounted in the heterogeneous accelerated computing system;

[0080] A public memory slice space determination module 12, specifically configured to determine a newly added public memory slice space from the newly added memory slice space memory;

[0081] The unified addressing module is specifically used to update the addressing space based on the newly added public memory slice space to obtain a new addressing space.

[0082] See also Figure 5 As shown, an embodiment of the present application discloses an electronic device 20, including a processor 21 and a memory 22; wherein the memory 22 is used to store a computer program; the processor 21 is used to execute the computer program, and the memory management method disclosed in the above embodiment.

[0083] For the specific process of the above memory management method, please refer to the corresponding content disclosed in the above embodiments, which will not be repeated here.

[0084] Furthermore, the memory 22 as a carrier for storing resources may be a read-only memory, a random access memory, a magnetic disk or an optical disk, etc., and the storage method may be temporary storage or permanent storage.

[0085] In addition, the electronic device 20 also includes a power supply 23, a communication interface 24, an input / output interface 25 and a communication bus 26; wherein the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and an external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input / output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0086] Furthermore, an embodiment of the present application also discloses a computer-readable storage medium for storing a computer program, wherein the computer program, when executed by a processor, implements the memory management method disclosed in the aforementioned embodiment.

[0087] For the specific process of the above memory management method, please refer to the corresponding content disclosed in the above embodiments, which will not be repeated here.

[0088] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0089] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0090] The above is a detailed introduction to a memory management method, device, equipment and medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for general technical personnel in this field, according to the idea of ​​the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A memory management method, characterized in that: include: The host memory of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device are sliced ​​separately to obtain corresponding memory slice space; Determine a common memory slice space from all the memory slice spaces; Performing unified address space addressing on all the public memory slice spaces to obtain corresponding addressing spaces; When executing an artificial intelligence algorithm computing task, the artificial intelligence algorithm model is deployed in the addressing space so that each processor accesses the corresponding public memory slice space in the addressing space to complete the artificial intelligence algorithm computing task; The step of determining a common memory slice space from all the memory slice spaces includes: Respectively determine the designated memory slice space among all the memory slice spaces corresponding to the host end and all the memory slice spaces corresponding to the onboard memory of each of the AI ​​accelerator devices as the private memory slice space; All non-specified memory slice spaces are determined as common memory slice spaces.

2. The memory management method according to claim 1, characterized in that: The onboard memory of each AI accelerator device in the heterogeneous accelerated computing system is sliced ​​separately, including: All AI accelerator devices in the heterogeneous accelerated computing system are traversed, and the onboard memory of each traversed AI accelerator device is sliced ​​respectively.

3. The memory management method according to claim 1, characterized in that: In the heterogeneous accelerated computing system, all the AI ​​accelerator devices are mounted to the host end based on the PCIe interface.

4. The memory management method according to claim 1, characterized in that: The heterogeneous accelerated computing system includes AI accelerator devices with different architectures.

5. The memory management method according to claim 1, characterized in that: The host-side memory of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device are sliced ​​to obtain corresponding memory slice spaces, including: According to the preset slice space size, the host-side memory of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device are sliced ​​separately to obtain the corresponding memory slice space.

6. The memory management method according to claim 1, characterized in that: Also includes: When it is detected that a new AI accelerator device is mounted in the heterogeneous accelerated computing system, the onboard memory of the new AI accelerator device is sliced ​​to obtain a newly added memory slice space; Determine a newly added public memory slice space from the newly added memory slice space memory; The addressing space is updated based on the newly added public memory slice space to obtain a new addressing space.

7. A memory management device, characterized in that: include: The memory slice processing module is used to slice the memory of the host side of the heterogeneous accelerated computing system and the onboard memory of each AI accelerator device to obtain the corresponding memory slice space; A public memory slice space determination module, used to determine a public memory slice space from all the memory slice spaces; A unified addressing module, used for performing unified address space addressing on all the public memory slice spaces to obtain corresponding addressing spaces; An intelligent algorithm computing task execution module is used to deploy the artificial intelligence algorithm model in the addressing space when executing the artificial intelligence algorithm computing task, so that each processor can access the corresponding public memory slice space in the addressing space to complete the artificial intelligence algorithm computing task; Wherein, the public memory slice space determination module is used for: Respectively determine the designated memory slice space among all the memory slice spaces corresponding to the host end and all the memory slice spaces corresponding to the onboard memory of each of the AI ​​accelerator devices as the private memory slice space; All non-specified memory slice spaces are determined as common memory slice spaces.

8. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the memory management method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the memory management method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Physical access control

    CA2814254A1

  • Segmentation and paging data storage space management method facing heterogeneous polynuclear system

    CN101008923A