Memory management method, device and system, storage medium and program product
By mapping the memory operations of hardware devices to memory allocators, the coupling problem of deploying large language models on different hardware devices is solved, achieving efficient and convenient memory management and improving the deployment efficiency and data throughput of the models.
Patent Information
- Application Number
- CN202511463881.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-09
AI Technical Summary
Different memory processing strategies on different hardware devices limit the deployment of large language models, creating coupling issues that affect deployment efficiency and convenience.
By configuring memory operation mappings of multiple hardware devices to corresponding memory allocators, and using multiple memory allocators to match the memory processing needs of different models, the memory control during model deployment is decoupled from the hardware devices.
It improves the efficiency and convenience of model deployment, reduces the use of GPU memory resources by temporary memory operations, lowers resource overhead, and increases data throughput.
Smart Images

Figure CN121301007A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, specifically to memory management methods, devices, systems, storage media, and program products. Background Technology
[0002] Large Language Models (LLMs), once scaled to a certain size, have shown significant application prospects in multiple fields, such as automatic text generation, machine translation, and programming.
[0003] Since the number of parameters in an LLM can reach hundreds of thousands or even millions, it places high demands on the memory processing performance of the inference environment. Currently, there are numerous hardware devices for large model inference, but different hardware devices have different built-in memory processing strategies, which are suitable for different LLM applications. In other words, there is a certain degree of coupling between the LLM and the hardware device, which makes the deployment of the LLM highly limited by the hardware device. Summary of the Invention
[0004] The purpose of this disclosure is to provide a memory management method, apparatus, system, and related devices to solve the technical problem of inconvenient model deployment under different hardware devices.
[0005] In a first aspect, embodiments of this disclosure provide a memory management method, the method comprising: Determine the session control object for the model; From a plurality of memory allocators, at least one target memory allocator is determined to be associated with the session control object of the model, wherein the plurality of memory allocators are determined based on a plurality of memory operations supported by a plurality of hardware devices, and different memory allocators correspond to different hardware devices and / or correspond to different memory operations; The memory operation tasks of the model are performed using the at least one target memory allocator.
[0006] Secondly, embodiments of this disclosure also provide a memory management device, the device comprising: The first determination module is used to determine the session control object of the model; The second determining module is used to determine at least one target memory allocator associated with the session control object of the model from a plurality of memory allocators, wherein the plurality of memory allocators are determined based on a plurality of memory operations supported by a plurality of hardware devices, and different memory allocators correspond to different hardware devices and / or correspond to different memory operations. An execution module is used to perform memory operation tasks of the model using the at least one target memory allocator.
[0007] Thirdly, this disclosure provides a memory management apparatus, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method described in the first aspect.
[0008] Fourthly, this disclosure provides a memory management system, including a memory management device as described in the second or third aspect and a plurality of memory allocators.
[0009] Fifthly, this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0010] In a sixth aspect, this disclosure provides a computer program product including computer instructions that, when executed by a processor, implement the steps of the method described in the first aspect.
[0011] In this disclosure, multiple hardware devices are configured, and each memory operation supported by each hardware device is mapped to a corresponding memory allocator. This allows multiple memory allocators to be used to match the different memory processing needs of different models during the deployment phase as much as possible, thereby decoupling memory control from hardware devices during model deployment and improving the efficiency and convenience of model deployment. Attached Figure Description
[0012] Preferred embodiments of the present disclosure are described below with reference to the accompanying drawings. The accompanying drawings, which are included to provide a further understanding of the present disclosure, and which, together with the following detailed description, are incorporated in and form a part of this specification and are used to explain the present disclosure. It should be understood that the drawings described below only relate to some embodiments of the present disclosure and are not intended to limit the present disclosure. In the drawings: Figure 1 This is a flowchart illustrating a memory management method provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of a device interface provided in an embodiment of this disclosure; Figure 3 This is a partial schematic diagram of a computation graph provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of the structure of a memory management device provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of another memory management device provided in an embodiment of this disclosure.
[0013] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals may be used in different drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation
[0014] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. The following description of the embodiments is merely illustrative and is in no way intended to limit this disclosure or its application or use. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0015] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.
[0016] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Furthermore, as used in this disclosure, the term "including" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". Therefore, "comprising" and "including" are synonymous. The term "based on" means "at least partially based on".
[0017] Throughout this specification, the terms "one embodiment," "some embodiments," or "embodiment" mean that a specific feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. For example, the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; and the term "some embodiments" means "at least some embodiments." Furthermore, the phrases "in one embodiment," "in some embodiments," or "in an embodiment" appearing in various places throughout the specification do not necessarily all refer to the same embodiment, but may refer to the same embodiment.
[0018] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.
[0019] It should be noted that the terms “a,” “a plurality of,” and “at least one” used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as “one or more.”
[0020] This disclosure provides a memory management method, such as... Figure 1 As shown, the memory management method includes steps 101 to 103.
[0021] Step 101: Determine the session control object for the model. In some embodiments, a session control object for the model may be created or constructed. For example, the session control object may be understood as a Session class or object.
[0022] In some embodiments, the model includes, but is not limited to, a large language model. Models involving memory management can all employ the memory management method of any embodiment of this disclosure for memory management or memory control.
[0023] Step 102: From a plurality of memory allocators, determine at least one target memory allocator associated with the session control object of the model.
[0024] The aforementioned memory allocators are determined or constructed based on the various memory operations supported by multiple hardware devices. Different memory allocators correspond to different hardware devices and / or different memory operations. Each hardware device supports multiple memory operations, and the memory operations supported by different hardware devices may be partially the same or completely different. For example, each memory allocator corresponds to one type of memory operation within a single hardware device.
[0025] In some embodiments, all memory operations involved in the model constitute a first set of operations, and all memory operations supported by the aforementioned multiple hardware devices constitute a second set of operations. The intersection-union ratio (IUU) between the first and second set of operations is greater than or equal to a preset threshold, such as 0.8 or 0.9, to ensure that the aforementioned multiple hardware devices fully support the memory operations involved in the model, thereby improving the versatility and universality of this method in practical applications. The multiple hardware devices refer to multiple different hardware devices supporting model deployment, including but not limited to hardware devices from different manufacturers or different types of hardware devices from the same manufacturer.
[0026] In some embodiments, at least one target memory allocator associated with the session control object of the model can be determined from a plurality of memory allocators in the following manner.
[0027] First, based on the session control object of the model, determine the model's identification information and the task description information of the model's memory operation tasks. For example, the session control object can carry the model's identification information and the task description information of the model's memory operation tasks.
[0028] Secondly, based on the correspondence between the model and the hardware device, the hardware device corresponding to the model's identification information is determined.
[0029] Then, based on the correspondence between hardware devices and memory operations, the memory operations supported by the hardware devices corresponding to the model's identification information are determined.
[0030] Finally, based on the task description information of the model's memory operation task, at least one memory allocator is selected from multiple memory allocators. This allocator is constructed based on the memory operations supported by the hardware device corresponding to the model's identification information and is used as at least one target memory allocator.
[0031] The following describes in detail the process of determining at least one target memory allocator associated with the session control object of the model from multiple memory allocators. This embodiment is only a specific implementation method and does not constitute a specific limitation on the memory management method disclosed herein.
[0032] For example, a device-memory operation-allocator mapping table and a model-device mapping table can be created to determine the at least one target memory allocator. The device-memory operation-allocator mapping table stores multiple first mapping entries, each defining a number of memory operations supported by a hardware device and the memory allocator corresponding to each memory operation supported by that hardware device. The model-device mapping table stores multiple second mapping entries, each defining a hardware device adapted to a model. The hardware device adapted to the model is the hardware device that supports the deployment of the model.
[0033] First, based on the aforementioned session control object, the model's identification information and the task description information of the model's memory operation tasks can be determined. Second, by looking up the identification information in the aforementioned model-device mapping table, the device identifier of the hardware device compatible with the model can be obtained. Then, by looking up the device identifier in the aforementioned device-memory operation-allocator mapping table, several memory operations supported by the corresponding hardware device can be determined, i.e., several memory operations associated with the model deployment. Finally, based on the task description information, a search is performed among the several memory operations associated with the model deployment to identify at least one memory operation matching the task description information. The memory allocator corresponding to the at least one memory operation matching the task description information is then determined as the at least one target memory allocator.
[0034] For example, such as Figure 2 As shown, multiple memory allocators can be deployed within a Device Application Programming Interface (Device API). As previously mentioned, these multiple memory allocators are abstracted / constructed based on various memory operations supported by multiple hardware devices (including but not limited to hardware devices from a first vendor, hardware devices from a second vendor, etc.). This abstraction can be understood as decoupling the various memory operations supported by different hardware devices from the hardware devices themselves, treating them as separate functional modules, such as those deployed within the memory allocator.
[0035] Step 103: Execute the memory operation task of the model using the at least one target memory allocator.
[0036] In some embodiments, the model's session control object is used to invoke the at least one target memory allocator to perform the model's memory operation tasks. For example, the model's session control object can instruct the model's memory operation tasks.
[0037] The memory operation tasks of the above model can be understood as the various memory processing tasks involved in the deployment of the model.
[0038] In some embodiments, the memory operation task includes at least one of the following: Tasks related to device-to-device (D2D) transfers; Tasks related to device-to-host (D2H) transfers; Tasks related to host-to-device (H2D) transfers; Tasks related to stream operations; Tasks related to memory allocation and / or deallocation.
[0039] In this disclosure, multiple hardware devices are configured, and each memory operation supported by each hardware device is mapped or abstracted into a corresponding memory allocator. This allows multiple memory allocators to be used to match the different memory processing needs of different models during the deployment phase as much as possible. This decouples memory control during model deployment from hardware devices, making the deployment of different models under different hardware devices more efficient and convenient, and improving the efficiency and convenience of model deployment.
[0040] In some embodiments, the at least one target memory allocator includes a first target memory allocator; the step of using the at least one target memory allocator to perform the memory operation task of the model includes: Using the first target memory allocator, temporary memory is requested for the first data of the model; After the first data is no longer in use, the temporary memory of the first data is released using the first target memory allocator; The model is deployed on a target device, and the target device's video memory does not include the temporary memory. The target device is one of the aforementioned hardware devices.
[0041] In this embodiment, by using the first target memory allocator to skip video memory and perform lightweight allocation and release of temporary memory, the calls to video memory resources by temporary memory-related operations can be reduced, thereby reducing unnecessary overhead caused by calling video memory resources and thus reducing the overall overhead of temporary memory calls.
[0042] In some embodiments, temporary memory allocation can be accomplished through a constructor, while temporary memory release can be accomplished through a destructor.
[0043] For example, if the first target memory allocator is set to a pool-based allocator, then for CUDA's top-k operator, it can be used in conjunction with PooledAllocator and C++'s Resource Acquisition Is Initialization (RAII) mechanism to complete the allocation and release of temporary memory corresponding to the top-k operator. The top-k operator is used to filter out the top K elements according to a set sorting relationship, and the temporary memory is used to store the filtered top K elements.
[0044] The memory management process related to the model's operators will be described below with reference to some examples.
[0045] Similar to general-purpose programming languages, the deployment of large language models can be divided into compiletime and runtime. The compiletime is responsible for rewriting and optimizing the large language model and performing some preparatory work before model inference execution to avoid additional overhead during execution. The runtime is responsible for executing the computation graph corresponding to the large language model and also supports a small number of Just-in-Time (JIT) compilation operations.
[0046] A computation graph can be understood as a graph structure composed of basic data structures (tensors) and basic operational units (operators). Nodes are used to represent operators in a computation graph, and directed edges between nodes represent tensor states, while also describing the dependencies between computations.
[0047] In applications, for large model inference frameworks that support dynamic size input requirements (e.g., VLLM and LightLLM frameworks), memory can be allocated according to a preset maximum size for each tensor in the computation graph. Before executing each operator during execution, the tensor is divided according to its actual size, so that each operator is calculated using the actual size, enabling the correct inference results to be obtained during execution. The above-mentioned dynamic size adjustment operations are completed between nodes in the computation graph.
[0048] In some embodiments, the at least one target memory allocator includes a second target memory allocator; the step of performing the memory operation task of the model using the at least one target memory allocator includes: When the execution of the first operator of the model ends, the second target memory allocator is used to reuse the data memory of the second data referenced by the first operator in the second operator of the model. The computation graph corresponding to the model includes the first operator and the second operator. The execution order of the second operator in the computation graph is lower than the execution order of the first operator in the computation graph, and the first operator is the last operator in the computation graph to reference the second data.
[0049] In this embodiment, when performing specific memory allocation operations within a node of the computation graph, if it is determined that the reference relationship of a certain data memory has ended (meaning the execution of the last operator using that data memory has ended), other operators following the last operator referencing that data memory can reuse that data memory (i.e., reuse it repeatedly, but the data in the data memory must be cleared before reuse). This reduces the frequency of data memory generation and recycling, which can further improve the memory usage efficiency of the model during deployment, reduce the resource overhead of memory allocation, and increase the number of user requests that the target device can process per unit time.
[0050] In some embodiments, prior to performing the memory operation task of the model using the at least one target memory allocator, the method further includes: The computation graph corresponding to the model is compiled to determine at least one reusable memory associated with the computation graph, wherein the at least one reusable memory includes the data memory of the second data. Reusable memory refers to data memory that can be reused by multiple operators of the computation graph.
[0051] By using at least one reusable memory that can be reused in the compile-time statistical computation graph, the memory reuse actually performed during the execution phase based on the at least one reusable memory is supported, ensuring the correct execution of the aforementioned memory reuse process.
[0052] In some embodiments, the memory used by the input data and the memory used by the output data in the computation graph can both be defined as reusable memory, and the activation time of each reusable memory can be set to the execution end time of the last operator that references the corresponding input data or output data.
[0053] It should be understood that the compilation phase only completes the statistical analysis of information on at least one reusable memory, and the execution phase will then actually perform the specific memory reuse operation based on the statistical information.
[0054] For example, a portion of the aforementioned computational graph is, for instance... Figure 3 As shown, Figure 3 The input tensor is processed by the two-dimensional convolution (Conv2D) operator to form output result 1. W-conv2d is used to represent the convolution weights used by the Conv2D operator when performing convolution. Output 1 is processed by the batch normalization operator to form Output 2, where W-batchnorm represents the batch normalization weights used by the batch normalization operator during batch normalization. Output 2 and the input tensor are summed by the Add operator to form output 3. Output 3 is then processed by the ReLU operator and used by other operators or output externally.
[0055] Since the Batchnorm operator is the last operator to reference output 1, to minimize the resource overhead of the memory allocation process, after the Batchnorm operator finishes execution, the memory used by output 1 can be retained (instead of being directly reclaimed) for use by other operators following the Batchnorm operator. For example, after the Batchnorm operator finishes execution, the memory used by output 1 can be used as the input data memory for the Add operator, or the memory used by output 1 can be used as the input data memory for the ReLU operator, etc.
[0056] For the model, the execution process is explicitly divided into two stages: Prefill and Decode. These two stages have different memory allocation requirements. For high performance considerations, related technologies often choose to implement separate memory allocation strategies for the operators in the Prefill and Decode stages. This means that high-frequency memory generation and release operations are performed in both the Prefill and Decode stages. These high-frequency memory operations greatly increase the resource overhead of the memory allocation process. However, by adopting the scheme described in this disclosure, and utilizing the aforementioned memory reuse mechanism, a batch of memory can be maintained across stages for different operators in the Prefill and Decode stages. This can meet the different memory allocation requirements of Prefill and Decode while greatly reducing the frequency of memory generation and release operations. Therefore, it can reduce the resource overhead caused by high-frequency memory generation and release operations and improve the data throughput of the model.
[0057] Those skilled in the art should understand that the description of the technical solution of this disclosure using CUDA as an example is only for better explaining the core concept of this disclosure and should not be construed as limiting the scope of this disclosure. The solution provided in this disclosure is also applicable to other parallel computing architectures.
[0058] See Figure 4 , Figure 4 This is a memory management device provided in the embodiments of this disclosure, such as... Figure 4 As shown, the memory management device 400 includes: The first determining module 401 is used to determine the session control object of the model; The second determining module 402 is used to determine at least one target memory allocator associated with the session control object of the model from a plurality of memory allocators, wherein the plurality of memory allocators are determined based on a plurality of memory operations supported by a plurality of hardware devices, and different memory allocators correspond to different hardware devices and / or correspond to different memory operations. Execution module 403 is used to perform memory operation tasks of the model using the at least one target memory allocator.
[0059] In some embodiments, the at least one target memory allocator includes a first target memory allocator; the execution module 403 includes: A temporary memory allocation unit is used to allocate temporary memory for the first data of the model using the first target memory allocator; A temporary memory release unit is used to release the temporary memory of the first data using the first target memory allocator after the first data has finished being used. The model is deployed on a target device, the target device's video memory does not include the temporary memory, and the target device is one of the plurality of hardware devices.
[0060] In some embodiments, the at least one target memory allocator includes a second target memory allocator; the execution module 403 includes: A memory reuse unit is used to reuse the data memory of the second data referenced by the first operator in the second operator of the model, based on the second target memory allocator, when the execution of the first operator of the model ends. The computation graph corresponding to the model includes the first operator and the second operator. The execution order of the second operator in the computation graph is lower than the execution order of the first operator in the computation graph, and the first operator is the last operator in the computation graph to reference the second data.
[0061] In some embodiments, the memory management device 400 further includes: A compilation module is used to compile the computation graph corresponding to the model and determine at least one reusable memory associated with the computation graph, wherein the at least one reusable memory includes the data memory of the second data.
[0062] In one embodiment, the memory operation task includes at least one of the following: Tasks related to device-to-device transfers; Tasks related to device-to-host transfers; Tasks related to host-to-device transfers; Tasks related to stream operations; Tasks related to memory allocation and / or deallocation.
[0063] The memory management device 400 provided in this embodiment can implement the various processes in the above-described memory management method embodiments, and will not be repeated here to avoid repetition.
[0064] According to embodiments of this disclosure, this disclosure also provides a memory management device and a readable storage medium.
[0065] Figure 5 A schematic block diagram of an example memory management device 500 that can be used to implement embodiments of the present disclosure is shown. The memory management device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The memory management device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0066] like Figure 5 As shown, the memory management device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0067] Multiple components in the memory management device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard or mouse; an output unit 507, such as various types of displays or speakers; a storage unit 508, such as a hard disk or optical disk; and a communication unit 509, such as a network interface card (NIC), a modem, or a wireless transceiver. The communication unit 509 allows the device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0068] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as memory management methods. For example, in some embodiments, the memory management method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the memory management device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the memory management methods described above may be performed. Alternatively, in other embodiments, computing unit 501 may be configured to perform memory management methods by any other suitable means (e.g., by means of firmware).
[0069] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0070] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0071] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0072] As used herein, the term "machine-readable medium" refers to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0073] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0074] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0075] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0076] This disclosure also provides a memory management system, including the memory management apparatus described in any embodiment of this disclosure and a plurality of memory allocators. The plurality of memory allocators are as described above and will not be repeated here.
[0077] This disclosure also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.
[0078] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0079] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A memory management method, characterized in that, include: Determine the session control object for the model; From a plurality of memory allocators, at least one target memory allocator is determined to be associated with the session control object of the model, wherein the plurality of memory allocators are determined based on a plurality of memory operations supported by a plurality of hardware devices, and different memory allocators correspond to different hardware devices and / or correspond to different memory operations; The memory operation tasks of the model are performed using the at least one target memory allocator.
2. The memory management method according to claim 1, characterized in that, The at least one target memory allocator includes a first target memory allocator; the step of using the at least one target memory allocator to perform the memory operation task of the model includes: Using the first target memory allocator, temporary memory is requested for the first data of the model; After the first data is no longer in use, the temporary memory of the first data is released using the first target memory allocator; The model is deployed on a target device, the target device's video memory does not include the temporary memory, and the target device is one of the plurality of hardware devices.
3. The memory management method according to claim 1, characterized in that, The at least one target memory allocator includes a second target memory allocator; the step of using the at least one target memory allocator to perform the memory operation task of the model includes: When the execution of the first operator of the model ends, the second target memory allocator is used to reuse the data memory of the second data referenced by the first operator in the second operator of the model. The computation graph corresponding to the model includes the first operator and the second operator. The execution order of the second operator in the computation graph is lower than the execution order of the first operator in the computation graph, and the first operator is the last operator in the computation graph to reference the second data.
4. The memory management method according to claim 3, characterized in that, Before performing the memory operation task of the model using the at least one target memory allocator, the method further includes: The computation graph corresponding to the model is compiled, and at least one reusable memory associated with the computation graph is determined, wherein the at least one reusable memory includes the data memory of the second data.
5. The memory management method according to any one of claims 1-4, characterized in that, The memory operation task includes at least one of the following: Tasks related to device-to-device transfers; Tasks related to device-to-host transfers; Tasks related to host-to-device transfers; Tasks related to stream operations; Tasks related to memory allocation and / or deallocation.
6. A memory management device, characterized in that, include: The first determination module is used to determine the session control object of the model; The second determining module is used to determine at least one target memory allocator associated with the session control object of the model from a plurality of memory allocators, wherein the plurality of memory allocators are determined based on a plurality of memory operations supported by a plurality of hardware devices, and different memory allocators correspond to different hardware devices and / or correspond to different memory operations. An execution module is used to perform memory operation tasks of the model using the at least one target memory allocator.
7. A memory management device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method as described in any one of claims 1 to 5.
8. A memory management system, characterized in that, include: The memory management device as described in claim 6 or 7; and Multiple memory allocators.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 5.
10. A computer program product, characterized in that, Includes computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 5.