Method and system for managing memory usage in data learning operations
Through the analysis-based memory allocation algorithm, collecting and analyzing object information of the training machine learning model and optimizing memory allocation and use, the challenges of memory management during the training process are solved, and more efficient memory usage and computational overhead are achieved.
Patent Information
- Application Number
- CN202411468580.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-10-21
- Publication Date
- 2025-05-20
AI Technical Summary
Memory management becomes a key challenge when training machine learning models, especially when data size increases and distributed systems collaboratively train, how to effectively manage memory resources to reduce computational overhead and memory consumption.
Using an analysis-based memory allocation algorithm, by analyzing machine learning parameters of the training machine learning model, the object size, memory allocation and de-allocation information are collected, and the memory usage is scheduled to optimize memory allocation, reduce memory fragmentation, and the same allocation information is applied to subsequent allocation steps.
It effectively reduces memory consumption and computing overhead, optimizes memory usage, avoids memory fragmentation, and supports more efficient data learning operations.
Smart Images

Figure CN120020695A_ABST
Abstract
Description
Technical Field
[0001] The embodiments described herein generally relate to managing data storage in data learning operations, such as training in data learning operations, e.g., training a machine learning model. More specifically, the embodiments described herein relate to methods and systems for managing memory usage in data learning operations, such as for training a machine learning model in a distributed training and / or fine-tuning system. Background Art
[0002] Memory management for training machine learning models has been an ongoing focus in the field of machine learning because each processing unit, e.g., a graphics processing unit (GPU) or a central processing unit (CPU), has a limited memory capacity. As the data for training machine learning models increases, e.g., training hundreds of billions of parameters, memory management remains a key challenge even when distributed systems are used to collaborate to train machine learning models together. Summary of the Invention
[0003] Features in the embodiments disclosed herein can support data learning operations, e.g., operations included in training a machine learning model, by improving the management of memory usage to reduce and / or optimize memory consumption and / or reduce the computational overhead of data learning operations. The memory management mechanisms discussed herein can be used in deep learning architectures, e.g., for natural language processing, computer vision, audio processing, multimodal processing, generative pre-training architectures, or directional encoder representations from transformers, where input data, model parameters, forward activations, backward gradients, optimizer states can use the processing device memory during the training of a machine learning model.
[0004] Features in the embodiments disclosed herein may include a memory management mechanism that includes an analytics-based memory allocation algorithm that utilizes information about allocated objects, which is pre-obtained through analysis, for machine learning parameters for training a machine learning model. Such a memory management mechanism can address the problems of prior art solutions by using analytics to collect at least the size, memory allocation, and memory deallocation of objects for training a machine learning model, e.g., via timestamps, to reduce memory consumption. The memory management mechanism may also include a scheduling step that uses the analyzed information to determine the size and address range of memory allocations, e.g., blocks in a machine learning model, and may then determine the total consumption of memory usage, e.g., GPU memory, which is optimized to reduce memory consumption. Then, the scheduling step can be applied to any subsequent allocations, e.g., execution steps or blocks for training a machine learning model. In this way, the memory allocation algorithm can include not only memory management that can be optimized by reducing memory consumption, e.g., by avoiding memory fragmentation, but also reduce computational overhead by applying the same allocation information or parameters to any subsequent allocations (e.g., execution steps or blocks) for training a machine learning model. In some embodiments, the analytics-based memory allocation algorithm can be a parallel ladder algorithm that is designed, programmed, or otherwise configured to balance memory consumption and computational complexity, which results in less memory consumption compared to existing suboptimal algorithms.
[0005] In one example embodiment, a method for managing memory usage in data learning operations is provided. The method includes analyzing one or more objects operating in a data learning operation, where the analysis includes determining an object size, a memory allocation timestamp, and a memory deallocation timestamp of the one or more objects; and scheduling memory usage for the one or more objects to determine a total size and an address range of one or more groups of the one or more objects. The scheduling includes: grouping the one or more objects into one or more groups based on the memory allocation timestamp and / or the memory deallocation timestamp of the one or more objects, and arranging the one or more objects in one or more groups in descending order in a memory space, where in the case where the one or more objects include two or more objects, the two or more objects are provided in the memory space in descending order from a first value to a second value less than the first value, where an object among the two or more objects having the earliest memory allocation timestamp is provided at the first value, and another object among the two or more objects having a later memory allocation timestamp is provided at the second value.
[0006] In another embodiment, a method for managing memory usage in a data learning operation is provided. The method includes analyzing a batch of objects, where each object in the batch includes a size, an allocation timestamp, and a deallocation timestamp; and arranging the analyzed batch of objects in a memory space to form a combined memory address range corresponding to the batch of objects. The arrangement includes determining whether the deallocation of one object in the batch of objects occurs after the deallocation of another object in the batch of objects based on the deallocation timestamp, and rearranging the order of the arranged batch of objects in the case where deallocation occurs, such that the memory usage of the combined memory address range for the batch of objects in the memory space is maximized.
[0007] In yet another embodiment, a data learning operation training system is provided. The system includes a memory for storing data learning operations; at least one processor for: analyzing one or more objects for training a data learning operation, where the analysis includes determining an object size, a memory allocation timestamp, and a memory deallocation timestamp of the one or more objects; and scheduling memory usage for the one or more objects to determine a total size and an address range of one or more groups of the one or more objects. The scheduling includes: grouping the one or more objects into one or more groups based on the memory allocation timestamp and / or the memory deallocation timestamp of the one or more objects, and arranging the one or more objects in the one or more groups in descending order in a memory space, where in the case where the one or more objects include two or more objects, the two or more objects are provided in the memory space in descending order from a first value to a second value less than the first value, where one object having the earliest memory allocation timestamp among the two or more objects is provided at the first value, and another object having a later memory allocation timestamp among the two or more objects is provided at the second value.
[0008] In yet another example embodiment, a non-transitory computer-readable medium storing computer-executable instructions is provided. When executed, these instructions cause one or more processors to perform operations including analyzing one or more objects for a training data learning operation, where the analysis includes determining an object size, a memory allocation timestamp, and a memory deallocation timestamp of the one or more objects; and scheduling the memory usage of the one or more objects to determine a total size and an address range of one or more groups of the one or more objects. The scheduling includes: grouping the one or more objects into one or more groups based on the memory allocation timestamp and / or the memory deallocation timestamp of the one or more objects, and arranging the one or more objects in the one or more groups in descending order in the memory space, where in the case where the one or more objects include two or more objects, the two or more objects are provided in the memory space in descending order from a first value to a second value less than the first value, where one object having the earliest memory allocation timestamp among the two or more objects is provided at the first value, and another object having a later memory allocation timestamp among the two or more objects is provided at the second value. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The drawings illustrate various embodiments of the systems, methods, and various other aspects of the present disclosure. Those of ordinary skill in the art will understand that the illustrated element boundaries in the drawings (e.g., boxes, groups of boxes, or other shapes) represent one example of the boundaries. In some examples, one element may be designed as multiple elements, or multiple elements may be designed as one element. In some examples, an element shown as an internal component of one element may be implemented as an external component of another element, and vice versa. The following drawings are described in a non-limiting and non-exhaustive manner. The components in the drawings are not necessarily drawn to scale, with the emphasis on illustrating the principles. In the following detailed description, the embodiments are described only by way of illustration, as various changes and modifications will become apparent to those skilled in the art from the following detailed description.
[0010] Figure 1 is a schematic diagram of an example memory management system arranged according to at least some embodiments described herein.
[0011] Figure 2 is a schematic diagram of an example processing flow for managing the memory usage for training a machine learning model arranged according to at least some embodiments described herein.
[0012] Figure 3 is a schematic diagram of an example object mapping process arranged according to at least some embodiments described herein.
[0013] Figure 4Schematic diagram of an analysis-based memory allocation algorithm arranged according to at least some embodiments described herein.
[0014] Figure 5A and 5B Schematic diagram of a memory space including objects with overlapping lifetimes according to at least some embodiments described herein.
[0015] Figure 6A and 6B Schematic diagram of a memory space including objects with overlapping lifetimes arranged using an analysis-based memory allocation algorithm according to at least some embodiments described herein.
[0016] Figures 7A - 7I Schematic diagram of a method for managing memory usage using an analysis-based memory allocation algorithm according to at least some embodiments described herein.
[0017] Figure 8 Schematic structural diagram of an example computer system suitable for implementing an electronic device arranged according to at least some embodiments described herein. Detailed Description
[0018] In the following detailed description, specific embodiments of the present disclosure are described with reference to the accompanying drawings, which form a part of the description. In this specification and the accompanying drawings, unless the context otherwise requires, the same reference numerals denote elements that can perform the same, similar, or equivalent functions. Additionally, unless otherwise stated, the description of each successive drawing can refer to features from one or more previous drawings to provide a clearer context and a more substantial explanation of the current example embodiments. However, the example embodiments described in the detailed description, the drawings, and the claims are not intended to be limiting. Other embodiments can be utilized and other changes can be made without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure generally described and illustrated in the drawings can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are explicitly contemplated herein.
[0019] It should be understood that the disclosed embodiments are merely examples of the present disclosure, which can be embodied in various forms. Well-known functions or configurations are not described in detail so as not to obscure the present disclosure with unnecessary details. Therefore, the specific structural and functional details disclosed herein should not be construed as restrictive, but merely as a basis for the claims and as a representative basis for teaching those skilled in the art to adopt the present disclosure in virtually any appropriate detailed structure differently.
[0020] In addition, the present disclosure may be described herein in terms of functional processing block components and various processing steps. It should be understood that such functional processing blocks may be implemented by any number of hardware and / or software components configured to perform the specified functions.
[0021] The scope of the present disclosure should be determined by the appended claims and their legal equivalents, rather than by the examples given herein. For example, the steps recited in any method claim may be performed in any order and are not limited to the order presented in the claim. Additionally, no element is essential for the practice of the present disclosure unless specifically described herein as "critical" or "necessary".
[0022] As used herein, "machine learning" or "data learning" is a technical term and may refer to computer or processor-related technologies by which decisions and / or actions are made, learned, and / or trained autonomously, without human intervention. Machine learning is a branch of artificial intelligence that focuses on using data and algorithms to mimic the way humans learn, gradually improving its accuracy. Machine learning may include software that supports machine learning, natural language understanding, natural language processing, speech recognition, computer vision, etc., i.e., algorithms and / or programs, hardware, or firmware, or any combination thereof. Also included within the scope of the machine learning functions and capabilities are the training and / or fine-tuning of machine learning models related to the embodiments disclosed, recited, and proposed herein.
[0023] As used herein, "model" or "machine learning model" is a technical term and may refer to software that supports machine learning, natural language understanding, natural language processing, speech recognition, computer vision, etc., such as algorithms and / or programs, hardware, or firmware, or any combination thereof. In an example embodiment, the process of training a model involves providing training data to a machine learning algorithm (e.g., a learning algorithm, etc.) for learning, and a machine learning model may refer to the model artifact created by the training process.
[0024] As used herein, "parameters" or "model parameters" of a model is a technical term and may refer to configuration variables within the model, the values of which may be estimated from given data. When making predictions, the model requires model parameters, and the model parameters can determine how the input data is transformed into the desired output. In an example embodiment, "weights" are model parameters that transform the input data within a (hidden) layer of the model and / or represent the connection strength between units or nodes of the model. In an example embodiment, "biases" are model parameters that represent the amount of difference between the predictions of the model and the target values compared to the training data.
[0025] As used herein, "forward" propagation, passing, or operation is a technical term and can refer to a function, operation, or algorithm that obtains or generates the actual output of a machine learning model. In an example embodiment, in a forward operation, input data can be fed into the model in a forward direction, e.g., by propagating the input data to an input layer, through (one or more) hidden layers and (one or more) successive layers, measuring the prediction of the model from an output layer, and calculating a model error based on the prediction made by the model. As used herein, "backward propagation" or "backward" propagation, passing, or operation is a technical term and can refer to a function, operation, or algorithm that traverses the model in the reverse order (from the output layer (through (one or more) hidden layers and (one or more) successive layers) to the input layer) and calculates the gradient with respect to the model parameters. In an example embodiment, in a backward operation, the flow (from the forward operation) is reversed by, e.g., propagating an error to the output layer until it reaches the input layer through (one or more) hidden layers. It should be understood that a "training step" or "fine-tuning step" is a technical term and can refer to a process that includes at least one forward operation and one backward operation based on a batch of input data.
[0026] As used herein, a "model execution step" can refer to one or a single training or fine-tuning step of executing the training of a machine learning model (using training data or non-training data), e.g., to collect or obtain desired or required information or data regarding the normal training execution of the model, such as, for example, the state of tensors and / or its modules, execution sequences or timestamps, execution phases (e.g., in a forward operation phase or a backward operation phase, etc.), hookable attributes, etc., the relationship between tensors and their modules, memory usage for each execution phase, etc. It should be understood that a model execution step is for collecting or obtaining desired or required information or data regarding the execution of the model, rather than for optimizing, training, or fine-tuning the model.
[0027] As used herein, an "optimizer" is a technical term and can refer to a function or algorithm that modifies the properties or parameters (e.g., weights, learning rates, etc.) of a machine learning process, method, or model. In an example embodiment, an optimizer can help reduce the overall loss and improve accuracy, minimize an error function (e.g., a loss function, etc.), and / or maximize productivity.
[0028] As used herein, the "survival" of an object or "survivability" can refer to the lifespan of an object, e.g., between the allocation timestamp and the deallocation timestamp of the object.
[0029] As used herein, a "long-lived object" may refer to an object that has a lifespan that transcends different phases or tiers. In some embodiments, a long-lived object may be an object that is allocated in the forward phase but deallocated only in the backward phase. In some embodiments, a long-lived object may be an object that is allocated in the backward phase but not deallocated in the same tier.
[0030] As used herein, a "normal object" may refer to an object that is not classified as a long-lived object. In some embodiments, a normal object may be an object that is allocated and deallocated in the same phase or tier.
[0031] As used herein, the term "object" is a technical term and may refer to one or more of data, variables, data structures, functions, code, etc., and / or may include one or more of values, references / identifiers, attributes, etc.
[0032] The features in the embodiments disclosed herein may reduce and / or optimize memory consumption by improving the management of memory usage and / or reduce the computational overhead of data learning operations to support data learning operations, e.g., operations included in training a machine learning model. The memory management mechanisms discussed herein may be used in deep learning architectures, e.g., for natural language processing, computer vision, audio processing, multimodal processing, generative pre-trained architectures, or directional encoder representations from transformers, where input data, model parameters, forward activations, backward gradients, optimizer states may use the processing device memory during the training of a machine learning model.
[0033] The features in the embodiments disclosed herein may include a memory management mechanism that includes an analysis-based memory allocation algorithm that utilizes information about allocated objects, which is pre-obtained through analysis, for machine learning parameters for training a machine learning model. Such a memory management mechanism can reduce memory consumption to address the problems of existing technical solutions by using analysis to collect at least the size, memory allocation, and memory deallocation of objects for training a machine learning model, e.g., via timestamps. The memory management mechanism may also include a scheduling step that uses the analyzed information to determine the size and address range of memory allocation, e.g., blocks in a machine learning model, and may then determine the total consumption of memory usage, e.g., GPU memory, which is optimized to reduce memory consumption. Then, the scheduling step can be applied to any subsequent allocation, e.g., an execution step or block for training a machine learning model. In this way, the memory allocation algorithm can not only include memory management that can be optimized by reducing memory consumption, e.g., by avoiding memory fragmentation, but also reduce computational overhead by applying the same allocation information or parameters to any subsequent allocation (e.g., execution step or block) for training a machine learning model. In some embodiments, the analysis-based memory allocation algorithm can be a parallel ladder algorithm that is designed, programmed, or otherwise configured to balance memory consumption and computational complexity, which results in less memory consumption compared to existing suboptimal algorithms.
[0034] Accordingly, the methods and systems discussed herein may have one or more of the following advantages:
[0035] Reduce computational overhead by leveraging the similarity between layers or blocks of a machine learning model (e.g., a model utilizing a transformer architecture), classifying objects into ordinary objects and long-lived objects (which may interfere with memory allocation), and grouping ordinary objects into small groups and long-lived objects into groups at once during the scheduling phase.
[0036] Reduce memory consumption by sorting objects based on allocation timestamps and deallocation timestamps to optimize the allocation of objects based on allocation / deallocation to maximize memory usage.
[0037] Reduce the occurrence of "out of memory" (OOM) and / or memory waste by using a suboptimal memory allocation algorithm that is designed, programmed, or otherwise configured to balance computational overhead and memory consumption, e.g., by managing objects according to memory blocks and releasing memory blocks back to the processing device as needed for optimized memory allocation (or usage).
[0038] Avoid memory fragmentation by reusing (releasing) memory as a contiguous memory space to maximize memory availability (e.g., for larger object sizes). For example, if contiguous memory space is not available, once the first object is deallocated, its space may not be immediately reused if adjacent objects are still active (e.g., memory fragmentation).
[0039] Balance computational overhead and memory consumption by using a suboptimal memory allocation algorithm or process that applies allocation information for one block to other blocks to train a machine learning model, e.g., for bulk allocations where the size and address of each allocation need not be determined.
[0040] Figure 1 Is a schematic diagram of an example memory management system in a data learning system 100 arranged according to at least some embodiments described herein.
[0041] System 100 may include devices 110, 120, 130, 140, 150 and network 160. It should be understood that Figure 1 Only an illustrative number of devices and / or networks are shown. The embodiments described herein are not limited to the number of devices and / or networks described. That is, the number of devices and / or networks described herein is for illustrative purposes only and is not intended to be limiting.
[0042] According to at least some example embodiments, devices 110, 120, 130, 140 and 150 may be various electronic devices. The various electronic devices may include, but are not limited to, mobile devices such as smart phones, tablet computers, e - book readers, laptop computers, desktop computers, servers, and / or any other suitable electronic devices.
[0043] According to at least some example embodiments, network 160 may be a medium for providing a communication link between devices 110, 120, 130, 140 and 150. Network 160 may be the Internet, a local area network (LAN), a wide area network (WAN), a local interconnect network (LIN), the cloud, etc. Network 160 may be implemented through various types of connections such as wired communication links, wireless communication links, fiber optic cables, etc.
[0044] According to at least some example embodiments, one or more of devices 110, 120, 130, 140 and 150 may be a server for providing various services to a user using one or more of the other devices. The server may be implemented by a distributed server cluster including multiple servers, or may be implemented by a single server.
[0045] Users can interact with each other via network 160 using one or more of devices 110, 120, 130, 140, and 150. Various applications or their localization interfaces, such as social media applications, online shopping services, data set operation services, machine learning services, etc., can be installed on devices 110, 120, 130, 140, and 150.
[0046] It should be understood that software applications or services according to the embodiments described herein and / or according to the services provided by service providers can be executed by devices 110, 120, 130, 140, and 150. Therefore, the apparatus for software applications and / or services can be arranged in devices 110, 120, 130, 140, and 150.
[0047] It should also be understood that when the service is not remotely executed, system 100 may not include network 160 and only include devices 110, 120, 130, 140, and / or 150.
[0048] Furthermore, it should be understood that devices 110, 120, 130, 140, and 150 may each include one or more processors, a memory, and a storage device storing one or more programs. Devices 110, 120, 130, 140, and / or 150 may also each include an Ethernet connector, a Wi-Fi receiver, etc. When executed by one or more processors, the one or more programs may cause the one or more processors to execute the methods described in any of the embodiments described herein. It should also be understood that a computer-readable non-volatile medium can be provided according to the embodiments described herein. The computer-readable medium stores a computer program. When the computer program is executed by a processor, it is used to execute the methods described in any of the embodiments described herein.
[0049] Furthermore, it should be understood that in the embodiments described herein, a device may refer to a computer system (e.g., 110, 120, 130, 140, 150, etc.) including at least one CPU, one GPU, and / or a combination thereof (see also Figure 4 for the description).
[0050] Figure 2 is a schematic diagram of an example processing flow 200 for managing memory usage in data learning operations (e.g., for training a machine learning model or operations therein) arranged according to at least some of the embodiments described herein.
[0051] It should be understood that a training model may refer to learning or determining desired or optimal model parameters (e.g., weights, biases, etc.) based on training data (e.g., based on one or more objects or a batch of objects for one or more data assignments). A fine-tuning model may refer to a method of transfer learning in which model parameters (e.g., weights, etc.) of a pre-trained model are trained on new training data (e.g., one or more objects or a batch of objects). Optimizing a model may refer to training and / or fine-tuning a model.
[0052] It should be understood that, unless otherwise specified, the processing flow 200 disclosed herein may be executed by one or more processors (e.g., Figure 1 the processors of one or more of the devices 110, 120, 130, 140, and 150, Figure 8 the CPU or GPU 805, and / or any other suitable processor).
[0053] It should also be understood that the processing flow 200 may include one or more operations, actions, or functions illustrated by one or more of the processing blocks 205, 220, 222, 224, 226, and 230. These various operations, functions, or actions may, for example, correspond to software, program code, or program instructions executable by a processor that causes the functions to be performed. Although illustrated as discrete processing blocks, obvious modifications may be made, e.g., two or more of the processing blocks may be reordered; further processing blocks may be added; and the various blocks may be divided into additional processing blocks, combined into fewer processing blocks, or eliminated, depending on the desired implementation. It should be understood that operations including initialization, etc. may be performed prior to the processing flow 200. For example, system parameters and / or application parameters may be initialized.
[0054] The processing flow 200 may include a memory allocation algorithm based on analysis, such as a parallel ladder algorithm, which is designed, programmed, or otherwise configured to balance memory consumption and computational complexity. In some embodiments, the processing flow 200 may include an analysis step or phase 205, a scheduling step or phase 220, and an execution step or phase 230. During the analysis step or phase 205, the processing flow 200 may collect the sizes, lifetimes, and the order of memory allocation and deallocation of one or more objects or a batch of objects. The lifetime of an object may be based on when the object is allocated and deallocated, e.g., based on timestamps and how long the object remains in memory. During the scheduling step or phase 220, the processing flow 200 may schedule the memory usage for allocation, and determine the size and address range of each memory allocation (e.g., a group of objects provided for the training of a block (or allocation) of a machine learning model), and subsequently determine the total consumption of memory for that block (or allocation), e.g., based on the total memory allocation of one or more objects or a batch of objects. As used herein, the term "block" may be a technical term referring to a neural network block, e.g., a transformer block, which may be a single layer or multiple layers in a machine learning model. During the execution step or phase 230, the processing flow 200 may apply the scheduling of the memory allocation determined at 220, e.g., the allocation information, and apply it to subsequent execution phases, e.g., one or more blocks of a machine learning algorithm, e.g., subsequent transformer building blocks, which may be scaled dot product attention units. In some embodiments, during the entire training process of a machine learning model, only a single analysis step or phase 205 and a single scheduling step or phase 220 may be provided, since the same amount of memory may be used in subsequent phases or blocks, e.g., because the memory usage may be deterministic. Thus, the analysis step or phase 205 and the single scheduling step or phase 220 may be utilized or applied to any subsequent block or layer of the training of the machine learning model, e.g., a transformer block, e.g., during the execution step or phase 230. For example, when a machine learning model is based on a transformer architecture (e.g., having an encoder and a decoder, each encoder and decoder containing multiple blocks, where each block includes a self-attention mechanism followed by a feed-forward neural network), the processing flow 200 is designed, programmed, or otherwise configured to use only the allocation information or pattern of the first block in the forward step or phase and the first block in the backward step or phase determined in the scheduling step or phase 220 for any other block (or layer) in the transformer architecture. That is, since in some embodiments, different blocks in a machine learning model may have the same allocation and / or deallocation pattern, the allocation information or pattern of a single block (e.g., from analysis and scheduling) may be used or applied to the scheduling of other blocks to reduce the number of allocations in the scheduling step or phase.
[0055] That is, in some embodiments, the processing flow 200 with an analytics-based memory allocation algorithm can be designed, programmed, or otherwise configured to reduce computational overhead and / or reduce memory consumption. For example, in some embodiments, computational overhead can be reduced by leveraging similarities between layers of a machine learning model (e.g., by leveraging or applying allocation information from a first block) to reduce the number of allocations in the scheduling step or phase 220. For example, for a machine learning model with 100 blocks (e.g., transformer blocks), the number of allocations can be reduced by approximately 198 times, e.g., 99 * 2, compared to a machine learning model that considers all allocations (or blocks) (in the forward and backward phases or steps). Additionally, computational overhead can be reduced by considering object classification between normal objects and long-lived objects, where only normal objects can be considered in the object map, but normal objects and long-lived objects are grouped and scheduled separately using an analytics-based memory allocation algorithm. Further, normal objects can be divided into small groups, e.g., based on allocation time, and small groups of objects are considered at once in the scheduling step or phase 220, as discussed further below.
[0056] Further, in some embodiments, memory consumption can be reduced by having the processing flow 200 with an analytics-based memory allocation algorithm be designed, programmed, or otherwise configured to arrange and / or rearrange objects into one or more groups to reduce memory consumption caused by allocation and deallocation, as discussed further below.
[0057] Return reference Figure 2 , at the processing block 205 (analytics), the processor can be designed, programmed, or otherwise configured to perform an analysis step or phase on the machine learning model. In an example embodiment, the analysis step or phase 205 can include an analysis of a single training or fine-tuning step that performs at least a forward operation and / or a backward operation on a single block of the machine learning model (e.g., a single allocation or transformer block). It should also be understood that the analysis results from the analysis step or phase (described in detail below) can include the same or substantially the same allocation information for all blocks or layers of the machine learning model, e.g., any remaining or subsequent blocks after a first training step.
[0058] It should be understood that the analysis step or phase (processing block 205) of the processing flow 200 is to collect or obtain detailed information or data of each object (e.g., one or more objects or a batch of objects) allocated in the first training step. The information or data collected or obtained (e.g., the analysis results from the analysis execution) can be used to guide memory allocation (and distribution) during the scheduling step or phase 220. For example, by grouping one or more objects into one or more groups, and / or guiding the placement or arrangement of one or more objects within the GPU memory or CPU memory of the device based on memory address ranges, allocation timestamps, deallocation timestamps, etc.
[0059] In an example embodiment, the information or data collected or obtained in the analysis step or phase 205 may include one or more of object sizes (e.g., the object size of each object in one or more objects or a batch of objects); object allocation timestamps and deallocation timestamps (e.g., execution sequences), for example, to determine whether there can be any object overlaps (conflicts) during the execution step or phase; the lifespan of the object, e.g., ordinary objects or long-lived objects; and / or the phase (forward or backward, and the model layer being executed) when the object is allocated or deallocated. In some embodiments, the information or data collected or obtained can be collected or obtained during the analysis step or phase 205 through one or more of the following: intercepting each memory allocation / deallocation call, e.g., in / for the memory allocator; collecting the allocation and deallocation timestamps of each object using a global counting variable (e.g., incrementing the global variable at each call or event); or collecting phase information (forward / backward) by using hooks or hook functions for the forward and backward phases (e.g., functions, operations, or algorithms that are executed or triggered when a condition is met, for example). The processing can proceed from processing block 205 to processing block 220.
[0060] It should be understood that the analysis step or phase 205 can be executed on the GPU or CPU. Executing the analysis step or phase 205 on the CPU may take longer than executing it on the GPU, and it may be difficult to determine the memory usage (e.g., increase, etc.) of the forward operation phase and the backward operation phase, e.g., for normal model execution.
[0061] At processing block 220 (scheduling), the processor (of each device) can be designed, programmed, or otherwise configured to perform a scheduling step or phase 220 on a machine learning model. In an example embodiment, a processing flow 200 with an analysis-based memory allocation algorithm can include scheduling to determine the total size of the memory to be used, the address range for each object, and / or object allocation / scheduling for execution. The scheduling step or phase 220 can include as inputs the size of the object, the allocation timestamp and deallocation timestamp of the object, the lifespan of the object, and / or object classification. In one embodiment, the scheduling step or phase 220 can include one or more of object classification processing block 222, e.g., classifying objects of the same block into two or more types, memory management processing block 224, for grouping, allocating, and / or scheduling objects based on the allocation timestamp and deallocation timestamp, and object mapping processing block 226, for object mapping of objects located in other blocks of the machine learning model (e.g., other transformer blocks).
[0062] At processing block 222 (object classification), a processor with an analysis-based memory allocation algorithm can be designed, programmed, or otherwise configured to classify objects into two or more types, including but not limited to long-lived objects and normal objects. In some embodiments, long-lived objects and normal objects can be grouped separately, e.g., in separate memory spaces. In some embodiments, long-lived objects can be moved or placed in different memory regions, e.g., not in the same memory region (or used) as normal objects. In some embodiments, long-lived objects can be pooled separately to avoid interfering with the memory allocation (and / or deallocation) of normal objects. In some embodiments, long-lived objects can be objects that are allocated in the forward phase but only deallocated in the backward phase, e.g., objects for gradient checkpointing, or objects that are allocated in the backward phase but not deallocated in the same layer. In some embodiments, normal objects are objects that are not classified as long-lived objects and / or objects that are allocated and deallocated in the same phase or layer. Processing can proceed from processing block 222 to processing block 224.
[0063] At processing block 224 (object allocation), the processor can be designed, programmed, or otherwise configured to utilize or apply an analytics-based memory allocation algorithm for object allocation to manage memory usage. Object allocation can include grouped long-lived objects and grouped normal objects and / or include further grouping of normal objects, e.g., based on the liveness of the last allocated object (e.g., based on allocation timestamp and deallocation timestamp), grouping one or more objects or a batch of objects from a single block or tier, sorting the objects first according to the allocation order (e.g., based on allocation timestamp), and then moving, arranging, or swapping the order of the objects based on the deallocation order (e.g., based on deallocation timestamp), as further discussed below.
[0064] During the grouping of normal objects, the objects are placed, grouped, or arranged into one or more groups, where objects with overlapping lifetimes (e.g., based on allocation timestamp and deallocation timestamp) are placed, grouped, or arranged in the same group because objects with overlapping lifetimes cannot be allocated simultaneously. That is, based on the allocation timestamp, all live objects (e.g., having overlapping lifetimes based on allocation timestamp and deallocation timestamp) will be placed in the same group. The grouping can be repeated until all objects are grouped into a series of groups, e.g., G0 (youngest group), G1, …, GN (oldest group). In some embodiments, the maximum size of the memory can be determined based on the maximum size of all groups (e.g., the memory addresses occupied by all groups).
[0065] After grouping, the processor can be designed, programmed, or otherwise configured to sort, arrange, move, or place the objects in each group. In some embodiments, the sorting of the objects can start from the earliest group, which can ensure that the allocation (or scheduling) of later groups does not conflict with the existing groups. In an example embodiment, the objects can be sorted in descending order based on their allocation timestamp, e.g., from the latest one to the earliest one, and placed in the memory space for scheduling, where their memory addresses are provided from low to high. Such a configuration can allow the analytics-based memory allocation algorithm to adapt to a stepped configuration for allocation, e.g., to maximize memory usage based on object allocation / deallocation.
[0066] In some embodiments, the processor may be designed, programmed, or otherwise configured to move, arrange, or swap the order of objects based on the deallocation timestamp. In one embodiment, objects within the same group may be moved, arranged, or swapped according to the deallocation timestamp if such movement, arrangement, or swap does not introduce a conflict with the objects of the previous group. For example, in some embodiments, for any pair of objects A and B, if A is placed to the right of B but A is deallocated later than B, then A may be swapped with B if the swap does not conflict (e.g., for allocation / deallocation), such as swapping with an object in the last group. Additionally, in some embodiments, after the movement, arrangement, or swap of objects, the memory offset gap may be reduced, moved, or eliminated, where each object may be pushed to the leftmost side (if possible), e.g., based on the memory availability in the memory space at the appropriate allocation time, reducing, removing, or decreasing the memory offset in the memory address range in the memory space (for allocation or scheduling of memory usage). Such a configuration may allow the memory allocation algorithm based on analysis to reduce the external memory fragmentation caused by deallocation, e.g., by maximizing memory usage to minimize the combined memory address range of objects.
[0067] Thus, based on the arrangement (and / or rearrangement) of objects in the memory space (for allocation or scheduling), including the processing flow 200 of the memory allocation algorithm based on analysis, it may be designed, programmed, or otherwise configured to determine the total size of the memory of all objects (e.g., one or more objects or a batch of objects), e.g., as the offset between the leftmost and rightmost objects in the group, e.g., based on the memory address positions of the objects in the memory space, and the scheduling (or allocation) of the allocation and deallocation of the objects, e.g., based on the allocation timestamp and the deallocation timestamp. After determining the total size of the memory, the processing flow 200 may then determine the address range of each object, where the object at the leftmost address may be set to 0, which may be maximized for the memory space of the object, e.g., based on the allocation and deallocation timestamps and the memory usage. The processing may proceed from processing block 224 to processing block 226.
[0068] At processing block 226 (object mapping), the processor may be designed, programmed, or otherwise configured to apply or utilize the same allocation rule (or information) for one or more objects or groups of objects for, e.g., the first block (or allocation) of block 0 of the machine learning model in the forward phase to any remaining blocks, e.g., blocks 1 to block (N), and / or apply or utilize the first block for, e.g., block (N) to any remaining blocks, e.g., blocks (N - 1) to block 0, in the reverse phase.
[0069] For example, Figure 3A schematic diagram showing an embodiment of an object mapping process that applies or utilizes the same allocation rule (or information) to remaining blocks in a machine learning model is presented. In one embodiment, the object mapping process may include an analysis step or phase 305 that can be performed for the allocation of one training step or a batch of data, and may include receiving or inputting a data batch block 302 of one or more objects or a batch of objects, a forward phase 310, a backward phase 315, and an optimizer 319. The data batch block 302 may be a batch of training data for a single block. A batch of training data may include one or more objects or a batch of objects, which may be the input or data for training the machine learning model collected or obtained by a processing stream (e.g., 200). In some embodiments, each transformer block may include 100 objects, 1000 objects, 10000 objects, 100000 objects, etc., but is not limited thereto. Instead, the number of objects depends on the number of objects in one step (or allocation) of training the machine learning model.
[0070] The forward phase 310 may be designed, programmed, or otherwise configured to obtain or generate the actual output of the machine learning model. In an example embodiment, in the forward phase 310, the data batch 302 may be fed into the machine learning model in the forward direction, e.g., by propagating the input data to the input layer, through the (multiple) hidden layers and (multiple) consecutive layers, measuring the prediction of the model from the output layer, and calculating the model error based on the prediction made by the model. The forward phase 310 may include a wte / wpe layer 311, blocks 312 0…N, and an ln_f / wte layer 313. The wte / wpe layer 311 may be, for example, a token and / or positional embedding layer for a transformer block architecture that receives the data batch and encodes the relationship between tokens and positions. Blocks 0…N may be transformer blocks that may include encoder / decoder blocks for processing the input data (e.g., objects in the data batch) of the machine learning model. It should be understood that as the machine learning model becomes more complex and larger (e.g., including millions, hundreds of millions, billions, etc., of parameters), a larger number of blocks may be used to train the model, e.g., 50 blocks, 100 blocks, 500 blocks, 1000 blocks, etc. The ln_f / wte layer 313 may be a normalization layer to normalize the input of each feature (e.g., the parameters of the machine learning model).
[0071] The backward phase 315 may be designed, programmed, or otherwise configured to receive the normalized input from the forward phase 310. In an example embodiment, during the backward phase, the machine learning model traverses in the reverse order, from the output layer (through the (multiple) hidden layers and (multiple) consecutive layers) to the input layer, and calculates the gradient with respect to the model parameters. In an example embodiment, in the backward phase, the process is reversed (from the forward operation) by, for example, propagating the error to the output layer until it reaches the input layer through the (multiple) hidden layers.
[0072] Optimizer 319 can be designed, programmed, or otherwise configured to receive results from the reverse phase 315 (and / or the forward phase 310) and modify attributes or parameters of the machine learning model (e.g., weights, learning rate, etc.). In an example embodiment, the optimizer can help reduce the overall loss and improve accuracy, minimize an error function (e.g., a loss function, etc.), and / or maximize productivity.
[0073] As Figure 3 shown, a processing flow (e.g., 200) can be designed, programmed, or otherwise configured to apply or utilize the same allocation rule (or information) to one or more objects or groups of objects of a first block (e.g., block 0) of a machine learning model, e.g., blocks 1 to block (N) in the forward phase and / or the first block (e.g., block (N)) to any remaining blocks (e.g., block (N-1) to block 0) in the reverse phase. That is, the processing flow 200 can be designed, programmed, or otherwise configured to utilize similarities (e.g., allocation information) between different blocks (or layers) in the machine learning model for common objects in other blocks (or layers). Processing can proceed from processing block 226 to processing block 230.
[0074] At processing block 230 (execution phase), the processor can be designed, programmed, or otherwise configured to apply or execute a schedule at a subsequent execution step or phase 230. In some embodiments, the execution phase can use the same address range for all objects determined during the analysis step or phase 205, where the analysis and scheduling are performed only for a batch of data in a first step (or one allocation), e.g., for block 0 in the forward phase or block (N) in the reverse phase. That is, for other training iterations of the machine learning model, the processor can be designed, programmed, or otherwise configured to follow all of the following processes of allocation and deallocation:
[0075] For each allocation request, the processor can be designed, programmed, or otherwise configured to use the phase, size, and lifetime (e.g., viability based on an allocation timestamp and a deallocation timestamp) as an identifier such that objects in subsequent blocks (or allocations) can be mapped to corresponding objects determined during the analysis step or phase and the scheduling. That is, the starting address of the allocation will use the same address determined in the analysis phase. For each deallocation request, no explicit operation is required because the content of any object will be overwritten by subsequent objects.
[0076] In this way, the processing flow 200 can be designed, programmed, or otherwise configured to balance memory consumption and computational complexity. That is, in some embodiments, the processing flow 200 can include an analysis-based memory allocation algorithm to reduce computational overhead and / or reduce memory consumption. For example, in some embodiments, computational overhead can be reduced by leveraging the similarity between layers of a machine learning model (e.g., by leveraging or applying allocation information from a first block) to reduce the number of allocations in the scheduling step or phase 220. Additionally, by considering object classification between normal objects and long-lived objects in the analysis-based memory allocation algorithm, computational overhead can be reduced, where only normal objects are object mapped, and normal objects are grouped into small object groups, and long-lived objects are grouped into object groups and scheduled accordingly. Furthermore, normal objects can be further divided into smaller groups, e.g., based on allocation time, and small object groups are considered one at a time in the scheduling step or phase 220 to reduce computational complexity and maximize memory usage by reducing memory consumption (e.g., minimizing the combined memory address range between the first allocation timestamp and the last deallocation timestamp).
[0077] Further details of processing flows with analysis-based memory allocation algorithms and memory usage scheduling are discussed below.
[0078] Figure 4 An example embodiment of a storage space 400 for a first group of objects 402 and a second group of objects 403 arranged according to a processing flow with an analysis-based memory allocation algorithm 400 (e.g., a parallel ladder algorithm) as discussed herein is shown. The first group of objects 402 includes objects arranged first based on allocation timestamps, e.g., the top sides of objects 402a, 402b, 402c, 402d indicate the allocation times of the corresponding objects, and then can be re-arranged or moved based on deallocation timestamps, e.g., the bottom sides of objects 402a, 402b, 402c, 402d indicate the deallocation times of the corresponding objects, e.g., along the object lifetime. The objects can be placed, arranged, or moved into the memory space of a processing device along a memory address range (e.g., with a memory offset indicating the distance (or displacement) between the first and last objects of the group).
[0079] That is, objects in the same group (e.g., the first group or the second group) can first be sorted by their respective allocation timestamps, where the earliest one (according to the allocation timestamp) will be placed in the rightmost position, e.g., at a memory address with the largest memory offset, and so on. Then, as long as re-sorting, moving, swapping, or re-arranging does not conflict with the placement of objects in any previous group, the objects can be sorted (e.g., re-sorted, moved, swapped, or re-arranged) from the right by their respective deallocation timestamps.
[0080] For example, based on the allocation timestamp, for the first group of objects 402 (e.g., a group of objects allocated first or before another group of objects), an object 402d of size 28 (KB, MB, GB, etc.) can be allocated first at the rightmost position (e.g., in the memory space with the largest memory offset) because the object 402d has the earliest allocation timestamp. Then an object 402c of size 36 can be arranged, moved, or placed near the left of the object 402d (in descending order of memory offset), and then an object 402b of size 8 and an object 402a of size 16 can be arranged, moved, or placed to the left of the object 402c in descending order of the allocation timestamp. Next, the objects 402a and 402b can be considered for rearrangement, movement, swapping, or placement based on the deallocation timestamp. In one embodiment, since the deallocation time of the object 402b is before the deallocation time of the object 402a, the object 402b can be placed to the right of the object 402a, e.g., in a memory space with a larger memory offset, so that the memory can be freed or released for subsequent use in contiguous memory.
[0081] In this way, as Figure 4 schematically shown, the objects in the first group can be organized as a ladder 404a (based on the allocation / deallocation timestamp as described above), while the objects in the second group 403 or the next group can be organized as a ladder 404b (based on the allocation timestamp (and / or deallocation timestamp)). For example, in some embodiments, since the object 403c of the second group 403 has an earlier allocation timestamp, the object 403c is arranged, moved, or placed at a position in the memory space along the memory address range at the associated allocation time, which is farthest to the rightmost position (e.g., the largest memory offset). Subsequently, the object 403b is arranged, moved, or placed at a position in the memory space along the memory address range corresponding to the allocation time of the object 403b, which maximizes the memory usage. Then, the object 403a is arranged, moved, or placed at a position in the memory space along the memory address range at the associated allocation time, where if there is a gap in the memory address range (e.g., a distance or value in the memory offset), the object 403a is shifted along the memory address range to remove (or reduce) any memory offset, e.g., corresponding to position 0. Thus, by arranging, rearranging, placing, or moving objects in a group based on the allocation and deallocation timestamps, compared with existing mechanisms for maximizing memory usage, less memory can be utilized, for example, by reducing the memory address range (e.g., the total size) of the memory space 400, for scheduling steps or phases (or memory usage allocation), and / or forming one or more contiguous memory blocks for reuse.
[0082] For example, Figure 5A and 5B schematically shows a memory space 500 including objects with overlapping lifetimes, e.g., objects that survive during concurrent allocation / deallocation times, which do not have a memory space 500 that is optimized or maximized for memory usage. As Figure 5A shown, in an example embodiment, three objects 502a, 502b, 502c can be arranged, moved, or placed within the memory space 500 along a memory address range and a time range for memory usage, e.g., scheduling / allocation. Object 502a can have a size of 36 bytes, object 502b can have a size of 10 bytes, and object 502c can have a size of 36 bytes. If the 10-byte object 502b is placed between the 36-byte objects 502a and 502c, then after object deallocation, the memory space is split into two ranges until the 10 bytes of object 502b are deallocated. Thus, when there is an object 502d (e.g., an object request with 72 bytes), as Figure 5B schematically shown, only the memory space of the 36-byte object 502c can be reused, and the memory space of the 36-byte object 502a cannot be reused until object 502b is allocated / deallocated. Therefore, since Figure 5A and 5B the memory space 500 can have separate memory ranges where one of the memory addresses in the memory address may not be accessed or reused, thus potentially forming memory fragmentation, which increases the memory address range of the object.
[0083] On the other hand, as Figure 6A and 6B schematically shown, when an analysis-based memory allocation algorithm (e.g., parallel staircase algorithm) as discussed herein is applied to arrange, rearrange, move, or place objects within a memory space 600, memory fragmentation can be avoided or reduced. For example, in one embodiment, as Figure 6A shown, if object 602b (which can correspond to Figure 5A and 5BIf the objects 502b) are rearranged, moved, or swapped such that objects with similar deallocation timestamps (e.g., objects 602a and 602c) are located adjacent to each other, then contiguous memory ranges can be made available. That is, by using or applying a process flow with an analysis-based memory allocation algorithm to rearrange, move, or swap objects in the memory space, after objects 602a and 602c are deallocated, contiguous memory can be allocated for subsequent objects (e.g., for object 602d that requires a 72-byte contiguous memory range). Thus, different from the Figure 5A , 5B embodiments, the availability of contiguous memory addresses can be made available for reuse in the memory space 600, e.g., for the reuse of object 602d and 72-byte contiguous / continuous memory, which can avoid memory fragmentation even if object 602b is not deallocated. Further embodiments are discussed below.
[0084] Figures 7A - 7I Schematically shows a method for managing memory usage by using an analysis-based memory allocation algorithm according to an example embodiment, where the scheduling or allocation of the memory space includes a memory offset (X-axis), where a memory address range can be determined, and a timestamp (Y-axis), where the start / end positions of allocation / deallocation are schematically shown in the vertical direction.
[0085] Initially, as Figure 7AAs shown, the method may include receiving a single or a first training step or a batch of objects for a data learning operation (e.g., a forward operation and / or a backward operation). In one embodiment, a batch of objects may include eight objects 702a, 702b, 702c, 702d, 702e, 702f, 702g, 702h. The method may further include an analysis step to collect or obtain detailed information or data for each object 702a, 702b, 702c, 702d, 702e, 702f, 702g, 702h (e.g., one or more objects in the batch of objects) assigned in the first training step. The information or data collected or obtained (e.g., the analysis result from the analysis performed) may be used to guide memory allocation (and deallocation) during a scheduling step or phase, e.g., by grouping one or more objects into one or more groups, and / or guiding the placement or arrangement of one or more objects within a memory space for GPU memory or within the CPU memory of the device based on a memory address range, an allocation timestamp, a deallocation timestamp, etc. for the scheduling (or allocation) of the objects. In an example embodiment, the information or data collected or obtained in the analysis step or phase may include one or more of object sizes (e.g., the object size of each object in one or more objects or a batch of objects); object allocation timestamps and deallocation timestamps (e.g., execution sequence), e.g., to determine if there can be any object overlap (conflict) during an execution step or phase; the lifespan of an object, e.g., a normal object or a long-lived object; and / or the phase (forward or backward, and the model layer being executed) when an object is allocated or deallocated.
[0086] In one embodiment, the analysis result from the analysis step or phase may include object 702a having a size of 32 bytes, an allocation timestamp of 0, and a deallocation timestamp of 1, object 702b having a size of 28 bytes, an allocation timestamp of 1, and a deallocation timestamp of 4, object 702c having a size of 36 bytes, an allocation timestamp of 2, and a deallocation timestamp of 5, object 702d having a size of 16 bytes, an allocation timestamp of 3, and a deallocation timestamp of 6, object 702e having a size of 8 bytes, an allocation timestamp of 4, and a deallocation timestamp of 5, object 702f having a size of 64 bytes, an allocation timestamp of 5, and a deallocation timestamp of 7, object 702g having a size of 10 bytes, an allocation timestamp of 7, and a deallocation timestamp of 8, and object 702h having a size of 40 bytes, an allocation timestamp of 6, and a deallocation timestamp of 8.
[0087] In some embodiments, an analysis-based memory allocation algorithm is designed, programmed, or otherwise configured to arrange objects 702a, 702b, 702c, 702d, 702e, 702f, 702g, 702h in a memory space 700 for scheduling (or allocation), as Figure 7B shown, to maximize memory usage (or reduce memory consumption) while reducing computational overhead, where the resulting allocation information (e.g., scheduling or allocation of objects) can further reduce computational overhead by reducing the number of allocations in the scheduling steps or phases.
[0088] As Figure 7C shown, at an initial step, an analysis-based memory allocation algorithm is designed, programmed, or otherwise configured to group objects based on the liveness of the last allocated object (e.g., the object with the most recent allocation timestamp). For example, in one embodiment, based on the allocation timestamp of object 702g, the liveness of other objects can be determined. That is, since object 702g has an allocation timestamp (e.g., at time 7), other objects with an overlap time of 7 between the allocation timestamp and the deallocation timestamp are determined. In this way, objects 702f and 702h are identified or determined to have liveness during the liveness of object 702g and are grouped together, e.g., group 703a. Similarly, based on the next allocated object, i.e., one or more objects with the most recent allocation timestamp after the removed object in the first group 703a (e.g., object 702e), which has an allocation timestamp at time 4, for example, other objects with an overlap time of 4 between the allocation timestamp and the deallocation timestamp are determined. In this way, objects 702b, 702c, and 702d are identified or determined to have liveness during the liveness of object 702e and are grouped together, e.g., group 703b. Finally, object 702a is grouped into group 703c because there are no remaining objects that are alive during the liveness of object 702a and have not been grouped yet. In this way, one or more objects (e.g., from a batch of objects or data batch) can be grouped into one or more groups based on the memory allocation timestamp such that the one or more objects are provided in descending order with respect to a memory address range, where the object with the earliest memory allocation timestamp (among the one or more groups) is placed, provided, or arranged more to the right (e.g., at the largest memory offset) to accommodate objects from the one or more groups of objects.
[0089] As Figure 7DAs shown, at the next step, the analyzed memory allocation algorithm is designed, programmed, or otherwise configured to arrange, move, or place objects into the memory space 700 for scheduling (or allocation). In one embodiment, the object groups 703a, 703b, 703c may be arranged, moved, or placed based on the earliest allocation timestamp. For example, since the object 702a of group 703c has the earliest allocation timestamp (e.g., time 0), group 703c may be arranged, moved, or placed by placing the object 702a at offset 0 in the memory space 700 (e.g., having an address range of [0, 32)), from timestamp 0 to 1. In some embodiments, the objects of a group are rearranged, moved, or swapped by considering the deallocation timestamp to maximize memory usage by reducing or minimizing the combined memory address range of the group. Since in this case, group 703c includes only one object, the object is not rearranged, moved, or swapped based on the deallocation timestamp.
[0090] Figure 7EIt is shown that an analysis-based memory allocation algorithm is designed, programmed, or otherwise configured to arrange, move, or place objects in the next set 703b in the memory space 700 in descending order based on the allocation timestamp, from a first value to a second value in the memory space 700. In one embodiment, objects (which may include two or more objects, e.g., objects 702b, 702c, 702d, 702e) may be arranged, moved, or placed in the memory space 700 based on their allocation timestamps. That is, object 702b may be arranged, moved, or placed in the rightmost position (e.g., the largest memory offset) in the memory space 700 to accommodate the group of objects in group 703b in the memory space at the first value, since object 702b has the earliest allocation timestamp (e.g., timestamps 1 to 4). Next, since object 702c has a later allocation timestamp, object 702c may be arranged, moved, or placed near or adjacent to the left of object 702b in the memory space at the second value from timestamp 2 to 5. Then object 702d may be arranged, moved, or placed near or adjacent to the left of object 702c in the memory space at the third value from timestamp 3 to 6 (since object 702d has a later allocation timestamp), and subsequently object 702e may be arranged, moved, or placed near or adjacent to the left of object 702d at the fourth value from timestamp 4 to 5 based on their respective timestamps (since object 702e has the most recent allocation timestamp). That is, after arranging the group of objects in the memory space based on the allocation timestamp, subsequent groups of one or more objects are arranged in the memory space in descending order from the first value (rightmost position) to the second value (leftmost position or at 0 memory offset), where the object with the earliest memory allocation timestamp is provided at the first value, and another object with a later memory allocation timestamp is provided at the second value.
[0091] In some embodiments, to maximize memory usage (or minimize the combined memory address range of the memory space), two or more objects of a group may be rearranged, moved, or swapped by considering the deallocation timestamp. In some embodiments, an analysis-based memory allocation algorithm may be designed, programmed, or otherwise configured to determine whether the deallocation of one object in a batch of objects occurs after another object in the batch of objects, and when deallocation occurs, rearrange the order of the arranged batch of objects such that the combined memory address range of the batch of objects is minimized between the first allocation timestamp and the last deallocation timestamp. For example, in an exemplary embodiment, as Figure 7F and 7GAs shown, since object 702d is placed to the right of object 702e but has a later deallocation timestamp, there may be memory fragmentations 704a and 704b. Thus, to reduce memory fragmentation by combining fragments 704a and 704b into contiguous memory, object 702d can be moved, positioned, or swapped with object 702e because moving, positioning, or swapping object 702e will not conflict with other (previous) objects in one or more other groups, e.g., based on memory space usage of allocation / deallocation timestamps and memory address ranges. Thus, object 702d can be moved, positioned, or swapped and placed to the left of object 702e and at offset 0 (at the second value), and object 702e can be placed at offset 16 to maximize memory usage, where memory fragmentations 704a, 704b are combined to form free space 704. Thus, objects 702b, 702c, 702d, 702e can be rearranged to minimize the size of the memory space (e.g., combined memory address ranges) between object 702b with the first (or earliest) allocation timestamp and object 702d with the last deallocation timestamp.
[0092] On the other hand, when not using an analysis-based memory allocation algorithm, as Figure 7H seen, object 702b can initially be placed next to object 702a, where memory fragmentation 704 can be provided or exist between object 702a and object 702b since object 702c cannot be placed in the memory space (due to a time conflict of object 702c, e.g., the memory size required at allocation time 2). Thus, object 702c can be arranged, moved, or placed after object 702b such that more memory of the memory space can be used compared to the memory usage method using an analysis-based memory allocation algorithm as discussed herein.
[0093] Finally, as Figure 7IAs shown, after the arrangement, movement, or placement of the objects in groups 703a, 703b, the remaining objects 702f, 702g, 702h in group 703c can be sorted as objects 702g, 702h, 702f based on the allocation timestamps of the objects. Each object can be pushed or moved to the leftmost (if possible) to maximize the memory address range of the objects in the memory space (with corresponding allocation timestamps and deallocation timestamps), such that object 702g is placed at offset 0 (at timestamps 7 to 8), object 702h is placed at offset 16 (at timestamps 6 to 8), and object 702f is placed at offset 60 (at timestamps 5 to 7). For example, where the memory usage may not be minimized based on the allocation timestamps of the objects. In some embodiments, since the objects 702f, 702g, 702h are already sorted in their deallocation order, the objects may not need to be moved, positioned, or swapped. Thus, since object 702f can be located in the memory space at the memory address range of [60, 125), the total memory size used by groups 703a, 703b, 703c is 124 bytes and can be arranged or stored in a 124 - byte memory block for scheduling memory usage and / or its allocation. Thus, the address range of each object can be determined as follows: object 702a: [0, 32), object 702b: [60, 88), object 702c: [24, 60), object 702d: [0, 16), object 702e: [16, 24), object 702f: [60, 124), object 702g: [0, 10), object 702h: [16, 56) in the corresponding timestamp range between 0 and 8.
[0094] Figure 8 is a schematic structural diagram of an example computer system 800 arranged according to at least some embodiments described herein and suitable for implementing an electronic device (e.g., Figure 1 one of the devices shown in Figure 8 ). It should be understood that
[0095] As shown, computer system 800 may include a central processing unit (CPU) or a graphics processing unit (GPU) 805. The CPU or GPU 805 can perform various operations and processes based on a program stored in a read - only memory (ROM) 810 or a program loaded from a storage device 840 into a random access memory (RAM) 815. The RAM 815 can also store various data and programs required for the operation of system 800. The CPU or GPU 805, ROM 810, and RAM 815 can be connected to each other via a bus 820. An input / output (I / O) interface 825 can also be connected to the bus 820.
[0096] Components connected to the I / O interface 825 may also include an input device 830, including a keyboard, a mouse, a digital pen, a graphics tablet, etc.; an output device 835, including a display, such as a liquid crystal display (LCD), a speaker, etc.; a storage device 840, including a hard disk, etc.; and a communication device 845, including a network interface card, such as a LAN card, a modem, etc. The communication device 845 can perform communication processing via a network such as the Internet, WAN, LAN, LIN, cloud, etc. In one embodiment, a drive 850 can also be connected to the I / O interface 825. A removable medium 855, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., can be installed on the drive 850 as needed, so that a computer program read from the removable medium 855 can be installed in the storage device 840.
[0097] It should be understood that the processes described in the flowchart (or process) and / or algorithm with reference Figures 2 - 7I can be implemented as a computer software program or hardware. A computer program product can include a computer program stored in a computer-readable non-volatile medium. The computer program includes program code for performing the methods shown in the flowchart and / or GUI. In this embodiment, the computer program can be downloaded and installed from a network via the communication device 845, and / or can be installed from the removable medium 855. When executed by a central processing unit (CPU) or a graphics processing unit (GPU) 805, the computer program can implement the above functions specified in the methods in the embodiments disclosed herein.
[0098] The features in the embodiments disclosed herein can support data learning operations by reducing and / or optimizing memory consumption and / or reducing the computational overhead of data learning operations (e.g., operations included in training a machine learning model), by improving the management of memory usage. The memory management mechanism discussed herein can be used in deep learning architectures, for example, for natural language processing, computer vision, audio processing, multi-modal processing, generative pre-training architectures, or directional encoder representations from transformers, where during the training of a machine learning model, input data, model parameters, forward activations, backward gradients, optimizer states can use the processing device memory.
[0099] Features in the embodiments disclosed herein may include a memory management mechanism that includes an analytics-based memory allocation algorithm that utilizes information about allocation objects (i.e., obtained in advance through analysis) to train machine learning parameters of a machine learning model. Such a memory management mechanism can reduce memory consumption by using analytics to collect at least the size, memory allocation, and memory deallocation (e.g., via timestamps) of objects used to train the machine learning model, thereby addressing the problems of existing solutions. The memory management mechanism may also include a scheduling step that uses the information from the analytics to determine the size and address range of a memory allocation (e.g., a block in a machine learning model), and may subsequently determine the total consumption of memory use (e.g., GPU memory), which is optimized to reduce memory consumption. The scheduling step can then be applied to any subsequent allocations, such as execution steps or blocks for training a machine learning model. Thus, the memory allocation algorithm can include not only memory management that can be optimized by reducing memory consumption (e.g., by avoiding memory fragmentation), but also reducing computational overhead by applying the same allocation information or parameters to a single allocation (or block) for any subsequent allocations, such as execution steps or blocks for training a machine learning model. In some embodiments, the analytics-based memory allocation algorithm can be a parallel ladder algorithm that is designed, programmed, or otherwise configured to balance memory consumption and computational complexity, which results in less memory consumption compared to existing suboptimal algorithms.
[0100] Thus, the methods and systems as discussed herein may have one or more of the following advantages:
[0101] Reducing computational overhead by leveraging similarities between layers or blocks of a machine learning model (e.g., a model utilizing a transformer architecture); classifying objects into normal objects and long-lived objects that may interfere with memory allocation; and grouping normal objects into small groups of objects at once during the scheduling phase, while grouping long-lived objects into separate groups during the scheduling phase.
[0102] Reducing memory consumption by sorting objects based on allocation timestamps and deallocation timestamps to optimize the allocation of objects based on allocation / deallocation, thereby maximizing memory utilization.
[0103] Reducing the occurrence of "out of memory" (OOM) and / or memory waste by using a suboptimal memory allocation algorithm that is designed, programmed, or otherwise configured to balance computational overhead and memory consumption, e.g., by managing objects according to memory blocks and releasing memory blocks back to the processing device as needed for optimized memory allocation.
[0104] Avoid memory fragmentation by reusing (releasing) memory as a contiguous memory space to maximize memory availability (e.g., for larger object sizes). For example, if contiguous memory space is not available, once the first object is deallocated, its space may not be immediately reused if the adjacent object is still active (e.g., memory fragmentation).
[0105] Balance computational overhead and memory consumption by using a suboptimal memory allocation algorithm or process that applies allocation information of one block to other blocks for training a machine learning model, e.g., for bulk allocation where the size and address of each allocation do not have to be determined.
[0106] It should be understood that the disclosed and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or combinations of one or more of them. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a substance composition that affects a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" includes all apparatus, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The apparatus may also include code that creates an execution environment for the relevant computer program in addition to the hardware, e.g., code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0107] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts in a markup language document), stored in a single file dedicated to the relevant program, or stored in multiple coordinated files (e.g., files that hold one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0108] The processes and logical flows described in this document can be performed by one or more programmable processors that execute one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, e.g., a field programmable gate array, an application specific integrated circuit, etc.
[0109] By way of example, processors suitable for the execution of a computer program include both general and special purpose microprocessors, and any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto - optical disks, or optical disks, or operatively coupled to receive data therefrom or transfer data thereto, or both. However, a computer need not have such devices. Computer - readable media suitable for storing computer program instructions and data include all forms of non - volatile memory, media and memory devices, including by way of example semiconductor memory devices, such as erasable programmable read only memory, electrically erasable programmable read only memory, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto - optical disks; and compact disc read only memory and digital video disc read only memory discs. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.
[0110] It should be understood that different features, variations, and multiple different embodiments have been shown and described in various details. What is sometimes described in this application in terms of specific embodiments is for illustrative purposes only and is not intended to limit or imply that the subject matter conceived is only a particular embodiment or specific embodiment. It should be understood that the present disclosure is not limited to any single specific embodiment or recited variation. Many modifications, variations, and other embodiments will occur to those skilled in the art, and these modifications, variations, and other embodiments are intended to and are in fact covered by the present disclosure. Indeed, the scope of the present disclosure should be determined by appropriate legal interpretation and construction of the present disclosure, including equivalents as understood by those skilled in the art as of the time of filing in reliance on the complete disclosure.
[0111] Aspects:
[0112] It should be understood that any of the aspects can be combined with each other.
[0113] Aspect 1. A method for managing memory usage in a data learning operation, the method comprising: analyzing one or more objects trained in the data learning operation, wherein the analysis includes determining an object size, a memory allocation timestamp, and a memory deallocation timestamp of the one or more objects; and scheduling the memory usage for the one or more objects to determine a total size and an address range of one or more groups of the one or more objects, wherein the scheduling includes: grouping the one or more objects into the one or more groups based on the memory allocation timestamp and / or the memory deallocation timestamp of the one or more objects, and arranging the one or more objects in the one or more groups in descending order in a memory space, wherein in a case where the one or more objects include two or more objects, the two or more objects are provided in the memory space in descending order from a first value to a second value less than the first value, wherein an object having an earliest memory allocation timestamp among the two or more objects is provided at the first value, and another object having a later memory allocation timestamp among the two or more objects is provided at the second value.
[0114] Aspect 2. The method according to aspect 1, further comprising training an operation in the data learning operation based on the scheduling of the memory usage.
[0115] Aspect 3. The method according to any one of aspects 1-2, wherein the scheduling of the memory usage further includes rearranging the two or more objects in the one or more groups based on the deallocation timestamp of the two or more objects such that an object having a later deallocation time among the two objects is moved to maximize the memory usage.
[0116] Aspect 4. The method according to aspect 3, wherein one of the two objects is moved to eliminate or reduce any gap in the memory offset.
[0117] Aspect 5. The method according to any one of aspects 1-4, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs in a forward phase of the training of the data learning operation, and the analysis is performed only on a first block of the forward phase, and wherein allocation information for the one or more blocks during the scheduling is determined for the first block of the forward phase and applied to any remaining blocks of the forward phase for the training of the data learning operation.
[0118] Aspect 6. The method according to any one of aspects 1-5, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs in the backward phase of the training of the data learning operation, and the analysis is performed only on the first block of the backward phase, and wherein the allocation information for the one or more blocks during the scheduling is determined for the first block of the backward phase and applied to any remaining blocks of the backward phase for the training of the data learning operation.
[0119] Aspect 7. The method according to any one of aspects 5-6, wherein during the training of the data learning operation, for each allocation request, the size, the memory allocation timestamp, and the memory deallocation timestamp of the one or more objects determined during the analysis are used to map any corresponding objects in each allocation request.
[0120] Aspect 8. The method according to any one of aspects 1-7, wherein the analysis further includes determining the lifetimes of the one or more objects, wherein determining the lifetimes includes classifying the objects as normal objects or long-lived objects based on the memory allocation timestamp and the memory deallocation timestamp, and wherein the scheduling includes grouping the normal objects and grouping the long-lived objects, and scheduling the normal object group and the long-lived object group.
[0121] Aspect 9. A method for managing memory usage in a data learning operation, comprising: analyzing a batch of objects, each object including a size, an allocation timestamp, and a deallocation timestamp; arranging the analyzed batch of objects in a memory space to form a combined memory address range corresponding to the batch of objects; wherein the arranging includes determining whether the deallocation of one object in the batch of objects occurs after another object in the batch of objects based on the deallocation timestamp, and in the case where the deallocation occurs, rearranging the order of the arranged batch of objects such that the memory usage for the combined memory address range of the batch of objects in the memory space is maximized.
[0122] Aspect 10. The method according to aspect 9, wherein arranging the analyzed batch of objects includes arranging the batch of objects according to the allocation timestamp, wherein the batch of objects is provided in descending order from a first value to a second value less than the first value, wherein one object having the earliest memory allocation timestamp in the batch of objects is provided at the first value, and another object having a later memory allocation timestamp in the group of objects is provided at the second value.
[0123] Aspect 11. The method according to any one of aspects 9-10, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs during the forward pass of the training of the data learning operation, and the analysis is performed only on the first block of the forward pass, and wherein the allocation information for the one or more blocks during the scheduling is determined for the first block of the forward pass and applied to any remaining blocks of the forward pass for the training of the data learning operation.
[0124] Aspect 12. The method according to any one of aspects 9-11, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs during the backward pass of the training of the data learning operation, and the analysis is performed only on the first block of the backward pass, and wherein the allocation information for the one or more blocks during the scheduling is determined for the first block of the backward pass and applied to any remaining blocks of the backward pass for the training of the data learning operation.
[0125] Aspect 13. The method according to any one of aspects 11-12, wherein during the training of the data learning operation, for each allocation request, the size, the memory allocation timestamp, and the memory deallocation timestamp of the one or more objects determined during the analysis are used to map any corresponding objects in each allocation request.
[0126] Aspect 14. A data learning operation training system, the system comprising: a memory for storing a data learning operation; at least one processor for: analyzing one or more objects for training in the data learning operation, wherein the analysis includes determining an object size, a memory allocation timestamp, and a memory deallocation timestamp of the one or more objects; and scheduling the memory usage for the one or more objects to determine a total size and an address range of one or more groups of the one or more objects, wherein the scheduling includes: grouping the one or more objects into the one or more groups based on the memory allocation timestamp and / or the memory deallocation timestamp of the one or more objects, and arranging the one or more objects in the one or more groups in descending order in a memory space, wherein in the case where the one or more objects include two or more objects, the two or more objects are provided in the memory space in descending order from a first value to a second value less than the first value, wherein one object having the earliest memory allocation timestamp among the two or more objects is provided at the first value, and another object having a later memory allocation timestamp among the two or more objects is provided at the second value.
[0127] Aspect 15. The system according to aspect 14 further includes a graphics processing unit.
[0128] Aspect 16. The system according to any one of aspects 14-15, wherein the at least one processor is configured to further train an operation in the data learning operation based on the schedule used by the memory.
[0129] Aspect 17. The system according to any one of aspects 14-16, wherein the schedule used by the memory further includes rearranging the two or more objects in the one or more groups based on the deallocation timestamps of the two or more objects, such that the one of the two objects having a later deallocation time is moved to maximize the memory usage.
[0130] Aspect 18. The system according to any one of aspects 14-17, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs in the forward phase of the training of the data learning operation, and the analysis is performed only on the first block of the forward phase, and wherein the allocation information for the one or more blocks during the schedule is determined for the first block of the forward phase and applied to any remaining blocks of the forward phase for the training of the data learning operation.
[0131] Aspect 19. The system according to any one of aspects 14-18, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs in the backward phase of the training of the data learning operation, and the analysis is performed only on the first block of the backward phase, and wherein the allocation information for the one or more blocks during the schedule is determined for the first block of the backward phase and applied to any remaining blocks of the backward phase for the training of the data learning operation.
[0132] Aspect 20. The system according to any one of aspects 14-19, wherein the analysis further includes determining the lifespan of the one or more objects, wherein determining the lifespan includes classifying an object as a normal object or a long-lived object based on the memory allocation timestamp and the memory deallocation timestamp, and wherein the schedule includes grouping the normal objects and grouping the long-lived objects, and scheduling the normal object group and the long-lived object group.
[0133] Aspect 21. A non-transitory computer-readable medium having computer-executable instructions stored thereon, the instructions when executed causing one or more processors to perform operations including analyzing one or more objects for training in the data learning operation, wherein the analyzing includes determining an object size, a memory allocation timestamp, and a memory deallocation timestamp of the one or more objects; and scheduling memory usage for the one or more objects to determine a total size and an address range of one or more groups of the one or more objects, wherein the scheduling includes: grouping the one or more objects into the one or more groups based on the memory allocation timestamp and / or the memory deallocation timestamp of the one or more objects, and arranging the one or more objects in the one or more groups in descending order in a memory space, wherein in a case where the one or more objects include two or more objects, the two or more objects are provided in the memory space in descending order from a first value to a second value less than the first value, wherein an object having an earliest memory allocation timestamp among the two or more objects is provided at the first value, and another object having a later memory allocation timestamp among the two or more objects is provided at the second value.
[0134] The terminology used in this specification is intended to describe particular embodiments and is not intended to be limiting. Unless otherwise expressly stated, the terms "a," "an," and "the" also include plural forms. When used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or components.
[0135] With respect to the foregoing description, it should be understood that changes may be made in detail, particularly in the building materials used and in the shape, size, and arrangement of the parts, without departing from the scope of the disclosure. This specification and the described embodiments are merely exemplary, and the true scope and spirit of the disclosure are indicated by the appended claims.
Claims
1. A method for managing memory usage in a data learning operation, the method comprising: analyzing one or more objects for training in the data learning operation, wherein the analyzing includes determining an object size, a memory allocation timestamp, and a memory deallocation timestamp for the one or more objects; as well as Scheduling the memory usage for the one or more objects to determine a total size and address range of one or more groups of the one or more objects, wherein the scheduling comprises: grouping the one or more objects into the one or more groups based on the memory allocation timestamp and / or the memory deallocation timestamp of the one or more objects, and In a memory space, the one or more objects in the one or more groups are arranged in descending order, wherein when the one or more objects include two or more objects, the two or more objects are provided in the memory space in descending order from a first value to a second value less than the first value, wherein one of the two or more objects having an earliest memory allocation timestamp is provided at the first value, and another of the two or more objects having a later memory allocation timestamp is provided at the second value.
2. The method according to claim 1, further comprising: Operations in the data learning operations are trained based on the schedule of the memory usage.
3. The method of claim 1 , wherein the scheduling of the memory usage further comprises: The two or more objects in the one or more groups are rearranged based on the deallocation timestamps of the two or more objects such that one of the two objects having a later deallocation time is moved to maximize the memory usage. 4 . The method of claim 3 , wherein the one of the two objects is moved to eliminate or reduce any gaps in memory offsets.
5. A method according to claim 1, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs in a forward phase of the training of the data learning operation, and the analysis is performed only on a first block of the forward phase, and wherein during the scheduling, the allocation information for the one or more blocks is determined for the first block of the forward phase and applied to any remaining blocks of the forward phase for the training of the data learning operation.
6. A method according to claim 1, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs in a reverse phase of the training of the data learning operation, and the analysis is performed only on a first block of the reverse phase, and wherein during the scheduling, allocation information for the one or more blocks is determined for the first block of the reverse phase and applied to any remaining blocks of the reverse phase for the training of the data learning operation.
7. A method according to claim 5, wherein during training of the data learning operation, for each allocation request, the size, the memory allocation timestamp and the memory deallocation timestamp of the one or more objects determined during the analysis are used to map any corresponding objects in each allocation request.
8. The method of claim 1, wherein the analyzing further comprises: Determine the lifespan of the one or more objects, wherein determining the lifespan comprises classifying the objects as common objects or long-lived objects based on the memory allocation timestamp and the memory deallocation timestamp, and wherein the scheduling comprises grouping the common objects and grouping the long-lived objects, and scheduling the groups of the common objects and the groups of the long-lived objects.
9. A method for managing memory usage in a data learning operation, comprising: analyzing a batch of objects, each of the objects comprising a size, an allocation timestamp, and a deallocation timestamp; Arranging the analyzed batch of objects in a memory space to form a combined memory address range corresponding to the batch of objects; The arrangement includes: determining whether the deallocation of one of the objects in the batch of objects occurs after another object in the batch of objects based on the deallocation timestamp, and rearranging the order of the arranged batch of objects if the deallocation occurs, so that the memory usage of the combined memory address range for the batch of objects in the memory space is maximized.
10. The method of claim 9, wherein arranging the batch of objects analyzed comprises: The batch of objects is arranged according to the allocation timestamp, wherein the batch of objects is provided in the memory space in descending order from a first value to a second value less than the first value, wherein one object in the batch of objects having an earliest memory allocation timestamp is provided at the first value, and another object in the group of objects having a later memory allocation timestamp is provided at the second value.
11. A method according to claim 9, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs in a forward phase of the training of the data learning operation, and the analysis is performed only on a first block of the forward phase, and wherein during the scheduling, the allocation information for the one or more blocks is determined for the first block of the forward phase and applied to any remaining blocks of the forward phase for the training of the data learning operation.
12. A method according to claim 9, wherein the data learning operation is based on a transformer architecture, wherein the analysis occurs in a reverse phase of the training of the data learning operation, and the analysis is performed only on the first block of the reverse phase, and wherein during the scheduling, the allocation information for the one or more blocks is determined for the first block of the reverse phase and applied to any remaining blocks of the reverse phase for the training of the data learning operation.
13. A method according to claim 11, wherein during training of the data learning operation, for each allocation request, the size, the memory allocation timestamp and the memory deallocation timestamp of the one or more objects determined during the analysis are used to map any corresponding objects in each allocation request.
14. A data learning operation training system, the system comprising: Memory: a storage device for storing one or more programs; at least one processor, Wherein, when the one or more programs are executed by the at least one processor, the at least one processor implements the method for managing memory usage in data learning operations according to any one of claims 1 to 8 or the method for managing memory usage in data learning operations according to any one of claims 9 to 13.