Large model training method and device, electronic equipment, storage medium and program product

By dividing the training data into segmented data and efficiently moving it between different storage levels, the problem of excessive memory usage in large model training is solved, efficient training is achieved in a single-machine environment, hardware costs are reduced, and resource utilization is improved.

CN120610798APending Publication Date: 2025-09-09MOORE THREADS TECH CO LTD

Patent Information

Application Number
CN202510743643.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

During the training of large models, video memory usage is too high, especially in a single GPU environment, which limits the model scale and training efficiency. Existing technologies rely on large-scale GPU clusters or complex distributed computing, which increases hardware costs and communication overhead.

Method used

The training data is divided into multiple segments and stored in non-volatile memory. Forward and backward propagation calculations are performed on the GPU. The activation values ​​and gradient data are efficiently moved between the video memory, CPU memory and non-volatile memory, optimizing the data transmission and calculation process.

Benefits of technology

Significantly reduces video memory usage, improves training efficiency, enables efficient training of large models in a single-machine environment, reduces hardware costs, and improves resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120610798A_ABST
    Figure CN120610798A_ABST
Patent Text Reader

Abstract

The invention relates to a large model training method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: for any item of training data of a target large model, segmenting the training data into a plurality of parts of segmented data, storing the plurality of parts of segmented data in a nonvolatile memory, and sequentially performing forward propagation calculation and back propagation calculation on the plurality of parts of segmented data; for any part of segmented data, reading the segmented data from a nonvolatile memory to a video memory, and executing forward propagation calculation on the segmented data through a GPU (Graphics Processing Unit) to obtain an activation value corresponding to the segmented data; and for any part of segmented data, executing back propagation calculation based on the activation value corresponding to the segmented data through the GPU to obtain gradient data corresponding to the segmented data, and moving the gradient data corresponding to the segmented data from the video memory to a nonvolatile memory or a CPU memory. The display memory occupation of the activation value can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a large model training method, a large model training device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In the fields of artificial intelligence and machine learning, large model training has become an important research direction. With the increase in the size of large models and the expansion of training data, the performance and capabilities of large models have been significantly improved, but this has also brought a series of technical challenges, especially in the utilization and optimization of hardware resources.

[0003] During the training of large models, video memory usage is a key factor limiting model scale and training efficiency. Video memory usage is mainly composed of the following components:

[0004] Model data: includes the model's weights and bias parameters, which need to be stored and updated during the training process.

[0005] Optimizer data: used to store the state of the optimization algorithm (such as SGD, Adam, etc.), such as the accumulated values ​​of momentum and squared gradient.

[0006] Model gradient: The gradient calculated during backpropagation is used to update model parameters.

[0007] Activation value: During the forward propagation of the model, the output (activation value) of each layer needs to be stored for use in the backward propagation.

[0008] As the length of the training text increases, the proportion of video memory occupied by activation values ​​increases significantly, which poses a challenge to the training of large models with very long contexts. Summary of the Invention

[0009] The present disclosure provides a technical solution for large-scale model training.

[0010] According to one aspect of the present disclosure, a large model training method is provided, comprising:

[0011] For any training data of the target large model, the training data is divided into multiple segmented data, and the multiple segmented data are stored in a non-volatile memory, wherein the multiple segmented data are sequentially subjected to forward propagation calculation and backpropagation calculation;

[0012] For any one of the multiple pieces of segmented data, read the segmented data from the non-volatile memory to a video memory, and perform a forward propagation calculation on the segmented data through the GPU to obtain an activation value corresponding to the segmented data;

[0013] For any one of the multiple pieces of segmented data, the GPU performs a backpropagation calculation based on the activation value corresponding to the segmented data to obtain gradient data corresponding to the segmented data, and moves the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or CPU memory.

[0014] In a possible implementation, the reading the segmented data from the non-volatile memory to a video memory includes:

[0015] Reading the segmented data from the non-volatile memory to the CPU memory;

[0016] The segmented data is read from the CPU memory to the video memory.

[0017] In a possible implementation, the method further includes:

[0018] In response to starting a new round of training iteration, reading the latest model parameters of the target large model from the non-volatile memory to the CPU memory;

[0019] The latest model parameters are read from the CPU memory to the video memory.

[0020] In one possible implementation, GPU computation and data transfer can be performed in parallel;

[0021] Wherein, the GPU calculation includes the forward propagation calculation and the back propagation calculation;

[0022] The data transmission includes at least some of the following types: data transmission between different GPUs, data transmission between a GPU and a CPU, and data transmission between a CPU and a non-volatile memory.

[0023] In one possible implementation, the GPU calculation is performed through a default stream, the data transmission is performed through a prefetch stream, the default stream domain and the prefetch stream can work in parallel, and the default stream performs a create empty tensor operation before performing GPU calculation on each segmented data.

[0024] In one possible implementation,

[0025] The method further includes: reading the query-key-value matrix corresponding to the segmented data from the non-volatile memory to the display memory;

[0026] The GPU performs a forward propagation calculation on the segmented data to obtain an activation value corresponding to the segmented data, including: performing an attention calculation on the segmented data based on a query-key-value matrix corresponding to the segmented data by the GPU to obtain a partial logarithm-sum-exponential operation result corresponding to the segmented data.

[0027] In one possible implementation, performing a forward propagation calculation on the segmented data by a GPU to obtain an activation value corresponding to the segmented data includes:

[0028] A multi-layer perceptron calculation is performed on the segmented data by a GPU to obtain a multi-layer perceptron output result corresponding to the segmented data.

[0029] In one possible implementation, performing multi-layer perceptron calculation on the segmented data by a GPU to obtain a multi-layer perceptron output result corresponding to the segmented data includes:

[0030] Segmenting the segmented data in a sequence dimension to obtain multiple data blocks of the segmented data;

[0031] Performing multi-layer perceptron calculations on the multiple data blocks in sequence through the GPU to obtain multi-layer perceptron output results corresponding to the multiple data blocks;

[0032] The multilayer perceptron output result corresponding to the segmented data is determined according to the multilayer perceptron output results corresponding to the multiple data blocks.

[0033] In one possible implementation, after obtaining the activation value corresponding to the segmented data, and before performing, by the GPU, backpropagation calculation based on the activation value corresponding to the segmented data and obtaining gradient data corresponding to the segmented data, the method further includes:

[0034] Moving the activation value corresponding to the segmented data from the video memory to the non-volatile memory or the CPU memory;

[0035] The activation value corresponding to the segmented data is read from the non-volatile memory or the CPU memory to the video memory.

[0036] In one possible implementation, performing back propagation calculations based on activation values ​​corresponding to the segmented data by the GPU to obtain gradient data corresponding to the segmented data includes:

[0037] Obtaining the total number of training data in the training data set used to train the target large model, and the sum of squares of the training data in the training data set;

[0038] determining a normalized root mean square of the segmented data based on the total number and the sum of squares;

[0039] Back propagation calculation is performed based on the activation value corresponding to the segmented data and the normalized root mean square of the segmented data to obtain gradient data corresponding to the segmented data.

[0040] In a possible implementation, after moving the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or the CPU memory, the method further includes:

[0041] Obtaining gradient data corresponding to the plurality of segmented data;

[0042] The model parameters of the target large model are updated according to the gradient data corresponding to the multiple pieces of segmented data, and the updated model parameters of the target large model are written into the non-volatile memory.

[0043] According to one aspect of the present disclosure, a large model training device is provided, comprising:

[0044] a segmentation module, configured to segment any training data of the target large model into a plurality of segmented data, and store the plurality of segmented data in a non-volatile memory, wherein the plurality of segmented data are sequentially subjected to forward propagation calculation and backpropagation calculation;

[0045] a forward propagation module, configured to read, for any one of the plurality of segmented data, the segmented data from the non-volatile memory into a video memory, and perform a forward propagation calculation on the segmented data through the GPU to obtain an activation value corresponding to the segmented data;

[0046] A back propagation module is used to perform a back propagation calculation on any one of the multiple segmented data based on the activation value corresponding to the segmented data through the GPU to obtain gradient data corresponding to the segmented data, and move the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or CPU memory.

[0047] In one possible implementation, the forward propagation module is used to:

[0048] Reading the segmented data from the non-volatile memory to the CPU memory;

[0049] The segmented data is read from the CPU memory to the video memory.

[0050] In a possible implementation, the forward propagation module is further configured to:

[0051] In response to starting a new round of training iteration, reading the latest model parameters of the target large model from the non-volatile memory to the CPU memory;

[0052] The latest model parameters are read from the CPU memory to the video memory.

[0053] In one possible implementation, GPU computation and data transfer can be performed in parallel;

[0054] Wherein, the GPU calculation includes the forward propagation calculation and the back propagation calculation;

[0055] The data transmission includes at least some of the following types: data transmission between different GPUs, data transmission between a GPU and a CPU, and data transmission between a CPU and a non-volatile memory.

[0056] In one possible implementation, the GPU calculation is performed through a default stream, the data transmission is performed through a prefetch stream, the default stream domain and the prefetch stream can work in parallel, and the default stream performs a create empty tensor operation before performing GPU calculation on each segmented data.

[0057] In a possible implementation, the forward propagation module is further configured to:

[0058] Reading the query-key-value matrix corresponding to the segmented data from the non-volatile memory to the display memory;

[0059] An attention calculation is performed on the segmented data based on a query-key-value matrix corresponding to the segmented data by a GPU to obtain a partial logarithm-sum-exponential operation result corresponding to the segmented data.

[0060] In one possible implementation, the forward propagation module is used to:

[0061] A multi-layer perceptron calculation is performed on the segmented data by a GPU to obtain a multi-layer perceptron output result corresponding to the segmented data.

[0062] In one possible implementation, the forward propagation module is used to:

[0063] Segmenting the segmented data in a sequence dimension to obtain multiple data blocks of the segmented data;

[0064] Performing multi-layer perceptron calculations on the multiple data blocks in sequence through the GPU to obtain multi-layer perceptron output results corresponding to the multiple data blocks;

[0065] The multilayer perceptron output result corresponding to the segmented data is determined according to the multilayer perceptron output results corresponding to the multiple data blocks.

[0066] In a possible implementation, the apparatus further includes:

[0067] A moving module, configured to move the activation value corresponding to the segmented data from the video memory to the non-volatile memory or the CPU memory;

[0068] A reading module is used to read the activation value corresponding to the segmented data from the non-volatile memory or the CPU memory to the video memory.

[0069] In one possible implementation, the back propagation module is used to:

[0070] Obtaining the total number of training data in the training data set used to train the target large model, and the sum of squares of the training data in the training data set;

[0071] determining a normalized root mean square of the segmented data based on the total number and the sum of squares;

[0072] Back propagation calculation is performed based on the activation value corresponding to the segmented data and the normalized root mean square of the segmented data to obtain gradient data corresponding to the segmented data.

[0073] In a possible implementation, the apparatus further includes an updating module, wherein the updating module is configured to:

[0074] Obtaining gradient data corresponding to the plurality of segmented data;

[0075] The model parameters of the target large model are updated according to the gradient data corresponding to the multiple pieces of segmented data, and the updated model parameters of the target large model are written into the non-volatile memory.

[0076] According to one aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0077] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented.

[0078] According to one aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0079] In an embodiment of the present disclosure, for any one training data of a target large model, the training data is divided into multiple segmented data, and the multiple segmented data are stored in a non-volatile memory, wherein the multiple segmented data are sequentially subjected to forward propagation calculation and back propagation calculation, and for any one segmented data among the multiple segmented data, the segmented data is read from the non-volatile memory to the video memory, and the forward propagation calculation is performed on the segmented data by the GPU to obtain an activation value corresponding to the segmented data, and for any one segmented data among the multiple segmented data, the back propagation calculation is performed by the GPU based on the activation value corresponding to the segmented data to obtain gradient data corresponding to the segmented data, and the gradient data corresponding to the segmented data is moved from the video memory to the non-volatile memory or CPU memory, thereby effectively managing the movement of data between different storage levels (video memory, CPU memory, non-volatile memory), so that the video memory occupancy of the activation value is significantly reduced. For example, 128MB of training data can be split into 1024 segments, each of which is 8KB in size. This means that the activation value of each segment only occupies 64MB and can be stored in a single GPU memory. This effectively solves the problem of insufficient GPU memory on a single GPU, making it possible to train large models on a single machine.

[0080] In the disclosed embodiments, by optimizing the data transmission and calculation processes, the video memory usage is significantly reduced and the training efficiency of large models is improved, thereby effectively solving the problem of activation values ​​occupying a large amount of video memory during the training of large models with very long contexts, and achieving efficient training of large models with very long contexts. This allows large models to be trained efficiently even with limited hardware resources, thereby reducing hardware costs and improving resource utilization.

[0081] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure.

[0082] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0084] Figure 1 A flowchart of the large model training method provided by an embodiment of the present disclosure is shown.

[0085] Figure 2 A schematic diagram of a memory exchange controller in a large model training method provided by an embodiment of the present disclosure is shown.

[0086] Figure 3 A block diagram of a large model training device provided by an embodiment of the present disclosure is shown.

[0087] Figure 4 A block diagram of an electronic device 1900 provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0088] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.

[0089] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0090] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of three situations: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.

[0091] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0092] Related technologies primarily address the memory usage issue of training large models with very long contexts by increasing the cluster size and using more GPUs (Graphics Processing Units) to share the computational load and memory requirements. Specifically, by increasing the number of GPUs, related technologies can distribute the model and data across more GPUs, thereby alleviating the memory pressure on a single GPU. More GPUs can process longer context data in parallel, but this also requires greater hardware investment and higher operating costs. Related technologies also expand the size of the computing cluster and utilize more nodes for distributed training. Each node processes a portion of the data, and data is exchanged and synchronized between nodes through efficient communication protocols (such as NCCL). This approach can support larger models and longer contexts, but it also requires expensive hardware resources and complex cluster management. Therefore, although increasing the cluster size and using more GPUs can, to a certain extent, address the memory usage issue of training large models with very long contexts, this approach comes at the cost of high hardware and operating expenses.

[0093] In summary, in related technologies, the challenges faced in training large models with very long contexts include:

[0094] High GPU memory requirements: Large model training methods used in related technologies place extremely high demands on GPU memory when processing extremely long text. Each layer of the model generates a large number of intermediate activation values ​​during forward and backward propagation. These activation values ​​require a large amount of GPU memory to store and calculate, leading to insufficient GPU memory. This limits the length of context that the model can process, especially in a single GPU (Graphics Processing Unit) environment.

[0095] Cluster computing reliance: Due to the limited memory of a single GPU, related technologies typically require large-scale GPU clusters for training. By distributing the model and data across multiple GPUs, cluster computing can provide sufficient memory and computing power to process extremely long texts. However, this approach relies on expensive hardware resources and a complex distributed computing architecture, increasing training costs and hindering widespread adoption. Furthermore, cluster computing introduces significant data transmission requirements, increasing communication overhead and reducing overall training efficiency.

[0096] Distributed memory management: To effectively utilize cluster resources, related technologies employ distributed memory management techniques, such as data parallelism and model parallelism. These techniques reduce the memory pressure on a single GPU by distributing and coordinating memory usage across multiple GPUs. However, they also introduce additional communication overhead and complexity, reducing overall training efficiency.

[0097] Memory-optimized algorithms: Some memory-optimized algorithms, such as gradient checkpointing, attempt to reduce memory usage by storing fewer intermediate results during forward propagation and recalculating them during backward propagation. However, while this approach reduces memory requirements, it increases the computational burden, resulting in a decrease in overall training speed.

[0098] In summary, when dealing with large model training with extremely long contexts, existing technologies have problems such as excessive video memory usage, reliance on large-scale GPU clusters, low data transmission efficiency, and low computing efficiency. A new method is urgently needed to optimize these aspects.

[0099] In order to solve technical problems similar to those described above, an embodiment of the present disclosure provides a large model training method, which divides any training data of a target large model into multiple segmented data and stores the multiple segmented data in a non-volatile memory, wherein the multiple segmented data are sequentially subjected to forward propagation calculation and back propagation calculation. For any segmented data among the multiple segmented data, the segmented data is read from the non-volatile memory to the video memory, and the GPU performs forward propagation calculation on the segmented data to obtain an activation value corresponding to the segmented data. For any segmented data among the multiple segmented data, the GPU performs back propagation calculation based on the activation value corresponding to the segmented data to obtain gradient data corresponding to the segmented data, and moves the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or CPU memory, thereby effectively managing the movement of data between different storage levels (video memory, CPU memory, non-volatile memory), so that the video memory occupancy of the activation value is significantly reduced. For example, 128MB of training data can be split into 1024 segments, each of which is 8KB in size. This means that the activation value of each segment only occupies 64MB and can be stored in a single GPU memory. This effectively solves the problem of insufficient GPU memory on a single GPU, making it possible to train large models on a single machine.

[0100] In the disclosed embodiments, by optimizing the data transmission and calculation processes, the video memory usage is significantly reduced and the training efficiency of large models is improved, thereby effectively solving the problem of activation values ​​occupying a large amount of video memory during the training of large models with very long contexts, and achieving efficient training of large models with very long contexts. This allows large models to be trained efficiently even with limited hardware resources, thereby reducing hardware costs and improving resource utilization.

[0101] The large model training method provided by the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.

[0102] Figure 1A flow chart of the large model training method provided by an embodiment of the present disclosure is shown. In one possible implementation, the execution subject of the large model training method may be a large model training device. For example, the large model training method may be executed by a terminal device or a server or other electronic device. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device or a wearable device, etc. In some possible implementations, the large model training method may be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 As shown, the large model training method includes steps S11 to S13.

[0103] In step S11, for any training data of the target large model, the training data is divided into multiple segmented data, and the multiple segmented data are stored in a non-volatile memory, wherein the multiple segmented data are sequentially subjected to forward propagation calculation and back propagation calculation.

[0104] In step S12, for any one of the multiple pieces of segmented data, the segmented data is read from the non-volatile memory to the video memory, and a forward propagation calculation is performed on the segmented data through the GPU to obtain an activation value corresponding to the segmented data.

[0105] In step S13, for any one of the multiple segmented data, the GPU performs back propagation calculation based on the activation value corresponding to the segmented data to obtain gradient data corresponding to the segmented data, and moves the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or CPU memory.

[0106] In a possible implementation, the target large model may be a large language model.

[0107] As an example of this implementation, the target large model may be a large language model with a very long context.

[0108] In the disclosed embodiments, any training data for a target large model can be split into multiple segments. For example, 128MB of training data can be split into 1024 segments, each 8KB in size. Thus, the activation value generated by each segment is 64MB in size, which can be stored in a single GPU memory. This effectively solves the problem of insufficient GPU memory on a single GPU, making large model training possible on a single machine.

[0109] The principles for segmenting the training data may include ensuring that the segmented data can fully utilize the GPU's performance, that the video memory can store the activation values ​​corresponding to two segments of data, and that video memory overflow does not occur. By segmenting the training data, each segment of data can fully utilize the GPU's performance when performing forward and backward propagation through the target large model. This means that the size of each segment of data should match the GPU's computing power so that optimal computational efficiency can be achieved when processing each segment of data. In the disclosed embodiments, the size of the segmented data can be determined through testing.

[0110] In the disclosed embodiment, relevant data for large model training can be stored in video memory, CPU memory, and non-volatile memory.

[0111] In one example, the size of the activation value corresponding to each segmented data is M. Two buffers with a capacity of M can be requested in the video memory. In the case of multiple GPUs and multiple video memories, two buffers with a capacity of M can be requested in each of the multiple video memories. Half of the CPU memory can be used to store the activation values, so that the other half of the CPU memory can still be used for the normal operation of the operating system and other programs, as well as other computing needs of deep learning frameworks and models. In this case, the number of buffers in the CPU memory can be: (CPU memory capacity / 2) / (number of GPUs on the same host × M).

[0112] In an embodiment of the present disclosure, after the training data is divided into a plurality of segmented data, the plurality of segmented data may be stored in a non-volatile memory.

[0113] In one possible implementation, a memory swap controller can be used to manage data movement between different storage levels (video memory, CPU memory, and non-volatile memory) to optimize video memory usage and improve data transmission efficiency. Figure 2 Schematic diagram of the memory exchange controller in the large model training method provided by the embodiment of the present disclosure is shown. Figure 2 As shown, the memory swap controller may include a non-volatile memory and CPU memory swapper (nc-swapper) and a CPU memory and video memory swapper (cg-swapper). The non-volatile memory and CPU memory swapper may be used to manage the buffers in the non-volatile memory and the buffers in the CPU memory, as well as the movement of data between the non-volatile memory and the CPU memory; the CPU memory and video memory swapper may be used to manage the buffers in the CPU memory and the buffers in the video memory, as well as the movement of data between the CPU memory and the video memory.

[0114] In one possible implementation, the non-volatile memory may be NVMe (Non-Volatile Memory Express) memory. Of course, other types of non-volatile memory may also be used, which is not limited here.

[0115] Since video memory is expensive and has limited capacity (such as 80GB), while CPU memory is cheap and can be expanded to a larger capacity (such as 1000GB) at a lower cost, non-volatile memory has a larger capacity, is easier to expand and has a lower cost, but the data transmission speeds of video memory, CPU memory and non-volatile memory decrease in sequence. Therefore, in the large model training scheme provided in the embodiment of the present disclosure, the storage location can be arranged according to the time sequence in which the data is about to be used. The data to be used is preferentially stored in the video memory to obtain the fastest computing speed, followed by the data to be used in the short term being stored in the CPU memory, and those data that are not used temporarily or will be used last are stored in the non-volatile memory, so as to optimize the utilization efficiency and cost of storage resources.

[0116] In a possible implementation, the reading the segmented data from the non-volatile memory to the video memory includes: reading the segmented data from the non-volatile memory to the CPU memory; and reading the segmented data from the CPU memory to the video memory.

[0117] As an example of this implementation, for any of the multiple pieces of segmented data, the segmented data can first be read from the non-volatile memory to the CPU memory via a non-volatile memory and CPU memory switch, and then read from the CPU memory to the video memory via a CPU memory and video memory switch. The CPU memory can provide a large buffer for temporarily storing the segmented data read from the non-volatile memory.

[0118] Data transfer from CPU memory to graphics memory is typically faster than data transfer from non-volatile memory to graphics memory. In this implementation, segmented data can be pre-read from non-volatile memory to CPU memory. This allows the GPU to retrieve the segmented data directly from CPU memory when needed, rather than from slower non-volatile memory. This pre-fetching mechanism reduces the idle time the GPU spends waiting for segmented data. In other words, the GPU can continuously process multiple segments without waiting for new data. This buffer management strategy helps maintain continuous GPU operation and improves overall throughput.

[0119] In one possible implementation, the method further includes: in response to starting a new round of training iterations, reading the latest model parameters of the target large model from the non-volatile memory to the CPU memory; and reading the latest model parameters from the CPU memory to the video memory.

[0120] In this implementation, after each training iteration, the model parameters (weights and biases, etc.) are updated and stored in non-volatile memory. When a new training iteration begins, the latest model parameters can be read from non-volatile memory into CPU memory, and then loaded from CPU memory into video memory for use in the new training iteration.

[0121] In this implementation, by storing model parameters in non-volatile memory and loading them into video memory only when needed, video memory usage can be reduced, freeing up space for other necessary data. This implementation makes more efficient use of limited video memory resources, especially when processing large models or in memory-constrained environments. By loading model parameters on demand, video memory is ensured to be used for the computational tasks that require it most.

[0122] In the disclosed embodiments, forward propagation calculations can be performed sequentially on multiple segments of data, where multiple segments of data can be loaded into the video memory sequentially for forward propagation calculations, without having to load all the segmented data into the video memory at once. Forward propagation refers to the process by which data is calculated from the input layer through each layer of the neural network to the output layer. This process may involve multiplication of the weight matrix with the input data, addition of bias terms, and application of activation functions.

[0123] In an embodiment of the present disclosure, for any one of the multiple pieces of segmented data, the segmented data can be read from the non-volatile memory to the video memory, and the GPU can perform forward propagation calculations on the segmented data to obtain the activation values ​​corresponding to the segmented data. In an embodiment of the present disclosure, for each piece of segmented data, the target large model can perform forward propagation calculations in sequence. That is, each piece of segmented data can be sent to the target large model separately to perform a series of calculations until the output result of the segmented data is generated. In the forward propagation process, in addition to the final output result, intermediate results are also generated, such as the activation value of each layer. These intermediate results are necessary for subsequent back propagation calculations because they are used to calculate the gradient of the loss function relative to the model parameters.

[0124] In one possible implementation, GPU computing and data transmission can be performed in parallel; wherein, the GPU computing includes the forward propagation computing and the backward propagation computing; and the data transmission includes at least some of the following types: data transmission between different GPUs, data transmission between a GPU and a CPU, and data transmission between a CPU and a non-volatile memory.

[0125] During deep learning model training, especially when processing large models with very long contexts, a single GPU may not be able to meet the computing and storage requirements because the model and data are often very large. Therefore, it is often necessary to distribute the model and data across multiple GPUs for training and computing. Data transfer between different GPUs can refer to the process of transferring data or synchronizing states between these distributed GPUs to ensure that computing on each GPU can be coordinated and consistent.

[0126] As an example of this implementation, the parallel execution of GPU computing and data transmission can be achieved through an overlay engine.

[0127] This implementation allows data transfer and GPU computation to proceed synchronously. During the forward and backward propagation of the target large model, GPU computation, data transfer between different GPUs, data transfer between the GPU and the CPU, and data transfer between the CPU and non-volatile memory are performed in parallel, thereby improving data transfer efficiency and, consequently, the overall training efficiency of the target large model.

[0128] As an example of this implementation, the GPU calculation is performed through a default stream, the data transfer is performed through a prefetch stream, the default stream domain and the prefetch stream can work in parallel, and the default stream performs a create empty tensor operation before performing GPU calculation on each segmented data.

[0129] In one example, the overlay engine can include two parts: a dynamic prefetcher, which can be used to prefetch and offload data during forward and backward propagation, and a computation and offload overlay mechanism, which can be used to perform data movement required for gradients and backward propagation computations in parallel.

[0130] For example, during backpropagation, the dynamic prefetcher can record the number of activation values ​​corresponding to the segmented data. For example, if there are N sets of activation values ​​corresponding to N segments, the activation values ​​corresponding to the first segmented data can be stored in the video memory, the activation values ​​corresponding to the second to ninth segments can be stored in the CPU memory, and the activation values ​​corresponding to the remaining segments can be stored in the non-volatile memory. The dynamic prefetcher can send an instruction to the CPU memory and video memory switch to prefetch the activation value corresponding to the second segmented data, and send an instruction to the non-volatile memory and CPU memory switch to prefetch the activation value corresponding to the tenth segmented data. After completing the back-propagation calculation of the activation value corresponding to the first piece of segmented data, the dynamic prefetcher can send an instruction to prefetch the activation value corresponding to the third piece of segmented data and an instruction to store the gradient data corresponding to the first piece of segmented data to the CPU memory and video memory switch, and can send an instruction to prefetch the activation value corresponding to the 11th piece of segmented data and an instruction to store the gradient data corresponding to the first piece of segmented data to the non-volatile memory and CPU memory switch (executed after the aforementioned CPU memory and video memory switch is completed).

[0131] The computation and offloading overlap mechanism can be achieved through the collaborative work of multiple execution streams in the GPU. For example, the default stream can be used to perform GPU computation, and a prefetch stream can be applied to perform prefetch and offload operations. The default stream can execute an empty tensor creation operation (such as the torch.empty(0) operation) before performing GPU computation on each segmented data. The empty tensor creation operation (such as the torch.empty(0) operation) does not actually perform any meaningful computation, but it will serve as a placeholder and occupy a position in the GPU's instruction queue. By doing this, the prefetch stream (the stream responsible for data prefetching and offloading) can have the opportunity to issue its kernel (i.e., the function or program executed on the GPU) before the default stream performs actual computation, thereby achieving parallel computation and data offloading, and improving overall training efficiency.

[0132] When calculating the root mean square error (RMSE), it is also performed directly on the original memory location of the data, without taking up additional memory space.

[0133] In one possible implementation, the method further includes: reading the query-key-value matrix corresponding to the segmented data from the non-volatile memory to the video memory; performing forward propagation calculation on the segmented data by the GPU to obtain the activation value corresponding to the segmented data, including: performing attention calculation on the segmented data based on the query-key-value matrix corresponding to the segmented data by the GPU to obtain the partial logarithm-sum-exponential operation result corresponding to the segmented data.

[0134] In this implementation, the forward propagation calculation may include an attention calculation, and the activation value corresponding to the segmented data may include a partial logarithm-sum-exponential operation result corresponding to the segmented data.

[0135] In this implementation, a query-key-value (QKV) matrix corresponding to the segmented data can be received from the memory exchange controller. A partial output result (O) corresponding to the segmented data can be obtained from the query-key-value matrix corresponding to the segmented data through a preset type of calculation (such as matrix multiplication or dot product). Based on the partial output result (O) corresponding to the segmented data, an attention calculation can be performed on the segmented data to obtain a partial logarithm-sum-exponential operation result (Log-Sum-Exp, LSE) corresponding to the segmented data. The attention calculation can involve weighted summation of the partial output result (O), and the weight can be determined by the similarity (such as dot product) between the query (Q) and the key (K). The partial logarithm-sum-exponential operation result corresponding to the segmented data can be stored in the memory exchange controller, and the query-key-value matrix corresponding to the next segmented data can be obtained. The above process can be repeated until the attention calculation of all segmented data is completed.

[0136] This implementation utilizes an in-place attention mechanism to process the query-key-value matrix in segments, performing attention calculations on each segment. This reduces the memory usage associated with one-time calculations and enables attention calculations on very long text without increasing memory usage. Furthermore, during the attention calculation process, data can be effectively managed and moved with the support of a memory exchange controller.

[0137] In-place is a computational strategy that allows algorithms to perform calculations and updates directly at existing memory locations, without requiring additional memory to store intermediate results. This computational strategy can reduce memory usage and improve computational efficiency, especially when processing large amounts of data or in resource-constrained environments. By adopting an in-place mechanism, the memory usage of activation values ​​can be effectively managed and reduced. The in-place attention mechanism refers to performing calculations and updates directly at the input memory locations during attention calculations, without requiring additional memory to store intermediate results. Similarly, the in-place multilayer perceptron mechanism mentioned below refers to performing calculations directly at the input memory locations during multilayer perceptron calculations, avoiding additional memory allocations. The in-place root mean square error (RMSE) mechanism mentioned below refers to performing RMS error calculations directly at the data's original memory locations, without occupying additional memory.

[0138] In one possible implementation, performing forward propagation calculations on the segmented data by a GPU to obtain activation values ​​corresponding to the segmented data includes: performing multi-layer perceptron calculations on the segmented data by a GPU to obtain multi-layer perceptron output results corresponding to the segmented data.

[0139] In this implementation, the forward propagation computation may include a multilayer perceptron computation, and the activation value corresponding to the segmented data may include the multilayer perceptron output. In this implementation, after obtaining the multilayer perceptron output corresponding to the segmented data, the multilayer perceptron output corresponding to the segmented data may be transmitted back to the memory exchange controller for subsequent processing.

[0140] In this implementation, an in-place multilayer perceptron (MLP) mechanism is used to process input data in segments and perform MLP calculations segment by segment, thereby reducing the video memory usage caused by one-time calculations. This allows MLP calculations of very long texts to be performed without increasing video memory usage.

[0141] As an example of this implementation, the performing of multilayer perceptron calculations on the segmented data by a GPU to obtain a multilayer perceptron output result corresponding to the segmented data includes: segmenting the segmented data in a sequence dimension to obtain multiple data blocks of the segmented data; performing multilayer perceptron calculations on the multiple data blocks in sequence by a GPU to obtain multilayer perceptron output results corresponding to the multiple data blocks; and determining the multilayer perceptron output result corresponding to the segmented data based on the multilayer perceptron output results corresponding to the multiple data blocks.

[0142] In this example, the segmented data can be further segmented in the sequence dimension, thereby dividing the segmented data into smaller data blocks to adapt to the limitations of computing resources, especially the limitations of video memory, and improve computing efficiency. Each segmented data block can be sent separately to the multilayer perceptron network for processing. The multilayer perceptron network is a feedforward neural network containing multiple hidden layers. It can perform nonlinear transformations on data to learn the complex features of the data. In this step, each data block can be independently passed through the multilayer perceptron network to obtain its own multilayer perceptron output result. After the multilayer perceptron calculation is performed, the multilayer perceptron output result of each data block can be passed to the memory exchange controller.

[0143] Since the multilayer perceptron output result of each data block is part of the multilayer perceptron output result of the segmented data, it is necessary to merge the multilayer perceptron output results of each data block to determine the output result of the entire segmented data after passing through the multilayer perceptron network, that is, the multilayer perceptron output result corresponding to the segmented data. The multilayer perceptron output results of all data blocks can be connected in the sequence dimension. For example, if the original segmented data is divided into N data blocks, the multilayer perceptron output results of each data block can be spliced ​​together in the sequence dimension to reconstruct the multilayer perceptron output result of the entire segmented data.

[0144] The above process can be repeated until all segmented data have completed the multilayer perceptron calculation.

[0145] In this example, the segmented data is segmented in the sequence dimension to obtain multiple data blocks of the segmented data, and the GPU performs multi-layer perceptron calculations on the multiple data blocks in sequence to obtain the multi-layer perceptron output results corresponding to the multiple data blocks. Based on the multi-layer perceptron output results corresponding to the multiple data blocks, the multi-layer perceptron output results corresponding to the segmented data are determined, thereby effectively managing and utilizing limited video memory resources.

[0146] In one possible implementation, after obtaining the activation value corresponding to the segmented data, and before performing backpropagation calculation based on the activation value corresponding to the segmented data by the GPU and obtaining the gradient data corresponding to the segmented data, the method further includes: moving the activation value corresponding to the segmented data from the video memory to the non-volatile memory or the CPU memory; and reading the activation value corresponding to the segmented data from the non-volatile memory or the CPU memory to the video memory.

[0147] In this implementation, after the forward propagation calculation is completed, each segmented data will generate a corresponding activation value. The activation value is the output of each layer in the target large model and is necessary for the subsequent backpropagation calculation because it will be used to calculate the gradient of the loss function with respect to the model parameters. Due to limited video memory resources, especially when processing large models or large datasets, it may not be possible to simultaneously retain the activation values ​​corresponding to all segmented data. Therefore, the activation values ​​corresponding to the segmented data need to be stored elsewhere so that they can be reloaded when needed.

[0148] In this implementation, the activation values ​​corresponding to the segmented data can be moved from the video memory to the non-volatile memory or CPU memory, thereby freeing up video memory space to process the forward propagation calculation of the next segmented data or perform other computing tasks.

[0149] Before backpropagation calculations need to be performed, the activation values ​​stored in non-volatile memory or CPU memory need to be read back into G memory. This is because backpropagation calculations require activation values ​​to calculate gradients, that is, the rate of change of the loss function relative to the model parameters. These gradients will be used to update the model parameters.

[0150] In this implementation, by temporarily storing activation values ​​in non-volatile memory or CPU memory, video memory usage can be effectively managed to avoid video memory overflow, especially when processing large data sets or complex models. Although the data transfer step is increased, this method allows the GPU to perform other tasks while waiting for the data transfer to complete, such as preprocessing the next segmented data or performing other computing tasks, thereby improving overall computing efficiency. This implementation provides greater flexibility, allowing data to be dynamically managed between different storage levels to adapt to different training requirements and hardware configurations.

[0151] In an embodiment of the present disclosure, back propagation calculations can be performed sequentially on multiple pieces of segmented data. During the back propagation process, a memory exchange controller can be used to manage the movement of activation values ​​and gradient data, so that activation values ​​can be loaded into the video memory in a timely manner, and gradient data can be unloaded to the CPU memory or non-volatile memory in a timely manner. In an embodiment of the present disclosure, for any piece of segmented data among the multiple pieces of segmented data, a back propagation calculation can be performed by the GPU based on the activation value corresponding to the segmented data to obtain the gradient data corresponding to the segmented data, and the gradient data corresponding to the segmented data can be moved from the video memory to the non-volatile memory or CPU memory.

[0152] In one possible implementation, the GPU performs back propagation calculations based on activation values ​​corresponding to the segmented data to obtain gradient data corresponding to the segmented data, including: obtaining the total number of training data in a training data set used to train the target large model, and the sum of squares of the training data in the training data set; determining the normalized root mean square of the segmented data based on the total number and the sum of squares; and performing back propagation calculations based on the activation values ​​corresponding to the segmented data and the normalized root mean square of the segmented data to obtain gradient data corresponding to the segmented data.

[0153] In this implementation, the in-place root mean square error (In-Place RMSE, In-Place Root Mean Square Error) mechanism can be used to realize the calculation of the normalized root mean square. Among them, the in-place root mean square error mechanism can include two rounds of calculation. In the first round of calculation, the total number of training data in the training data set used to train the target large model (that is, the global number of samples used for normalized root mean square) and the sum of the squares of the training data can be recorded. In the second round of calculation, the normalized root mean square of each segmented data can be calculated using the total number of training data in the training data set recorded in the first round and the sum of the squares of the training data. After calculating the normalized root mean square of the segmented data, back propagation calculation can be performed based on the normalized root mean square of the segmented data and the activation value corresponding to the segmented data to obtain the gradient data corresponding to the segmented data.

[0154] In this implementation, efficient normalization processing can be achieved without increasing video memory usage through segmented processing and segment-by-segment calculation.

[0155] In one possible implementation, after moving the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or CPU memory, the method further includes: obtaining the gradient data corresponding to the multiple pieces of segmented data; updating the model parameters of the target large model based on the gradient data corresponding to the multiple pieces of segmented data, and writing the updated model parameters of the target large model into the non-volatile memory.

[0156] In this implementation, after backpropagation calculations are completed for all segmented data, the gradient data corresponding to all segmented data can be aggregated. Optimizers (such as SGD and Adam) can be used to update model parameters in CPU memory and write the updated model parameters back to non-volatile memory for use in the next round of calculations.

[0157] In this implementation, by moving gradient data to non-volatile memory or CPU memory, GPU memory resources can be efficiently managed, avoiding GPU memory overflow, especially when training large models. Performing parameter updates in CPU memory provides greater flexibility, as CPU memory is typically much larger than GPU memory and can handle more complex optimization algorithms. Furthermore, saving updated model parameters to non-volatile memory prevents data loss during training, ensuring the continuity and stability of model training.

[0158] The large model training method provided by the embodiments of the present disclosure can be applied to technical fields such as artificial intelligence and machine learning, and is not limited here.

[0159] The following describes the large model training method provided by the embodiment of the present disclosure through a specific application scenario. In this application scenario, the large model training method may include four components: a memory exchange controller, an in-situ attention mechanism, an in-situ multi-layer perceptron mechanism, and an in-situ root mean square error mechanism. The four components work together to achieve high efficiency and video memory optimization for large model training with ultra-long contexts. The memory exchange controller may include a non-volatile memory and CPU memory exchanger and a CPU memory and video memory exchanger.

[0160] In this application scenario, the large model training method may include the following steps:

[0161] The first step is data loading and segmentation.

[0162] In this step, for any training data of the target large model, the training data can be divided into N segments, and the N segments are stored in a non-volatile memory, where N is an integer greater than 1.

[0163] The first segmented data can be read from the non-volatile memory to the CPU memory through the non-volatile memory and CPU memory switch, and the segmented data can be loaded from the CPU memory to the video memory through the CPU memory and video memory switch to prepare for forward propagation calculation.

[0164] The second step is forward propagation.

[0165] In this step, forward propagation calculation is performed on each segmented data in sequence.

[0166] When a new round of training iterations begins, the latest model parameters of the target large model can be read from the non-volatile memory to the CPU memory via the non-volatile memory and CPU memory switch, and the latest model parameters can be read from the CPU memory to the video memory via the CPU memory and video memory switch. Segmented data can be read from the non-volatile memory to the CPU memory via the non-volatile memory and CPU memory switch, and segmented data can be read from the CPU memory to the video memory via the CPU memory and video memory switch.

[0167] The forward propagation calculation may include attention calculation and multi-layer perceptron calculation. The attention calculation may adopt an in-situ attention mechanism, and the multi-layer perceptron calculation may adopt an in-situ multi-layer perceptron mechanism.

[0168] In which, the query-key-value matrix corresponding to the segmented data input by the memory exchange controller can be received. The partial output result (O) corresponding to the segmented data can be obtained from the query-key-value matrix corresponding to the segmented data through a preset type of calculation (such as matrix multiplication or dot product). Based on the partial output result (O) corresponding to the segmented data, attention calculation can be performed on the segmented data to obtain the partial logarithm-sum-exponential operation result corresponding to the segmented data. In which, the attention calculation can involve weighted summation of the partial output result (O), and the weight can be determined by the similarity (such as dot product) between the query (Q) and the key (K). The partial logarithm-sum-exponential operation result corresponding to the segmented data can be stored in the memory exchange controller, and the query-key-value matrix corresponding to the next segmented data can be obtained. The above process can be repeated until the attention calculation of all segmented data is completed.

[0169] For any segmented data, the segmented data can be segmented along a sequence dimension to obtain multiple data blocks of the segmented data. Multilayer perceptron calculations can be performed sequentially on the multiple data blocks to obtain corresponding multilayer perceptron output results, which are then stored sequentially in a memory exchange controller. Based on the corresponding multilayer perceptron output results, the corresponding multilayer perceptron output result can be determined for the segmented data. This process can be repeated until the multilayer perceptron calculations for all segmented data are completed.

[0170] After each segmented data is processed, the corresponding video memory space can be released to prepare for the loading and calculation of the next segmented data.

[0171] The third step is back propagation.

[0172] In this step, backpropagation calculation is performed on each segmented data in sequence.

[0173] In back propagation, an in-situ root mean square error mechanism can be used to realize the calculation of normalized root mean square. Among them, the in-situ root mean square error mechanism can include two rounds of calculation. In the first round of calculation, the total number of training data in the training data set used to train the target large model (that is, the global number of samples used for normalized root mean square) and the sum of squares of the training data can be recorded. In the second round of calculation, the normalized root mean square of each segmented data can be calculated using the total number of training data in the training data set recorded in the first round and the sum of squares of the training data. After calculating the normalized root mean square of the segmented data, back propagation calculation can be performed based on the normalized root mean square of the segmented data and the activation value corresponding to the segmented data to obtain the gradient data corresponding to the segmented data.

[0174] The gradient data corresponding to the segmented data may be moved from the video memory to the non-volatile memory or the CPU memory.

[0175] Step 4: Gradient update.

[0176] After the backpropagation calculation of all segmented data is completed, the gradient data corresponding to all segmented data can be aggregated. The optimizer (such as SGD, Adam, etc.) can be used to update the model parameters in the CPU memory and write the updated model parameters back to the non-volatile memory for use in the next round of calculation.

[0177] By adopting the above process, efficient training of the 7B parameter Llama model can be achieved in a single-machine 8-GPU environment, effectively solving the graphics memory bottleneck problem in ultra-long context training.

[0178] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.

[0179] In addition, the present disclosure also provides a large model training device, an electronic device, a computer-readable storage medium, and a computer program product, all of which can be used to implement any large model training method provided by the present disclosure. The corresponding technical solutions and technical effects can be found in the corresponding records in the method section and will not be repeated here.

[0180] Figure 3 FIG. 1 is a block diagram of a large model training device provided by an embodiment of the present disclosure. Figure 3 As shown, the large model training device includes:

[0181] A segmentation module 31 is configured to segment any training data of a target large model into multiple segments, and store the multiple segments in a non-volatile memory, wherein the multiple segments are sequentially subjected to forward propagation calculations and backpropagation calculations;

[0182] A forward propagation module 32 is configured to read, for any one of the plurality of segmented data, the segmented data from the non-volatile memory to a video memory, and perform a forward propagation calculation on the segmented data through the GPU to obtain an activation value corresponding to the segmented data;

[0183] The back propagation module 33 is used to perform a back propagation calculation on any one of the multiple segmented data based on the activation value corresponding to the segmented data through the GPU to obtain the gradient data corresponding to the segmented data, and move the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or CPU memory.

[0184] In one possible implementation, the forward propagation module 32 is configured to:

[0185] Reading the segmented data from the non-volatile memory to the CPU memory;

[0186] The segmented data is read from the CPU memory to the video memory.

[0187] In a possible implementation, the forward propagation module 32 is further configured to:

[0188] In response to starting a new round of training iteration, reading the latest model parameters of the target large model from the non-volatile memory to the CPU memory;

[0189] The latest model parameters are read from the CPU memory to the video memory.

[0190] In one possible implementation, GPU computation and data transfer can be performed in parallel;

[0191] Wherein, the GPU calculation includes the forward propagation calculation and the back propagation calculation;

[0192] The data transmission includes at least some of the following types: data transmission between different GPUs, data transmission between a GPU and a CPU, and data transmission between a CPU and a non-volatile memory.

[0193] In one possible implementation, the GPU calculation is performed through a default stream, the data transmission is performed through a prefetch stream, the default stream domain and the prefetch stream can work in parallel, and the default stream performs a create empty tensor operation before performing GPU calculation on each segmented data.

[0194] In a possible implementation, the forward propagation module 32 is further configured to:

[0195] Reading the query-key-value matrix corresponding to the segmented data from the non-volatile memory to the display memory;

[0196] An attention calculation is performed on the segmented data based on a query-key-value matrix corresponding to the segmented data by a GPU to obtain a partial logarithm-sum-exponential operation result corresponding to the segmented data.

[0197] In one possible implementation, the forward propagation module 32 is configured to:

[0198] A multi-layer perceptron calculation is performed on the segmented data by a GPU to obtain a multi-layer perceptron output result corresponding to the segmented data.

[0199] In one possible implementation, the forward propagation module 32 is configured to:

[0200] Segmenting the segmented data in a sequence dimension to obtain multiple data blocks of the segmented data;

[0201] Performing multi-layer perceptron calculations on the multiple data blocks in sequence through the GPU to obtain multi-layer perceptron output results corresponding to the multiple data blocks;

[0202] The multilayer perceptron output result corresponding to the segmented data is determined according to the multilayer perceptron output results corresponding to the multiple data blocks.

[0203] In a possible implementation, the apparatus further includes:

[0204] A moving module, configured to move the activation value corresponding to the segmented data from the video memory to the non-volatile memory or the CPU memory;

[0205] A reading module is used to read the activation value corresponding to the segmented data from the non-volatile memory or the CPU memory to the video memory.

[0206] In a possible implementation, the back propagation module 33 is used to:

[0207] Obtaining the total number of training data in the training data set used to train the target large model, and the sum of squares of the training data in the training data set;

[0208] Determining a normalized root mean square of the segmented data based on the total number and the sum of squares;

[0209] Back propagation calculation is performed based on the activation value corresponding to the segmented data and the normalized root mean square of the segmented data to obtain gradient data corresponding to the segmented data.

[0210] In a possible implementation, the apparatus further includes an updating module, wherein the updating module is configured to:

[0211] Obtaining gradient data corresponding to the plurality of segmented data;

[0212] The model parameters of the target large model are updated according to the gradient data corresponding to the multiple pieces of segmented data, and the updated model parameters of the target large model are written into the non-volatile memory.

[0213] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. Its specific implementation and technical effects can refer to the description of the above method embodiments. For the sake of brevity, they will not be repeated here.

[0214] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the above method. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium.

[0215] The embodiment of the present disclosure further provides a computer program, comprising a computer-readable code. When the computer-readable code is executed in an electronic device, a processor in the electronic device executes the above method.

[0216] An embodiment of the present disclosure further provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0217] An embodiment of the present disclosure also provides an electronic device, comprising: one or more processors; a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0218] The electronic device may be provided as a terminal, a server, or other forms of devices.

[0219] Figure 4FIG1 shows a block diagram of an electronic device 1900 provided by an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal. Figure 4 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.

[0220] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), a graphical user interface operating system launched by Apple (MacOS X TM ), a multi-user, multi-process computer operating system (Unix TM ), a free and open source Unix-like operating system (Linux TM ), an open-source Unix-like operating system (FreeBSD TM ) or similar.

[0221] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0222] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0223] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0224] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0225] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0226] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0227] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0228] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0229] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0230] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0231] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0232] If the technical solutions of the embodiments of the present disclosure involve personal information, the products applying the technical solutions of the embodiments of the present disclosure have clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solutions of the embodiments of the present disclosure involve sensitive personal information, the products applying the technical solutions of the embodiments of the present disclosure have obtained the individual's separate consent before processing the sensitive personal information, and at the same time meet the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information. The personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0233] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A large model training method, characterized in that: include: For any training data of the target large model, the training data is divided into multiple segmented data, and the multiple segmented data are stored in a non-volatile memory, wherein the multiple segmented data are sequentially subjected to forward propagation calculation and backpropagation calculation; For any one of the multiple pieces of segmented data, read the segmented data from the non-volatile memory to a video memory, and perform a forward propagation calculation on the segmented data through the GPU to obtain an activation value corresponding to the segmented data; For any one of the multiple pieces of segmented data, the GPU performs a backpropagation calculation based on the activation value corresponding to the segmented data to obtain gradient data corresponding to the segmented data, and moves the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or CPU memory.

2. The method according to claim 1, characterized in that The step of reading the segmented data from the non-volatile memory to the video memory comprises: Reading the segmented data from the non-volatile memory to the CPU memory; The segmented data is read from the CPU memory to the video memory.

3. The method according to claim 1, characterized in that The method further comprises: In response to starting a new round of training iteration, reading the latest model parameters of the target large model from the non-volatile memory to the CPU memory; The latest model parameters are read from the CPU memory to the video memory.

4. The method according to claim 1, wherein GPU calculation and data transfer can be performed in parallel; Wherein, the GPU calculation includes the forward propagation calculation and the back propagation calculation; The data transmission includes at least some of the following types: data transmission between different GPUs, data transmission between a GPU and a CPU, and data transmission between a CPU and a non-volatile memory.

5. The method according to claim 4, characterized in that The GPU calculation is performed through a default stream, and the data transmission is performed through a prefetch stream. The default stream domain and the prefetch stream can work in parallel, and the default stream performs a creation of an empty tensor operation before performing GPU calculation on each segment data.

6. The method according to any one of claims 1 to 5, characterized in that The method further includes: reading the query-key-value matrix corresponding to the segmented data from the non-volatile memory to the display memory; The GPU performs a forward propagation calculation on the segmented data to obtain an activation value corresponding to the segmented data, including: performing an attention calculation on the segmented data based on a query-key-value matrix corresponding to the segmented data by the GPU to obtain a partial logarithm-sum-exponential operation result corresponding to the segmented data.

7. The method according to any one of claims 1 to 5, characterized in that The performing a forward propagation calculation on the segmented data by the GPU to obtain an activation value corresponding to the segmented data includes: A multi-layer perceptron calculation is performed on the segmented data by a GPU to obtain a multi-layer perceptron output result corresponding to the segmented data.

8. The method according to claim 7, characterized in that The performing multi-layer perceptron calculation on the segmented data by the GPU to obtain the multi-layer perceptron output result corresponding to the segmented data includes: Segmenting the segmented data in a sequence dimension to obtain multiple data blocks of the segmented data; Performing multi-layer perceptron calculations on the multiple data blocks in sequence through the GPU to obtain multi-layer perceptron output results corresponding to the multiple data blocks; The multilayer perceptron output result corresponding to the segmented data is determined according to the multilayer perceptron output results corresponding to the multiple data blocks.

9. The method according to any one of claims 1 to 5, characterized in that After obtaining the activation value corresponding to the segmented data, and before performing back propagation calculation based on the activation value corresponding to the segmented data by the GPU and obtaining gradient data corresponding to the segmented data, the method further includes: Moving the activation value corresponding to the segmented data from the video memory to the non-volatile memory or the CPU memory; The activation value corresponding to the segmented data is read from the non-volatile memory or the CPU memory to the video memory.

10. The method according to any one of claims 1 to 5, characterized in that The performing back propagation calculation based on the activation value corresponding to the segmented data by the GPU to obtain gradient data corresponding to the segmented data includes: Obtaining the total number of training data in the training data set used to train the target large model, and the sum of squares of the training data in the training data set; Determining a normalized root mean square of the segmented data based on the total number and the sum of squares; Back propagation calculation is performed based on the activation value corresponding to the segmented data and the normalized root mean square of the segmented data to obtain gradient data corresponding to the segmented data.

11. The method according to any one of claims 1 to 5, characterized in that After moving the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or the CPU memory, the method further includes: Obtaining gradient data corresponding to the plurality of segmented data; The model parameters of the target large model are updated according to the gradient data corresponding to the multiple pieces of segmented data, and the updated model parameters of the target large model are written into the non-volatile memory.

12. A large model training device, characterized in that: include: a segmentation module, configured to segment any training data of the target large model into a plurality of segmented data, and store the plurality of segmented data in a non-volatile memory, wherein the plurality of segmented data are sequentially subjected to forward propagation calculation and backpropagation calculation; a forward propagation module, configured to read, for any one of the plurality of segmented data, the segmented data from the non-volatile memory into a video memory, and perform a forward propagation calculation on the segmented data through the GPU to obtain an activation value corresponding to the segmented data; A back propagation module is used to perform a back propagation calculation on any one of the multiple segmented data based on the activation value corresponding to the segmented data through the GPU to obtain gradient data corresponding to the segmented data, and move the gradient data corresponding to the segmented data from the video memory to the non-volatile memory or CPU memory.

13. An electronic device, characterized in that: include: one or more processors; a memory for storing executable instructions; The one or more processors are configured to call the executable instructions stored in the memory to execute the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, characterized in that: When the computer-readable code is executed in an electronic device, a processor in the electronic device executes the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Model training method and related device

    CN115114927A

  • Resource-constrained large model heterogeneous training method, computer equipment and storage medium

    CN119597469A

  • Model training method and apparatus, and computing device

    WO2025060619A1

Cited By

  • Model updating method and device, electronic equipment, medium and product

    CN121480594A

  • Training method and system, parameter updating method, electronic equipment and storage medium

    CN122334383A