Large model distribution training method and device based on large-scale data

By splitting the large model into sub-models and deploying it on multiple devices, the problem of excessive device performance requirements for large model training is solved, memory and performance requirements are reduced, and the model security is ensured through encrypted devices.

CN120163188APending Publication Date: 2025-06-17LONGWOOD VALLEY MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510228831.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Large model training requires too high equipment performance, which causes many devices to fail to meet this requirement.

Method used

By splitting the large model into multiple sub-models and deploying it on multiple computer devices, each sub-model contains several continuous neural network layers, and connecting the computer devices based on large-scale sample data, the model training of the large model to be trained is realized.

Benefits of technology

The computing complexity of each computer device is reduced, and the memory and performance requirements of the device training for model are greatly reduced. At the same time, the core sub-model is deployed through encrypted devices, ensuring the security of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163188A_ABST
    Figure CN120163188A_ABST
Patent Text Reader

Abstract

The invention provides a large-scale data-based large model distribution training method and apparatus. The method comprises the steps of obtaining a to-be-trained large model; a to-be-trained large model is split into a plurality of sub-models, the sub-models are deployed to a plurality of computer devices respectively, and each sub-model comprises a plurality of continuous neural network layers; acquiring large-scale sample data; and based on the large-scale sample data, connecting the computer equipment to realize model training of the to-be-trained large model. According to the method and the device, the large model is split and deployed to different computer devices, so that the calculation complexity of each computer device is reduced, and the device memory requirement and the performance requirement of model training are greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of model training, and more specifically, to a large model distributed training method and device based on large-scale data. Background Art

[0002] Large models have an extremely high number of parameters and can capture complex patterns and features, enabling them to generate high-quality content to meet the corresponding needs in the medical field.

[0003] On the one hand, in the current large model training process, there are extremely high requirements for device memory and device performance, and many devices cannot meet these requirements. Summary of the Invention

[0004] The problem solved by this application is that large model training has excessively high requirements for device performance.

[0005] To solve the above problems, in the first aspect of this application, a large model distributed training method based on large-scale data is provided, including:

[0006] Obtain the large model to be trained;

[0007] Split the large model to be trained into multiple sub-models and deploy them to multiple computer devices respectively, where each sub-model contains several consecutive neural network layers;

[0008] Obtain large-scale sample data;

[0009] Based on the large-scale sample data, connect the computer devices to implement the model training of the large model to be trained.

[0010] In the second aspect of this application, a large model distributed training device based on large-scale data is provided, which includes:

[0011] A model acquisition module for obtaining the large model to be trained;

[0012] A model splitting module for splitting the large model to be trained into multiple sub-models and deploying them to multiple computer devices respectively, where each sub-model contains several consecutive neural network layers;

[0013] A data acquisition module for obtaining large-scale sample data;

[0014] A model training module for connecting the computer devices based on the large-scale sample data to implement the model training of the large model to be trained.

[0015] In the third aspect of this application, an electronic device is provided, which includes: a memory and a processor;

[0016] The memory is used for storing programs;

[0017] The processor, coupled to the memory, is configured to execute the program for:

[0018] Obtain the large model to be trained;

[0019] Split the large model to be trained into multiple sub-models and deploy them to multiple computer devices respectively, where each sub-model contains a number of consecutive neural network layers;

[0020] Obtain large-scale sample data;

[0021] Based on the large-scale sample data, connect the computer devices to implement the model training of the large model to be trained.

[0022] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement the above-mentioned large model distributed training method based on large-scale data.

[0023] In the present application, by splitting the large model and deploying it to different computer devices respectively, the computational complexity of each computer device is reduced, and the device memory requirements and performance requirements for model training are greatly reduced.

[0024] In the present application, by setting up an encryption device to deploy the core sub-model, the security of the core sub-model during training and use is guaranteed, thereby meeting the security requirements of the entire large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a flowchart of the large model distributed training method based on large-scale data according to an embodiment of the present application;

[0026] Figure 2 It is a structural block diagram of the large model distributed training device based on large-scale data according to an embodiment of the present application;

[0027] Figure 3 It is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0028] To make the above objects, features, and advantages of the present application more obvious and understandable, the following detailed description of the specific embodiments of the present application is provided in conjunction with the accompanying drawings. Although the exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.

[0029] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should have the ordinary meanings understood by those skilled in the art to which this application belongs.

[0030] In view of the above problems, this application provides a new large model distributed training solution based on large-scale data, which can solve the problem of excessive requirements for device performance in large model training by deploying to different computer devices respectively.

[0031] An embodiment of this application provides a large model distributed training method based on large-scale data. The specific solution of this method is Figure 1 as shown. This method can be executed by a large model distributed training device based on large-scale data. The large model distributed training device based on large-scale data can be integrated in electronic devices such as computers, servers, computers, server clusters, and data centers. Combining Figure 1 as shown, it is a flowchart of a large model distributed training method according to an embodiment of this application; wherein, the large model distributed training method based on large-scale data includes:

[0032] S101, obtain the large model to be trained;

[0033] S102, split the large model to be trained into multiple sub-models and deploy them to multiple computer devices respectively. Each sub-model contains several consecutive neural network layers;

[0034] S103, obtain large-scale sample data;

[0035] In this application, data is loaded from public data sets (such as Wikipedia, Common Crawl) or custom data sources. Preprocessing operations such as word segmentation, padding, and truncation are performed on the data. The data set is divided into multiple data shards, and each shard is assigned to a group of computing devices.

[0036] S104, based on the large-scale sample data, connect the computer devices to implement the model training of the large model to be trained.

[0037] In this application, by splitting the large model and deploying it to different computer devices respectively, the computational complexity of each computer device is reduced, and the memory requirements and performance requirements of the device for model training are greatly reduced.

[0038] In this application, different parts of the model are distributed to multiple devices to solve the problem of insufficient memory of a single device.

[0039] In one implementation, at least one of the computer devices is an encryption device, and at least one sub-model is a core sub-model; the core sub-model is deployed on the encryption device.

[0040] In this application, an encryption device is set up to deploy the core sub-model, thereby ensuring the security of the core sub-model during training and use, and further meeting the security requirements of the entire large model.

[0041] In this application, the core sub-model is an indispensable part of the large model. Without these parts, the model will not be able to complete its designed functions.

[0042] In this application, by deploying the core sub-model on the encryption device, the security of the large model during training and use can be significantly improved.

[0043] In this application, when transmitting data between the encryption device and other devices, an encryption protocol (such as TLS) is used to protect data security.

[0044] In this application, during training and inference, it is ensured that the calculations and data storage of the core sub-model are always carried out on the encryption device; for example, during training, the encryption device is used to calculate gradients and update the parameters of the core sub-model.

[0045] In one implementation, the large model to be trained is a large-scale language model based on the Transformer architecture.

[0046] The large model to be trained is a large model based on Transformer, which contains multiple Transformer layers, and each Transformer layer is composed of a self-attention mechanism (Self-Attention) and a feed-forward neural network (Feed-Forward Network, FFN).

[0047] The large model to be trained is split by layer, and each layer or several adjacent layers are a sub-model; they are distributed to multiple devices. For example, 24 layers of Transformer are evenly distributed to 6 devices, and each device is responsible for 4 layers.

[0048] In this application, the large model to be trained includes an input embedding layer, a self-attention mechanism, a feed-forward neural network, and an output layer. Any part of the input embedding layer, the self-attention mechanism, and the feed-forward neural network can be used as the core sub-model.

[0049] In one implementation, based on the large-scale sample data, connecting the computer devices to implement the model training of the large model to be trained includes:

[0050] Dividing the training data set into multiple data shards, and each data shard contains several training samples;

[0051] Constructing a forward propagation link and a backward propagation link between computer devices;

[0052] Train the data shards and calculate gradients according to the forward propagation link and the backward propagation link;

[0053] Update the model parameters according to the gradients until the model converges.

[0054] In this application, the computer devices are divided into several groups. Each group of computer devices contains the same number of computer devices, and a large model to be trained is deployed on each group of computer devices; that is, the sub-models of the large model to be trained are deployed in a group of computer devices.

[0055] In this application, data sharding: divide the training data set into multiple subsets, and each subset contains several training samples.

[0056] In this application, each group of computer devices processes one data shard and independently calculates the gradients.

[0057] In this application, ensure that the amount of data processed by each group of computer devices is similar to avoid resource waste.

[0058] In this application, construct the forward propagation link and the backward propagation link between computer devices; construct the forward propagation link and the backward propagation link between a group of computer devices; a link is also set between groups, and this link is used to summarize and share the gradients between different groups.

[0059] During specific training, each group of computer devices is assigned a data shard, and independent training is performed according to the forward propagation link and the backward propagation link to calculate the corresponding gradients; then the gradients of different groups are summarized, the overall gradient is calculated, and distributed to each group, so that each group updates the model parameters according to the overall gradient.

[0060] For example, 24 computer devices are divided into 4 groups, with 6 devices in each group. After the large model is split into different sub-models, they are respectively deployed on 6 devices; then the data shards are also divided into 4 pieces, and each data shard is assigned to a group of devices. Construct the forward propagation link and the backward propagation link between 6 devices in each group, perform one training, and calculate the corresponding gradients; summarize the gradients of 4 groups, calculate the overall gradient, and then distribute it to each device in 4 groups for parameter update.

[0061] In one implementation, the training of the data shards and the calculation of gradients according to the forward propagation link and the backward propagation link include:

[0062] On each computing device, use the corresponding sub-model to perform forward propagation calculation on the data shard to obtain the local output;

[0063] Based on the forward propagation link, pass the local output to the next computing device, and sequentially complete the forward propagation calculation of all sub-models to obtain the prediction data;

[0064] Calculate the loss function according to the prediction data;

[0065] Each computing device calculates the gradient of its corresponding sub-model according to the loss function;

[0066] Based on the backpropagation link, it is passed to the previous computing device to complete the gradient calculation of all sub-models in sequence.

[0067] This step is the gradient calculation of a group of computer devices.

[0068] In this application, the forward propagation link refers to the path through which data is sequentially passed between multiple computing devices; ensuring that each computing device can use its corresponding sub-model to perform forward propagation calculation on the data shard and pass the result to the next computing device.

[0069] In this application, the backpropagation link is a link with the opposite propagation direction to the forward propagation link, which is used to ensure that each computing device can calculate the gradient of its corresponding sub-model according to the loss function and pass the gradient to the previous computing device.

[0070] In this application, in the forward propagation link, each computing device uses its corresponding sub-model to perform forward propagation calculation on the data shard to obtain a local output. For example, device 0 uses the input embedding layer to perform forward propagation calculation on data shard 1 to obtain a local output; the local output is passed to the next computing device. For example, device 0 passes the local output to device 1, and device 1 uses the first layer of Transformer to calculate the local output to obtain a new local output and passes it to device 2, and so on; the forward propagation calculation of all sub-models is completed in sequence to obtain the final prediction data.

[0071] In this application, each computing device calculates the gradient of its corresponding sub-model according to the loss function. For example, device 4 calculates the gradient of the output layer, and device 3 calculates the gradient of the 20th - 24th layer of Transformer, and so on; the gradient is passed to the previous computing device. For example, device 4 passes the gradient to device 3, and device 3 passes the gradient to device 2, and so on; the gradient calculation of all sub-models is completed in sequence.

[0072] In one implementation, according to the forward propagation link and the backpropagation link, training and gradient calculation of the data shard further includes:

[0073] Quantize and compress the gradient of the computer device, compressing the 32-bit floating-point number gradient into an 8-bit integer gradient;

[0074] Synchronize the 8-bit integer gradients calculated by all computer devices to all computer devices;

[0075] Decompress the synchronized 8-bit integer gradients into 32-bit floating-point gradients.

[0076] On each computer device, compress the 32-bit floating-point gradients into 8-bit integer gradients. Synchronize the 8-bit integer gradients on all devices to a preset location (or synchronize to all devices), and decompress them into 32-bit floating-point gradients.

[0077] In this application, through gradient quantization compression, the communication overhead in distributed training can be significantly reduced, thus accelerating the training process.

[0078] Among them, synchronize the 8-bit integer gradients calculated by all computer devices to all computer devices, including: devices within the same data shard group synchronize their corresponding sub-model gradients through All Reduce; global gradients are aggregated through All Reduce among all data shard groups.

[0079] Among them, devices within each data shard group perform All Reduce to aggregate the gradients within the group; all data shard groups (including single-device groups) perform global All Reduce to average the gradients to all devices.

[0080] In one implementation, after dividing the training dataset into multiple data shards, each data shard containing several training samples, it further includes:

[0081] Adopt pipeline parallelism technology to divide the input data into multiple micro-batches;

[0082] Process the micro-batches sequentially on each computing device, and overlap computing and communication;

[0083] During forward propagation and backward propagation, ensure seamless connection of the computing and communication tasks of each computing device.

[0084] In this application, pipeline parallelism overlaps computing and communication by dividing the input data into multiple micro-batches and processing these micro-batches sequentially on multiple computing devices. The specific steps include: data division: divide the training dataset into multiple data shards, and each data shard is further divided into multiple micro-batches. Model splitting: split the model by layer or module and distribute it to multiple computing devices. Pipeline scheduling: process the micro-batches sequentially on each device and overlap computing and communication. Seamless connection: ensure seamless connection of computing and communication tasks during forward propagation and backward propagation.

[0085] In this application, each data shard is further divided into multiple micro-batches (for example, each shard is divided into 8 micro-batches); the size of the micro-batches is adjusted according to the hardware performance and model complexity.

[0086] In this application, forward propagation: Each device processes mini - batches sequentially and passes the output to the next device. For example: Device 0 processes the forward propagation of mini - batch 1 and passes the result to Device 1; Device 1 processes the forward propagation of mini - batch 1 while Device 0 processes the forward propagation of mini - batch 2; and so on until all mini - batches are processed.

[0087] In this application, backpropagation: Each device processes the backpropagation of mini - batches sequentially and passes the gradients to the previous device.

[0088] For example: Device 3 processes the backpropagation of mini - batch 1 and passes the gradients to Device 2; Device 2 processes the backpropagation of mini - batch 1 while Device 3 processes the backpropagation of mini - batch 2; and so on until all mini - batches are processed.

[0089] In this application, asynchronous communication (such as NCCL) and CUDA streams are used to overlap computation and communication.

[0090] For example: When Device 0 processes the forward propagation of mini - batch 1, it asynchronously sends the output of mini - batch 1 to Device 1; When Device 1 processes the forward propagation of mini - batch 1, it simultaneously receives the data sent by Device 0.

[0091] In this application, a scheduler is implemented to ensure seamless connection of the computing and communication tasks of each device.

[0092] For example: Events or barriers are used to synchronize tasks between devices.

[0093] In this application, by combining mini - batches and pipeline parallelism techniques, the efficiency of distributed training can be significantly improved.

[0094] In one implementation, an adaptive adjustment module is further provided after the output layer of the large model to be trained. The processing process of this adaptive model includes:

[0095] Partition the image of the output layer of the large model to obtain independent blocks;

[0096] For each independent block, obtain the first neighborhood blocks and the second neighborhood blocks with different spacings;

[0097] Generate the first feature block based on the independent block and the first neighborhood blocks;

[0098] Generate the second feature block based on the independent block and the second neighborhood blocks;

[0099] Perform feature compression on the first feature block and the second feature block to obtain compressed blocks;

[0100] Traverse all independent blocks and generate a predicted medical image based on the obtained compressed blocks.

[0101] In this application, partitioning the image of the output layer of the large model means dividing the image of the output layer of the large model into corresponding image blocks through a checkerboard; among them, the image blocks can be at the pixel level (that is, each pixel is an image block), or at other levels, and the specific partitioning depends on the actual processing situation.

[0102] In this application, a sliding window or a fixed step size is used to divide the image into blocks of the same size.

[0103] It should be noted here that if the image of the output layer of the large model is a two-dimensional image, it is directly divided into a checkerboard, and each grid is an image block; if the image of the output layer of the large model is a three-dimensional image, a plane is selected for checkerboard partitioning, and each grid is a strip-shaped grid with a lot of depth (the depth is the depth of the three-dimensional image), and this strip-shaped grid is an image block.

[0104] Preferably, in this application, each image block has 100 - 1000 pixels, so as to perform more feature calculations between local regions on the basis of ensuring the generation accuracy and reducing the calculation amount.

[0105] In this application, an image block is selected as an independent block, and the adjacent image blocks above, below, to the left, and to the right of the independent block are the first neighborhood blocks; the image blocks separated by one grid above, below, to the left, and to the right of the independent block are the second neighborhood blocks. The distances between the first neighborhood blocks and the second neighborhood blocks and the independent block are different.

[0106] In this application, neighborhood information is extracted for each independent block to capture local structures.

[0107] In this application, generating the first feature block means generating a local feature representation using the independent block and its first neighborhood blocks, and specifically, it can be: performing convolutional layer and attention layer processing on the independent block and the first neighborhood blocks to obtain the first feature block.

[0108] In this application, the specific structures and specific parameters of the convolutional layer and the attention layer can be obtained according to the training data or determined according to the actual situation.

[0109] It should be noted that in this application, there are four first neighborhood blocks and multiple first feature blocks.

[0110] In this application, the independent block and the first neighborhood blocks are processed through a convolutional layer and an attention layer to obtain the first feature block. The specific process is as follows: The independent block and four neighborhood blocks are concatenated together to form a multi-channel input, and a convolutional layer is used to extract features from the concatenated blocks; The self-attention mechanism or the channel attention mechanism is used to enhance important features, calculate the attention weights, and weight the output of the convolutional layer to enhance important features; The output of the attention layer is split into multiple feature blocks, and each feature block corresponds to the processing results of the independent block and at least one neighborhood block.

[0111] In this application, the second feature block is generated to generate a more extensive local feature representation using the independent block and its second neighborhood blocks. The specific generation process is the same as that of the first feature block, except that the parameters of the convolutional layer and the attention layer are different.

[0112] In this application, the generated feature blocks are compressed into a more compact representation to reduce the computational amount and retain key information. Pooling operations (such as max pooling or average pooling) or fully connected layers are used for feature compression.

[0113] In this way, through compression, multiple first feature blocks and second feature blocks are compressed into a compressed block, which corresponds to the independent block in size and position and is used to replace the independent block. All image blocks are replaced by compressed blocks to obtain the predicted medical image.

[0114] In this application, each image block of the output layer image of the large model is traversed in a traversal manner to obtain the corresponding compressed block.

[0115] In this application, for the image blocks / independent blocks near the edge, their first neighborhood blocks and second neighborhood blocks are incomplete. At this time, the first neighborhood blocks and second neighborhood blocks in the relative positions are copied for complementation. For example, if the first neighborhood block above the independent block does not exist, the first neighborhood block below is copied and used as the block above.

[0116] In this application, through complementation, the processing accuracy of the edge image blocks is greatly improved.

[0117] In this application, the adaptive adjustment module is used to capture the similarity relationship between local regions, thereby enhancing the feature representation. For images with rich textures or complex structures, high-quality medical images can be generated.

[0118] In this application, by adding the adaptive adjustment module, the large model can overcome its own defects and achieve the generation of high-quality medical images. In this way, through distributed training, a large model that can output high-quality medical images can be obtained.

[0119] The embodiments of the present application provide a large model distributed training device based on large-scale data, which is used to execute a large model distributed training method described in the above content of the present application. The following is a detailed description of the large model distributed training device based on large-scale data.

[0120] As Figure 2 shown, the large model distributed training device based on large-scale data includes:

[0121] A model acquisition module 101, which is used to acquire the large model to be trained;

[0122] A model splitting module 102, which is used to split the large model to be trained into multiple sub-models and deploy them to multiple computer devices respectively. Each sub-model contains several consecutive neural network layers;

[0123] A data acquisition module 103, which is used to acquire large-scale sample data;

[0124] A model training module 104, which is used to connect the computer devices based on the large-scale sample data to implement the model training of the large model to be trained.

[0125] In one implementation, at least one of the computer devices is an encryption device, and at least one sub-model is a core sub-model; the core sub-model is deployed on the encryption device.

[0126] In one implementation, the large model to be trained is a large-scale language model based on the Transformer architecture.

[0127] In one implementation, the model training module 104 is further used for:

[0128] Dividing the training data set into multiple data shards, each data shard contains several training samples; constructing a forward propagation link and a backward propagation link between computer devices; training and gradient calculation are performed on the data shards according to the forward propagation link and the backward propagation link; the model parameters are updated according to the gradient until the model converges.

[0129] In one implementation, the model training module 104 is further used for:

[0130] On each computing device, the corresponding sub-model is used to perform forward propagation calculation on the data shard to obtain a local output; the local output is passed to the next computing device based on the forward propagation link to complete the forward propagation calculation of all sub-models in sequence, and the prediction data is obtained; the loss function is calculated according to the prediction data; each computing device calculates the gradient of the corresponding sub-model according to the loss function; and it is passed to the previous computing device based on the backward propagation link to complete the gradient calculation of all sub-models in sequence.

[0131] In one embodiment, the model training module 104 is further configured to:

[0132] Quantize and compress the gradients of the computer device, compress the 32-bit floating-point number gradients into 8-bit integer gradients; aggregate the 8-bit integer gradients calculated by all computer devices and synchronize them to all computer devices; decompress the synchronized 8-bit integer gradients into 32-bit floating-point number gradients.

[0133] In one embodiment, the model training module 104 is further configured to:

[0134] Adopt pipeline parallel technology to divide the input data into multiple micro-batches; process the micro-batches sequentially on each computing device, and overlap the computing and communication; ensure seamless connection of the computing and communication tasks of each computing device during the forward propagation and backward propagation processes.

[0135] The large model distributed training device based on large-scale data provided by the above embodiments of the present application and the large model distributed training method based on large-scale data provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.

[0136] The internal functions and structures of a large model distributed training device based on large-scale data are described above. As Figure 3 shown, in practice, the large model distributed training device based on large-scale data can be implemented as an electronic device, including: a memory 301 and a processor 303.

[0137] The memory 301 can be configured to store programs.

[0138] In addition, the memory 301 can also be configured to store various other data to support operations on the electronic device. Examples of these data include instructions for any application program or method for operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc.

[0139] The memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0140] The processor 303 is coupled to the memory 301 and is configured to execute the programs in the memory 301 for:

[0141] Obtain the large model to be trained;

[0142] The large model to be trained is split into multiple sub-models and deployed to multiple computer devices respectively, and each sub-model contains several consecutive neural network layers;

[0143] Obtain a large amount of sample data;

[0144] Based on the large amount of sample data, connect the computer devices to implement the model training of the large model to be trained.

[0145] In one implementation, at least one of the computer devices is an encryption device, and at least one sub-model is a core sub-model; the core sub-model is deployed on the encryption device.

[0146] In one implementation, the large model to be trained is a large-scale language model based on the Transformer architecture.

[0147] In one implementation, the processor 303 is further configured to:

[0148] Divide the training data set into multiple data shards, each data shard contains several training samples; construct the forward propagation link and the backward propagation link between computer devices; perform training and gradient calculation on the data shards according to the forward propagation link and the backward propagation link; update the model parameters according to the gradient until the model converges.

[0149] In one implementation, the processor 303 is further configured to:

[0150] On each computing device, use the corresponding sub-model to perform forward propagation calculation on the data shard to obtain a local output; transmit the local output to the next computing device based on the forward propagation link, and sequentially complete the forward propagation calculation of all sub-models to obtain prediction data; calculate the loss function according to the prediction data; each computing device calculates the gradient of the corresponding sub-model according to the loss function; transmit it to the previous computing device based on the backward propagation link, and sequentially complete the gradient calculation of all sub-models.

[0151] In one implementation, the processor 303 is further configured to:

[0152] Quantize and compress the gradients of the computer devices, compress the 32-bit floating-point number gradients into 8-bit integer gradients; combine the 8-bit integer gradients calculated by all computer devices and synchronize them to all computer devices; decompress the synchronized 8-bit integer gradients into 32-bit floating-point number gradients.

[0153] In one implementation, the processor 303 is further configured to:

[0154] Using pipelining parallel technology, the input data is divided into multiple micro-batches; the micro-batches are processed sequentially on each computing device, and the computing and communication are overlapped; during the forward propagation and backward propagation, it is ensured that the computing and communication tasks of each computing device are seamlessly connected.

[0155] In this application, Figure 3 only some components are schematically shown, which does not mean that the electronic device only includes Figure 3 the components shown.

[0156] The electronic device provided in this embodiment is based on the same inventive concept as a large model distributed training method based on large-scale data provided in the embodiments of this application, and has the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0157] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CDROM, optical storage, etc.) containing computer-usable program code.

[0158] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0159] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0160] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0161] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0162] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (FlashRAM). The memory is an example of computer-readable media.

[0163] This application also provides a computer-readable storage medium corresponding to a large model distributed training method based on large-scale data provided in the foregoing embodiment. A computer program (i.e., program product) is stored thereon. When the computer program is run by a processor, it will execute a large model distributed training method based on large-scale data provided in any of the foregoing embodiments.

[0164] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CDROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0165] The computer-readable storage medium provided in the above embodiments of this application and a large model distributed training method based on large-scale data provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

[0166] It should be noted that in the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this description.

[0167] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0168] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A large-model distributed training method based on large-scale data, characterized in that: include: Get the large model to be trained; Split the large model to be trained into multiple sub-models and deploy them on multiple computer devices respectively. Each sub-model contains several consecutive neural network layers. Obtain large-scale sample data; Based on large-scale sample data, the computer device is connected to realize model training of the large model to be trained.

2. The large-scale model distribution training method based on large-scale data according to claim 1 is characterized in that: At least one of the computer devices is an encryption device, and at least one sub-model is a core sub-model; the core sub-model is deployed on the encryption device.

3. The large-scale model distribution training method based on large-scale data according to claim 1 is characterized in that: The large model to be trained is a large-scale language model based on the Transformer architecture.

4. The large-scale model distribution training method based on large-scale data according to any one of claim 13, characterized in that: The method of connecting the computer device based on the large-scale sample data to realize the model training of the large model to be trained includes: Divide the training data set into multiple data shards, each of which contains several training samples; Construct forward propagation links and reverse propagation links between computer devices; According to the forward propagation link and the reverse propagation link, training and gradient calculation are performed on the data slices; Update the model parameters according to the gradient until the model converges.

5. The large-scale model distribution training method based on large-scale data according to claim 4 is characterized in that: The training and gradient calculation of the data slices according to the forward propagation link and the reverse propagation link include: On each computing device, use the corresponding sub-model to perform forward propagation calculations on the data shards to obtain local outputs; Based on the forward propagation link, the local output is passed to the next computing device, and the forward propagation calculations of all sub-models are completed in sequence to obtain the predicted data; Calculate the loss function based on the predicted data; Each computing device calculates the gradient of the corresponding sub-model according to the loss function; Based on the back-propagation link passed to the previous computing device, the gradient calculation of all sub-models is completed in sequence.

6. The large-scale model distribution training method based on large-scale data according to claim 5 is characterized in that: The training and gradient calculation of the data slices according to the forward propagation link and the reverse propagation link also includes: Quantize and compress the gradient of the computer device, compressing the 32-bit floating point gradient into an 8-bit integer gradient; The 8-bit integer gradients calculated by all computer devices are synchronized to all computer devices; Decompresses synchronized 8-bit integer gradients into 32-bit floating point gradients.

7. The large-scale model distribution training method based on large-scale data according to claim 4 is characterized in that: After dividing the training data set into a plurality of data slices, each data slice comprising a plurality of training samples, the method further includes: Using pipeline parallel technology, the input data is divided into multiple micro-batches; Process micro-batches sequentially on each computing device, overlapping computation and communication; During the forward and backward propagation processes, ensure that the computing and communication tasks of each computing device are seamlessly connected.

8. A large-scale model distributed training device based on large-scale data, characterized in that: include: A model acquisition module, which is used to obtain a large model to be trained; A model splitting module is used to split the large model to be trained into multiple sub-models and deploy them to multiple computer devices respectively. Each sub-model contains several consecutive neural network layers. A data acquisition module, which is used to acquire large-scale sample data; The model training module is used to connect the computer device to realize model training of the large model to be trained based on large-scale sample data.

9. An electronic device, characterized in that: include: Memory and processor; The memory is used to store programs; The processor, coupled to the memory, is configured to execute the program to: Get the large model to be trained; Split the large model to be trained into multiple sub-models and deploy them on multiple computer devices respectively. Each sub-model contains several consecutive neural network layers. Obtain large-scale sample data; Based on large-scale sample data, the computer device is connected to realize model training of the large model to be trained.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by the processor to implement a large-model distributed training method based on large-scale data as described in any one of claim 17.