Distributed Training Method of Recommendation Model Based on Double-Layer Index Embedding Layer and GPU

By constructing a double-layer index on the server node and the computing node and optimizing the reading and writing of the embedding layer vectors using ping-pong buffers and data prefetch pipelines, the problem of GPU-accelerated embedding layer computing difficulties in the prior art is solved, and the throughput and scalability of model training is improved.

CN114021736BActive Publication Date: 2025-07-18SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111284302.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-01
Publication Date
2025-07-18
Estimated Expiration
2041-11-01

AI Technical Summary

Technical Problem

Existing machine learning systems cannot effectively utilize GPU to accelerate embedded layer computing in distributed training, resulting in low model training throughput, poor scalability, and communication hotspot problems caused by sparse access mode.

Method used

A distributed training method for the recommended model based on the two-layer index embedding layer is adopted. By constructing static and dynamic indexes in the server node and the computing node, a sampler is introduced to formulate a sharding strategy, and the reading and writing of the embedding layer vectors are optimized in the computing node using a ping-pong buffer and data prefetch pipeline.

Benefits of technology

Without reducing the prediction performance of the model, the total throughput of recommended model training is improved, the scalability of distributed training is enhanced, and the training of large-scale embedded layer recommendation models is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114021736B_ABST
    Figure CN114021736B_ABST
Patent Text Reader

Abstract

The present invention provides a distributed training method and GPU for a recommendation model based on a double-layer index embedding layer. The method includes: constructing two layers of indexes for the embedding layer vectors in the server node and the computing node respectively, wherein the index in the server node is a static index, and the index in the computing node is a dynamic index; introducing a sampler in the server node and formulating a sharding strategy according to the sampled data; introducing a ping-pong buffer in the computing node to read and write the embedding layer vectors, and prefetching the embedding layer vectors based on the input of adjacent iterations to form a data prefetching pipeline. The present invention can improve the total throughput of the recommendation model training, enhance the scalability of distributed training, and effectively support the training of large-scale embedding layer recommendation models on the premise of ensuring that the model prediction performance does not decline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of distributed computing and machine learning, and particularly to a distributed training method for a recommendation model based on a double-layer index embedding layer and a GPU. Background Art

[0002] With the wide popularization of the Internet, the user scale and the amount of information are increasing day by day, and the problem of information overload has become increasingly prominent. The recommendation system is one of the solutions to solve the problem of information overload. By analyzing the historical behavior information of users, the recommendation system guides users to discover their own information needs and provides personalized recommendation services for users. It has become an important part of people's online life and is a research direction widely concerned by academia and industry. Industrial recommendation models usually have the characteristics of huge training data volume, complex models, and strong timeliness. Distributed training and computing hardware acceleration have become necessary conditions for solving large-scale recommendation models. However, the current machine learning systems have the following problems: First, when the one-hot encoded input features (such as user IDs, product IDs, etc.) reach hundreds of millions of dimensions, the scale of the recommendation model based on the embedding layer far exceeds the size of the GPU video memory, and the GPU cannot be used to accelerate the calculation of the embedding layer; Second, the parameters of the embedding layer; Third, the access pattern of the embedding layer parameters is sparse compared with common network layers such as the fully connected layer and the convolutional layer, that is, the embedding layer vectors corresponding to a small number of input features are accessed and updated in each training iteration. This access pattern will cause hot communication problems and affect the training throughput.

[0003] Currently, mainstream deep learning frameworks (TensorFlow, PyTorch, MXNet, etc.) do not use GPU to accelerate the calculation of the embedding layer during distributed training, and do not optimize the redundant communication and sparse access pattern of the embedding layer, resulting in problems such as low model training throughput and poor scalability in the distributed scenario. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a distributed training method for a recommendation model based on a double-layer index embedding layer and a GPU, which is used to solve the technical problems of low model training throughput and poor scalability in the existing centralized and distributed scenarios.

[0005] To achieve the above and other related objectives, the present invention provides a distributed training method for a recommendation model based on a double-layer index embedding layer. The method includes: constructing two layers of indexes for the embedding layer vectors at the server node and the computing node respectively, wherein the index within the server node is a static index, and the index within the computing node is a dynamic index; introducing a sampler at the server node and formulating a sharding strategy according to the sampled data; introducing a ping-pong buffer at the computing node to read and write the embedding layer vectors, and prefetching the embedding layer vectors based on the input of adjacent iterations to form a data prefetching pipeline.

[0006] In an embodiment of the present invention, the constructing two layers of indexes for the embedding layer vectors at the server node and the computing node respectively includes: initializing the complete embedding layer at the server node through a parameter server architecture, and constructing the first layer of index at the server node for the parameter synchronization update of the embedding layer; preallocating video memory space related to the input dimension and the number of worker nodes within the computing node, and aggregating the inputs of each worker node within each iteration period to dynamically construct the second layer of index to support the embedding layer calculation.

[0007] In an embodiment of the present invention, it further includes: configuring the communication method based on the index structure: aggregating the inputs of each worker node first within the computing node, and initializing a communication buffer to replace the communication between the worker node and the server node.

[0008] In an embodiment of the present invention, the formulating a sharding strategy according to the sampled data includes: the sampler sampling the training set according to the sampling rule to obtain the access frequency of the embedding layer; then screening out the hot parameters of the embedding layer through a preset threshold, and recording the hot parameters through a hot bitmap; according to the access frequency of the hot parameters, re-initializing the hot parameters at each server node through a replaceable hot-spot-aware sharding strategy, and establishing a sharding mapping table from the embedding layer index to the server node; during training, judging whether the current embedding vector is a hot parameter based on the hot bitmap, and if so, pulling and updating the embedding vector according to the sharding mapping table and bypassing the original mapping.

[0009] In an embodiment of the present invention, the introducing a ping-pong buffer at the computing node to read and write the embedding layer vectors includes: establishing a ping-pong buffer according to the preset input dimension, including a readable buffer and a writable buffer; the readable buffer serves the embedding layer calculation of the current iteration, and the writable buffer writes the corresponding embedding layer vectors based on the input of the next iteration.

[0010] In an embodiment of the present invention, the data prefetching pipeline prefetches the unupdated embedding layer vectors in advance according to the input of the current iteration and the next iteration.

[0011] In an embodiment of the present invention, the process of the data prefetch pipeline includes: when calculating each working node within a computing node, first read in the input of the current iteration to construct the corresponding second-layer index and place it in the readable buffer, and then read in the input of the next iteration to construct an index and place it in the writable buffer; when communicating with the server node, pull the embedding vectors corresponding to all the indexes in the readable buffer into the readable buffer and start model calculation; during model calculation, start a communication thread to communicate with the server node for the second time, and request the indexes in the writable buffer that have not been accessed from the server node.

[0012] In an embodiment of the present invention, the server node counts the first communication. When the number of communications is equal to the number of working nodes, it allows processing the second communication of each working node; thereafter, the server node checks the access bitmap to determine whether it was accessed in the previous iteration, returns the embedding layer parameters that have not been accessed among them, and counts the second communication. When it is equal to the number of working nodes, the bitmap is reset.

[0013] In an embodiment of the present invention, the communication thread of the computing node holds a lock during the second communication process until the unaccessed embedding layer is returned in the next iteration request and copied to the corresponding position in the writable buffer before releasing the lock, and the calculation of the next iteration needs to wait for the lock to be released before proceeding.

[0014] An embodiment of the present invention also provides a GPU, which applies the distributed training method of the recommendation model based on the double-layer index embedding layer as described above.

[0015] As described above, a distributed training method and a GPU of a recommendation model based on a double-layer index embedding layer of the present invention have the following beneficial effects:

[0016] The present invention can improve the total throughput of the recommendation model training, enhance the scalability of distributed training, and effectively support the training of large-scale embedding layer recommendation models on the premise of ensuring that the model prediction performance does not decline. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It shows the overall flowchart of the distributed training method of the recommendation model based on the double-layer index embedding layer in an embodiment of the present invention.

[0018] Figure 2 It shows the hardware system architecture diagram applied by the distributed training method of the recommendation model based on the double-layer index embedding layer in an embodiment of the present invention.

[0019] Figure 3 It shows the double-layer index structure diagram of the distributed training method of the recommendation model based on the double-layer index embedding layer in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention.

[0021] The purpose of the embodiments of the present invention is to provide a distributed training method and server for a recommendation model based on a double-layer index embedding layer, which is used to solve the technical problems of low throughput and poor scalability in model training in existing centralized distributed scenarios.

[0022] The embodiments of the present invention relate to a dense neural network connection, a recurrent neural block, and a hierarchical multi-scale update mechanism to better learn the long-term dependencies and hierarchical structures in the text, thereby improving the accuracy of text classification.

[0023] An embodiment of the present invention provides a distributed training method for a recommendation model based on a double-layer index embedding layer, including: designing an embedding layer and related operators with a double-layer index structure, which accelerates the calculation of the embedding layer and reduces the redundant communication of the embedding layer; designing a sampler and a hot-spot-aware sharding strategy based on the parameter server architecture, which avoids the communication delay caused by hot-spot access to the embedding layer in distributed training; designing a data prefetch pipeline and a ping-pong buffer based on the double-layer index structure, which reduces the synchronous waiting delay for pulling parameters. The embodiments of the present invention can improve the total throughput of the recommendation model training and enhance the scalability of distributed training on the premise of ensuring that the model prediction performance does not decline, and effectively support the training of large-scale embedding layer recommendation models.

[0024] The following will elaborate in detail on the principle and implementation of a distributed training method and server for a recommendation model based on a double-layer index embedding layer in this embodiment, so that those skilled in the art can understand a distributed training method and server for a recommendation model based on a double-layer index embedding layer in this embodiment without creative labor.

[0025] As Figure 1 shown, this embodiment provides a distributed training method for a recommendation model based on a double-layer index embedding layer. The distributed training method for a recommendation model based on a double-layer index embedding layer includes:

[0026] Step S100, constructing two layers of indexes for the embedding layer vectors on the server node and the computing node respectively, where the index inside the server node is a static index and the index inside the computing node is a dynamic index;

[0027] Step S200, introducing a sampler on the server node and formulating a sharding strategy according to the sampled data;

[0028] Step S300: Introduce ping-pong buffers in the computing node to read and write the embedding layer vectors, and prefetch the embedding layer vectors based on the inputs of adjacent iterations to form a data prefetch pipeline.

[0029] The following combines Figure 2 and Figure 3 to elaborate on the above steps S100 to S300 of this embodiment in detail.

[0030] Step S100: Construct two layers of indexes for the embedding layer vectors in the server node and the computing node respectively. Among them, the index in the server node is a static index, and the index in the computing node is a dynamic index.

[0031] In this embodiment, as Figure 2 and Figure 3 shown, the construction of two layers of indexes for the embedding layer vectors in the server node and the computing node respectively includes: initializing the complete embedding layer in the server node through a parameter server architecture, and constructing the first layer of index in the server node for parameter synchronization update of the embedding layer; pre-allocating video memory space related to the input dimension and the number of worker nodes in the computing node, and aggregating the inputs of each worker node in each iteration cycle to dynamically construct the second layer of index to support embedding layer calculation.

[0032] In distributed training, each computing node may have multiple GPUs, and each GPU participates in distributed training as a logically working node and communicates independently with the server node. Repeated input features among the working nodes within the node will cause redundant transmissions, affecting the throughput of distributed training.

[0033] Specifically, in this embodiment, it further includes: configuring the communication method based on the index structure: aggregating the inputs of each worker node in the computing node first, and initializing the communication buffer to replace the communication between the worker node and the server node to reduce communication redundancy.

[0034] Computing-related operators: Include the Lookup operator, Sort&Unique operator, and Backward operator related to the embedding layer. The Lookup operator performs parallel binary search on the index through the GPU based on the second layer of index in the computing node; the Sort&Unique operator provides the function of sorting and de-duplicating according to the input to dynamically construct the second layer of index. It merges the inputs of each computing unit through an aggregation communication primitive, and then calls the Sort and Unique operations of CUDA Thrust to complete the construction; the Backward operator calculates the gradient corresponding to the embedding layer vector in the GPU based on the second layer of index within the node, and accumulates it by looking up the corresponding gradient storage location within the computing node through the index.

[0035] Step S200: Introduce a sampler at the server node and formulate a sharding strategy based on the sampled data.

[0036] In this embodiment, the sampler and the hot-spot aware sharding strategy based on the parameter server architecture: Introduce a sampler at the server node and formulate a sharding strategy based on the sampled data to avoid communication hot-spot problems.

[0037] Specifically, in this embodiment, the formulating the sharding strategy based on the sampled data includes: The sampler samples the training set according to the sampling rule to obtain the access frequency of the embedding layer; then filters out the hot-spot parameters of the embedding layer through a preset threshold, and records the hot-spot parameters through a hot-spot bitmap; according to the access frequency of the hot-spot parameters, re-initialize the hot-spot parameters at each server node through a replaceable hot-spot aware sharding strategy, and establish a sharding mapping table from the embedding layer index to the server node; during training, judge whether the current embedding vector is a hot-spot parameter based on the hot-spot bitmap, if so, pull and update the embedding vector according to the sharding mapping table, and bypass the original mapping.

[0038] That is, the sampler samples the training set according to the sampling rule to obtain the access frequency of the embedding layer, then filters out the hot-spot parameters of the embedding layer (the embedding layer vectors with access frequencies higher than the threshold are regarded as hot-spot parameters) through a preset threshold, and records the hot-spot parameters through a bitmap. According to the access frequency of the hot-spot parameters, re-initialize the hot-spot parameters at each server node through a replaceable hot-spot aware sharding strategy, and establish a sharding mapping table from the embedding layer index → server node. During training, first check the hot-spot bitmap to judge whether the current embedding vector is a hot-spot parameter, if so, pull and update the embedding vector according to the sharding mapping table, and bypass the original mapping.

[0039] In this embodiment, the hot-spot aware sharding algorithm is adopted in formulating the sharding strategy based on the sampled data:

[0040] First, sort the access frequencies of the hot-spot parameters from large to small. After sorting, calculate the best average access frequency as the available space according to the number of worker nodes, and it is necessary to judge whether the access frequency of each embedding vector can be loaded into the available space. If it is greater than the average access frequency, after allocating it to a certain server node, no remaining embedding vectors will be allocated to the same server node. If it is less than the available space of the current server node, allocate it to the current server and subtract the available space.

[0041] Step S300: Introduce a ping-pong buffer at the computing node to read and write embedding layer vectors, and prefetch the embedding layer vectors based on the input of adjacent iterations to form a data prefetch pipeline.

[0042] In this embodiment, the introduction of the ping-pong buffer for reading and writing the embedding layer vectors in the computing node includes: establishing a ping-pong buffer according to a preset input dimension, including a readable buffer and a writable buffer; the readable buffer serves the embedding layer calculation of the current iteration, and the writable buffer writes the corresponding embedding layer vectors based on the input of the next iteration.

[0043] In this embodiment, the data prefetch pipeline pulls the unupdated embedding layer vectors in advance according to the inputs of the current iteration and the next iteration, reduces the synchronization waiting time, and improves the scalability of distributed training.

[0044] In this embodiment, the process of the data prefetch pipeline includes: when each worker node in the computing node, first reads the input of the current iteration to construct the corresponding second-layer index and puts it into the readable buffer, and then reads the input of the next iteration to construct the index and puts it into the writable buffer; when communicating with the server node, pulls the embedding vectors corresponding to all the indexes in the readable buffer into the readable buffer and starts the model calculation; during the model calculation, starts a communication thread to communicate with the server node for the second time, and requests the indexes in the writable buffer that have not been accessed (the server node stores an access bitmap).

[0045] In this embodiment, the server node counts the first communication. When the communication count is equal to the number of worker nodes, it allows processing the second communication of each worker node; then, the server node checks the access bitmap to determine whether it was accessed in the previous iteration, returns the embedding layer parameters that have not been accessed, and counts the second communication. When it is equal to the number of worker nodes, it resets the bitmap.

[0046] In this embodiment, the communication thread of the computing node holds the lock during the second communication, and does not release it until the unaccessed embedding layer is returned in the next iteration request and copied to the corresponding position in the writable buffer, and the calculation of the next iteration needs to wait for the lock to be released before it can proceed.

[0047] To enable those skilled in the art to further understand the distributed training method of the recommendation model based on the double-layer index embedding layer in this embodiment, the implementation process of the distributed training method of the recommendation model based on the double-layer index embedding layer in this embodiment is described below.

[0048] In the initialization stage of the model, the initialization process of the recommendation model with an embedding layer is divided into two parts: the initialization of the embedding layer and the initialization of the remaining network. The remaining network is initialized locally at the worker nodes. The embedding layer only allocates storage space for the indices and the ping-pong buffer (whose size is equal to the minimum of the input dimension and the embedding layer dimension) corresponding to the embedding layer locally, and the meta-information of the embedding layer is transmitted to the scheduler module of the parameter server through the master worker node in the parameter server architecture. At the same time, the sampler samples the training set to obtain the access frequencies of each embedding vector. The scheduler filters the hot parameters according to the access frequencies of the sampled embedding vectors and constructs an access bitmap to bypass the original mapping relationship. The scheduler first designates a server node to initialize the corresponding embedding layer parameters according to the default hash sharding strategy, and then re-initializes them on the server node according to the access frequencies of the hot parameters, ensuring that the access heat of each node is averaged to avoid the access and calculation of hot parameters becoming a bottleneck. In addition, the initialization stage of the model also includes the initialization of the communication buffers within and between the computing nodes, with a focus on the initialization of the communication buffers for the aggregation operations of each worker node within the computing node and the initialization of the communication buffers for the double-layer index structure between the computing nodes.

[0049] In the training stage of the model, in each iteration cycle, each worker node first reads the input of one iteration, aggregates and sorts it within the computing node, and removes duplicates to generate the indices corresponding to the embedding layer vectors. The parameter manager needs to look up the bitmap of the hot parameters within the computing node according to the indices of the current iteration, communicate with the server node, and pull the corresponding embedding vectors and copy them to the readable ping-pong buffer. During the local model calculation, the readable buffer of the ping-pong buffer provides the embedding layer indices and the corresponding embedding vectors, and the operators in the execution engine additionally support lookups and calculations based on the double-layer index. At the same time, the execution engine manages the data prefetch pipeline and, at the start of the local model calculation, instructs the parameter manager to request the server node to prefetch the data for the next iteration. The server node receives the request for the next iteration, checks the access bitmap, and returns the embedding vectors that have not been accessed in the requested indices. After receiving the response, the parameter manager writes the unaccessed embedding vectors to the corresponding positions in the writable buffer of the ping-pong buffer according to the indices. After the local model completes the backpropagation and updates the parameters, the parameter manager aggregates the parameters to complete the training cycle of the current iteration.

[0050] In addition, an embodiment of the present invention also provides a GPU, which applies the distributed training method of the recommendation model based on the double-layer index embedding layer as described above. The distributed training method of the recommendation model based on the double-layer index embedding layer has been described in detail above and will not be elaborated here.

[0051] In summary, the present invention can maximize the system throughput on the premise that the user can perceive it; the results of the present invention can indirectly provide support for the scheduling technology of GPUs with multiple computing units potentially configured; the results of the present invention can enable the construction of a dynamic task scheduling system for GPU chips with multiple computing units configured, which is of commercial significance, and provide task scheduling services for users. Therefore, the present invention effectively overcomes various drawbacks in the prior art and has high industrial utilization value.

[0052] The above embodiments are only illustrative of the principles and effects of the present invention, and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A distributed training method for a recommendation model based on a double-layer index embedding layer, characterized in that: The method includes: Constructing two - layer indexes for the embedding - layer vectors at the server node and the computing node respectively, where the index in the server node is a static index and the index in the computing node is a dynamic index; Introducing a sampler at the server node and formulating a sharding strategy according to the sampled data; Introducing a ping - pong buffer at the computing node to read and write the embedding - layer vectors, and pre - fetching the embedding - layer vectors based on the inputs of adjacent iterations to form a data pre - fetching pipeline; The formulating the sharding strategy according to the sampled data includes: The sampler samples the training set according to the sampling rules to obtain the access frequencies of the embedding layer; Then, filtering out the hot parameters of the embedding layer through a preset threshold, and recording the hot parameters through a hot - spot bitmap; According to the access frequencies of the hot parameters, re - initializing the hot parameters at each server node through a replaceable hot - spot - aware sharding strategy, and establishing a sharding mapping table from the embedding - layer index to the server nodes; During training, judging whether the current embedding vector is a hot parameter based on the hot - spot bitmap. If so, pulling and updating the embedding vector according to the sharding mapping table and bypassing the original mapping; The introducing a ping - pong buffer at the computing node to read and write the embedding - layer vectors includes: Establishing a ping - pong buffer according to the preset input dimension, including a readable buffer and a writable buffer; The readable buffer serves the embedding - layer calculation of the current iteration, and the writable buffer writes the corresponding embedding - layer vectors based on the input of the next iteration; The data pre - fetching pipeline pre - fetches the un - updated embedding - layer vectors in advance according to the inputs of the current iteration and the next iteration.

2. The distributed training method of the recommendation model based on the double-layer index embedding layer according to claim 1, wherein: The constructing two - layer indexes for the embedding - layer vectors at the server node and the computing node respectively includes: Initializing the complete embedding layer at the server node through a parameter - server architecture, and constructing the first - layer index at the server node for the parameter synchronization and update of the embedding layer; Pre - allocating video - memory space related to the input dimension and the number of worker nodes in the computing node, and aggregating the inputs of each worker node at each iteration cycle to dynamically construct the second - layer index to support the embedding - layer calculation.

3. The distributed training method of the recommendation model based on the double-layer index embedding layer according to claim 1 or 2, characterized in that: It also includes: Configuring the communication method based on the index structure: aggregating the inputs of each worker node first in the computing node, and initializing a communication buffer to replace the communication between the worker nodes and the server node.

4. The distributed training method of the recommendation model based on the double-layer index embedding layer according to claim 1, wherein: The process of the data pre - fetching pipeline includes: When each worker node in the computing node, first reads in the input of the current iteration to construct the corresponding second - layer index and puts it into the readable buffer, and then reads in the input of the next iteration to construct the index and puts it into the writable buffer; When communicating with the server node, pulling the embedding vectors corresponding to all the indexes in the readable buffer into the readable buffer and starting the model calculation; During the model calculation, starting a communication thread to communicate with the server node for the second time, and requesting the indexes in the writable buffer that have not been accessed from the server node.

5. The distributed training method of the recommendation model based on the double-layer index embedding layer according to claim 4, characterized in that: The server node counts the first communication. When the communication count is equal to the number of worker nodes, it allows processing the second communication of each worker node; afterwards, the server node checks the access bitmap to judge whether it was accessed in the previous iteration, returns the un - accessed embedding - layer parameters, and counts the second communication. When it is equal to the number of worker nodes, it resets the bitmap.

6. The distributed training method of the recommendation model based on the double-layer index embedding layer according to claim 1, characterized in that: The communication thread of the computing node holds the lock during the second communication process and releases it only after the next iteration request returns an unaccessed embedding layer and copies it to the corresponding position in the writable buffer. Moreover, the calculation of the next iteration needs to wait for the lock to be released before it can proceed.

7. A GPU, characterized in that: The GPU application is the distributed training method for the recommendation model based on the double-layer index embedding layer as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Information recommendation model training method and information recommendation method and device

    CN110069715A

  • Information recommendation model training method and device, information recommendation method and device, and equipment

    CN112464100A