Method and device for training deep learning recommendation model
By offloading the storage table to the memory device and using the dynamic random access memory of the computing fast-link memory module for prefetching and parallel search, the memory occupancy problem of storage tables is solved, and the training time is reduced and efficiency is improved.
Patent Information
- Application Number
- CN202510156757.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-07-08
AI Technical Summary
When training deep learning recommended models, the storage table occupies a lot of memory, resulting in consumer-grade GPUs being unable to train. Multi-GPU training increases costs and long communication time, and the random access speed of solid-state drives is slow, which affects training performance.
By offloading the storage table to the memory device, the dynamic random access memory of the computing fast-link memory module is used for prefetching and parallel search, reducing training time and improving training efficiency.
The overlap between the training time and the storage table transmission time is achieved, which reduces the training time and improves the training efficiency.
Smart Images

Figure CN120278294A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies. More specifically, the present disclosure relates to a method and apparatus for training a deep learning recommendation model. Background Art
[0002] The scale of machine learning models is growing at a very fast pace, not only for natural language processing (NLP) models such as Generative Pre-Trained Transformer (GPT) models, but also for recommendation models such as deep learning recommendation models (DLRM). Deep learning recommendation models can be large-scale recommendation models and can be applied to data center applications such as search, social media, and entertainment.
[0003] Deep learning recommendation models can be much larger than traditional deep neural networks (DNN). Storage tables (which can refer to, for example, embedding tables, EMB for short) can occupy a large amount of memory and can store the features of raw data during training, and their total size can range from GB (Gigabyte) to dozens of TB (Terabyte). Since storage tables (such as EMB tables) are expected to be stored in high bandwidth memory (HBM) during training, consumer-grade graphics processing units (GPUs) may not be able to train deep learning recommendation models because they may not be able to store such large storage tables (such as EMB tables). These large-scale storage tables (such as EMB tables) may be difficult to deploy in a single graphics processing unit with limited resources.
[0004] The first related solution can deploy all storage tables (such as EMB tables) in dynamic random access memory (DRAM). When a storage table (such as an EMB table) is lost during the training process, the central processing unit (CPU) can transfer the storage table (such as an EMB table) to the graphics processing unit. However, since storage tables (such as EMB tables) may require a large amount of memory, using only dynamic random access memory will make these solutions very costly. This may have a great impact on other applications that require a large amount of dynamic random access memory.
[0005] The second related solution can deploy all storage tables (e.g., EMB tables) separately in multiple graphics processors. The multiple graphics processors run in parallel. The storage tables (e.g., EMB tables) can be divided into multiple parts and stored separately in multiple graphics processors. One graphics processor only processes the part of the storage table (e.g., EMB table) stored therein. However, first, training using multiple graphics processors may increase a lot of costs; second, allocating the same computation to multiple graphics processors may waste the computing resources of the graphics processors. In addition, multiple graphics processors need to communicate with each other, which will hinder the computing of the graphics processors. The more graphics processors there are, the longer the communication time will be.
[0006] The third related solution can deploy all storage tables (e.g., EMB tables) in solid state drives (SSD for short). Just like deploying in dynamic random access memory, when a storage table (e.g., EMB table) is missing during the training process, the central processing unit is allowed to transfer the storage table (e.g., EMB table) to the graphics processor. However, since the random access speed of the solid state drive may be nearly two orders of magnitude slower than that of the dynamic random access memory, training a large-scale deep learning recommendation model may take more time, which may lead to a decline in training performance.
[0007] The fourth related solution can deploy all storage tables (e.g., EMB tables) in a caching deep learning recommendation model (cDLRM for short). A look-ahead window can be set in the central processing unit (dynamic random access memory), and a cache can be set in the graphics processor (high-bandwidth memory). The look-ahead window can be a prefetch module for prefetching storage tables (e.g., EMB tables) of n batch sizes and transferring them to the cache of the graphics processor. After the current training iteration ends, the caching deep learning recommendation model can update the trained storage tables (e.g., EMB tables) to the central processing unit for possible use next time. However, ① all storage tables (e.g., EMB tables) after each prefetch may need to be transferred to the graphics processor without additional parsing, and hot data may be transferred repeatedly; ② the prefetch efficiency of a single central processing unit may not be high, which may cause a large amount of time overhead; ③ the caching deep learning recommendation model may not consider the fluctuations in the training time of the deep learning recommendation model and may not be able to adapt to the training iterations of the model. Summary of the Invention
[0008] An exemplary embodiment of the present disclosure is to provide a method and device for training a deep learning recommendation model DLRM to reduce the training time and improve the training efficiency.
[0009] According to an exemplary embodiment of the present disclosure, a method for training a deep learning recommendation model executed by a first processor is provided, including: unloading one or more storage tables to a memory device; training the deep learning recommendation model based on training data, wherein during the training of the deep learning recommendation model, a storage table prefetched from the memory device is loaded according to a training stage; and storing feature data obtained during training into the loaded storage table.
[0010] Optionally, the memory device may include a dynamic random access memory CMM-D based on a Compute Express Link memory module, and wherein the storage table includes an embedding table.
[0011] Optionally, the prefetched storage table may be loaded based on determining that a second processor prefetches a storage table from the memory device according to a training stage.
[0012] According to an exemplary embodiment of the present disclosure, a method for training a deep learning recommendation model DLRM executed by a second processor is provided, including: prefetching a storage table from a memory device according to a training stage associated with the training of the deep learning recommendation model by a first processor, wherein the storage table is configured to store feature data obtained during training; and transmitting the prefetched storage table to the first processor based on the prefetched storage table.
[0013] Optionally, the storage table may correspond to the next training stage after the current training stage.
[0014] Optionally, the prefetched storage table may include: parallelly searching for a plurality of storage tables corresponding to the next training stage in the memory device; and prefetching a storage table from the memory device based on the plurality of storage tables.
[0015] Optionally, the parallel search may include: selecting corresponding data corresponding to the next training stage in the training data for training the deep learning recommendation model; and parallelly searching for each storage table in the plurality of storage tables in the memory device based on at least a part of the corresponding data.
[0016] Optionally, the prefetched storage table may include: selecting a first storage table from the plurality of storage tables, wherein the popularity associated with the first storage table is less than a popularity threshold; selecting a second storage table whose similarity to the first storage table is greater than or equal to a first similarity threshold; and prefetching the second storage table to replace the first storage table if the popularity of the second storage table is greater than or equal to the popularity threshold.
[0017] Optionally, based on the current storage table replacement rate being less than the replacement rate threshold, the second storage table may be prefetched. Among them, for the prefetched storage table, it also includes: based on the current storage table replacement rate being greater than or equal to the replacement rate threshold, determining not to perform the storage table replacement operation and prefetching the first storage table.
[0018] Optionally, the prefetched storage may include: based on determining that an approximate storage table has been transmitted to the first processor, where the approximate storage table has an approximation degree greater than or equal to a second approximation threshold with a third storage table among the multiple storage tables, determining not to prefetch the third storage table and determining to use the approximate storage table to replace the third storage table.
[0019] Optionally, the determination not to prefetch the third storage table may include: based on the current storage table replacement rate being less than the replacement rate threshold, determining not to prefetch the third storage table. Among them, for the prefetched storage table, it may also include: based on the current storage table replacement rate being greater than or equal to the replacement rate threshold, determining not to perform the storage table replacement operation and prefetching the third storage table.
[0020] According to an exemplary embodiment of the present disclosure, there is provided an apparatus for training a deep learning recommendation model DLRM, including: a storage table unloading unit configured to unload one or more storage tables to a memory device; a model training unit configured to train the deep learning recommendation model based on training data; a storage table loading unit configured to, during the process of training the deep learning recommendation model, load the storage table prefetched from the memory device according to the training phase; and a storage table updating unit configured to store the feature data obtained during training into the loaded storage table.
[0021] Optionally, the memory device may include a dynamic random access memory CMM-D based on a compute express link memory module, and among them, the storage table may include an embedding table.
[0022] Optionally, the prefetched storage table may be loaded based on determining that a second processor prefetches the storage table from the memory device according to the training phase.
[0023] According to an exemplary embodiment of the present disclosure, there is provided an apparatus for training a deep learning recommendation model DLRM, including: a storage table prefetching unit configured to prefetch a storage table from a memory device according to a training phase associated with the training of the deep learning recommendation model by a first processor, where the storage table is configured to store the feature data obtained during training; and a storage table transmitting unit configured to transmit the prefetched storage table to the first processor based on the prefetched storage table.
[0024] Optionally, the storage table may correspond to the next training phase after the current training phase.
[0025] Optionally, the storage table prefetch unit may further be configured to: parallelly search for multiple storage tables corresponding to the next training phase in the memory device; and prefetch the storage tables from the memory device based on the multiple storage tables.
[0026] Optionally, the storage table prefetch unit may be configured to: select corresponding data corresponding to the next training phase in the training data for training the deep learning recommendation model; and based on at least a part of the corresponding data, parallelly search for each storage table in the multiple storage tables in the memory device.
[0027] Optionally, the storage table prefetch unit may be configured to: select a first storage table from the multiple storage tables, where the popularity associated with the first storage table is less than a popularity threshold; select a second storage table whose similarity to the first storage table is greater than or equal to a first similarity threshold; and prefetch the second storage table to replace the first storage table based on the popularity of the second storage table being greater than or equal to the popularity threshold.
[0028] Optionally, the storage table prefetch unit may further be configured to: prefetch the second storage table based on the current storage table replacement rate being less than a replacement rate threshold; and determine not to perform a storage table replacement operation and prefetch the first storage table based on the current storage table replacement rate being greater than or equal to the replacement rate threshold.
[0029] Optionally, the storage table prefetch unit may be configured to: based on determining that an approximate storage table has been transmitted to a first processor, where the approximate storage table has a similarity greater than or equal to a second similarity threshold to a third storage table in the multiple storage tables, determine not to prefetch the third storage table and determine to use the approximate storage table to replace the third storage table.
[0030] Optionally, the storage table prefetch unit may further be configured to: determine not to prefetch the third storage table based on the current storage table replacement rate being less than a replacement rate threshold; and determine not to perform a storage table replacement operation and prefetch the third storage table based on the current storage table replacement rate being greater than or equal to the replacement rate threshold.
[0031] According to an exemplary embodiment of the present disclosure, a system for training a deep learning recommendation model DLRM is provided, including: a memory device; a first processor configured to: unload one or more storage tables to the memory device, train the deep learning recommendation model based on training data, and during the process of training the deep learning recommendation model, load the storage tables prefetched from the memory device according to the training phase, a second processor configured to: prefetch storage tables from the memory device according to the training phase, based on the prefetched storage tables, transmit the prefetched storage tables to the first processor, and the first processor is further configured to store the feature data obtained during training into the loaded storage tables.
[0032] Optionally, the first processor may be a graphics processing unit (GPU), and the second processor may be a central processing unit (CPU).
[0033] Optionally, the memory device may include a Compute Express Link Memory Module - Dynamic Random Access Memory (CMM - D), and wherein the storage table may include an embedded table.
[0034] According to an exemplary embodiment of the present disclosure, there is provided a computer - readable storage medium having a computer program stored thereon, which when executed by a processor, implements a method for training a deep - learning recommendation model according to an exemplary embodiment of the present disclosure.
[0035] According to an exemplary embodiment of the present disclosure, there is provided a computing device, including: at least one processor; at least one memory storing a computer program, which when executed by the at least one processor, implements a method for training a deep - learning recommendation model according to an exemplary embodiment of the present disclosure.
[0036] According to an exemplary embodiment of the present disclosure, there is provided a computer program product, and instructions in the computer program product can be executed by a processor of a computer device to complete a method for training a deep - learning recommendation model according to an exemplary embodiment of the present disclosure.
[0037] A method, device, and system for training a Deep - Learning Recommendation Model (DLRM) according to an exemplary embodiment of the present disclosure, by offloading one or more storage tables to a memory device and training the deep - learning recommendation model based on training data, wherein during the training of the deep - learning recommendation model, a storage table prefetched from the memory device is loaded according to a training stage, and feature data obtained during training is stored in the loaded storage table, thereby overlapping the training time and the transfer time through the prefetch of the storage table, reducing the training time, and thus improving the training efficiency.
[0038] A method, device, and system for training a Deep - Learning Recommendation Model (DLRM) according to an exemplary embodiment of the present disclosure, by prefetching a storage table from a memory device according to a training stage associated with the training of the deep - learning recommendation model by a first processor, wherein the storage table is configured to store feature data obtained during training, and based on the prefetched storage table, transferring the prefetched storage table to the first processor, thereby realizing the prefetch of the storage table and reducing the training time.
[0039] Additional aspects and / or advantages of the general concept of the present disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the general concept of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The above and other objects and features of the exemplary embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the drawings that exemplarily show the embodiments, wherein: Figure 1A A flowchart showing a method for training a deep learning recommendation model DLRM executed by a first processor according to an exemplary embodiment of the present disclosure; Figure 1B A flowchart showing a method for training a deep learning recommendation model DLRM executed by a second processor according to an exemplary embodiment of the present disclosure; Figure 2 A schematic diagram showing a fine-grained lock according to an exemplary embodiment of the present disclosure; Figure 3 A schematic diagram showing a parallel lookup according to an exemplary embodiment of the present disclosure; Figure 4A A block diagram showing an apparatus for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure; Figure 4B A block diagram showing an apparatus for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure; Figure 5 A schematic diagram showing a system for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure; and Figure 6 A schematic diagram showing a computing device according to an exemplary embodiment of the present disclosure. Detailed Description of the Embodiments
[0041] Reference will now be made in detail to the exemplary embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein like reference numerals always refer to like elements. Some specific embodiments will be described below by referring to the drawings to explain various aspects of the present disclosure.
[0042] As is conventional in the art, embodiments are described and illustrated in the figures in terms of functional blocks, units, and / or modules. Those skilled in the art will understand that these blocks, units, and / or modules are physically implemented by electronic (or optical) circuits such as logic circuits, discrete components, microprocessors, hardwired circuits, memory elements, wiring connections, etc., which may be formed using semiconductor-based manufacturing technologies or other manufacturing technologies. In cases where the blocks, units, and / or modules are implemented by a microprocessor or the like, the microprocessor or the like may be programmed with software (e.g., microcode) to perform the various functions discussed herein and may optionally be driven by firmware and / or software. Alternatively, each block, unit, and / or module may be implemented by dedicated hardware or as a combination of dedicated hardware that performs some functions and a processor (e.g., one or more programmed microprocessors and associated circuitry) that performs other functions. Additionally, without departing from the scope hereof, each block, unit, and / or module of an embodiment may be physically divided into two or more interacting and discrete blocks, units, and / or modules. Further, without departing from the scope hereof, the blocks, units, and / or modules of an embodiment may be physically combined into more complex blocks, units, and / or modules.
[0043] Figure 1A A flowchart showing a method for training a deep learning recommendation model DLRM executed by a first processor according to an exemplary embodiment of the present disclosure. Figure 1B A flowchart showing a method for training a deep learning recommendation model executed by a second processor according to an exemplary embodiment of the present disclosure. Here, the first processor may be a graphics processing unit GPU, and the second processor may be a central processing unit CPU, but the embodiments are not limited thereto.
[0044] Referring to Figure 1A , in step S101, the storage table may be unloaded to a memory device. Here, after the deep learning recommendation model initializes the storage table, the first processor (e.g., a graphics processing unit) may unload the storage table from the first processor (e.g., a graphics processing unit) to the memory device. Here, before training the deep learning recommendation model, the training data may be loaded into the deep learning recommendation model located in the first processor (e.g., a graphics processing unit). Loading the training data into the deep learning recommendation model located in the first processor (e.g., a graphics processing unit) may mean that the deep learning recommendation model is to be trained based on the training data. Based on this, the storage table may be unloaded to the memory device in step S101.
[0045] In an exemplary embodiment of the present disclosure, the memory device may include dynamic random access memory based on Compute Express Link (CXL) memory module DRAM (referred to as CMM-D, where Compute Express Link is abbreviated as CXL), and the storage table may include an embedded table (EMB table). For example, the memory device may be dynamic random access memory based on a Compute Express Link memory module, and the storage table may be an embedded table.
[0046] In step S102, the deep learning recommendation model may be trained based on training data. Here, the training data may be used for training the deep learning recommendation model.
[0047] In an exemplary embodiment of the present disclosure, the method further includes: during the training of the deep learning recommendation model, loading a storage table pre-fetched from the memory device according to the training stage, so that the training time can overlap with the transmission time, and the training progress of the deep learning recommendation model in the GPU can be not slowed down, thereby reducing the training time.
[0048] In an exemplary embodiment of the present disclosure, loading the storage table pre-fetched from the memory device according to the training stage may include: based on determining that the second processor pre-fetches a storage table from the memory device according to the training stage, loading the pre-fetched storage table, thereby achieving loading the storage table pre-fetched from the memory device.
[0049] In step S103, the feature data during training may be stored in the loaded storage table. Here, the feature data during the training process are respectively stored in the corresponding storage tables.
[0050] Referring to Figure 1B , in step S111, a storage table is pre-fetched from the memory device according to the training stage associated with the training of the deep learning recommendation model by the first processor for storing the feature data during training.
[0051] In order to overlap the training time with the transmission time of the storage table and not slow down the training progress of the deep learning recommendation model, pre-fetch parameters (for example, but not limited to, the size of each pre-fetch batch, how many batches to pre-fetch, etc.) may be set in advance for each training stage in the present disclosure. Then, during the training process, for example, t batches of storage tables may be fetched from the memory device (for example, dynamic random access memory based on a Compute Express Link memory module) in advance according to the training stage and transmitted to the first processor (for example, in a graphics processor).
[0052] In an exemplary embodiment of the present disclosure, the prefetch storage table may include: prefetching from the memory device the storage table to be used in the next training phase of the current training phase (e.g., the next training phase after the current training phase). Here, by prefetcing in the current training phase the storage table to be used in the next training phase, the waiting between the training process (e.g., the main process) and the prefetch process can be reduced or eliminated, so that the training time of the deep learning recommendation model can overlap with the transmission time of the storage table, without slowing down the training progress of the deep learning recommendation model.
[0053] In the present disclosure, a fine-grained lock can be designed such that there is almost no waiting between the main process of training the deep learning recommendation model and the prefetch process of the storage table.
[0054] Figure 2 A schematic diagram showing a fine-grained lock according to an exemplary embodiment of the present disclosure. As Figure 2 shown, the training phase T of the training process may correspond to the prefetch phase T + 1 of the prefetch process, the training phase T + 1 of the training process may correspond to the prefetch phase T + 2 of the prefetch process, and the training phase T + 2 of the training process may correspond to the prefetch phase T + 3 of the prefetch process. For example, when the first processor (e.g., a graphics processing unit) executes the training phase T of the training process, the second processor (e.g., a central processing unit) may be executing the prefetch phase T + 1 of the prefetch process to prefetch the storage table to be used in (or corresponding to) the training phase T + 1 of the training process. When the first processor (e.g., a graphics processing unit) executes the training phase T + 1 of the training process, the second processor (e.g., a central processing unit) may be executing the prefetch phase T + 2 of the prefetch process to prefetch the storage table to be used in (or corresponding to) the training phase T + 2 of the training process. When the first processor (e.g., a graphics processing unit) executes the training phase T + 2 of the training process, the second processor (e.g., a central processing unit) may be executing the prefetch phase T + 3 of the prefetch process to prefetch the storage table to be used in (or corresponding to) the training phase T + 3 of the training process. For example, when T = 1, the training phase 1 of the training process may correspond to the prefetch phase 2 of the prefetch process, and when T = 0, the training phase 0 of the training process may correspond to the prefetch phase 1 of the prefetch process. In some embodiments, the first prefetch phase of the prefetch process may be completed at or before the start of the training process (e.g., before the training phase 1 of the training process).
[0055] In an exemplary embodiment of the present disclosure, prefetching a storage table to be used in a next training phase (or corresponding to the next training phase) after the current training phase from the memory device may include: parallelly searching for a plurality of storage tables to be used in the next training phase (or corresponding to the next training phase) in the memory device; and prefetching the storage table from the memory device based on the plurality of storage tables. Here, by performing the parallel search operation, the time for performing the search operation can be saved, so that the training time of the deep learning recommendation model overlaps with the transmission time of the storage table.
[0056] Searching or a search operation refers to an operation on a storage table, which means searching for a storage table according to input data. Usually, it is processed item by item. Due to the independence of each storage table, that is, each storage table can be processed by itself and is not affected by other storage tables, these storage tables can be processed in parallel according to the embodiment. Considering these characteristics, the present disclosure may be related to parallel search, which means that multiple second processor cores (for example, central processing unit cores) can be used to process the search operation of the storage table.
[0057] Figure 3 A schematic diagram showing parallel search according to an exemplary embodiment of the present disclosure. As Figure 3 shown, multiple second processor cores 301 (for example, central processing unit cores, that is, Figure 3 the CPU cores in) are used to respectively and simultaneously search for the storage table corresponding to one input data of a plurality of input data. For example, when using multiple second processor cores (for example, central processing unit cores) to respectively and simultaneously search for the storage tables corresponding to 3 input data, the following operations are performed simultaneously: the first second processor core 301A (for example, central processing unit core) can be used to search for the storage table corresponding to the first input data 302A, the second second processor core 301B (for example, central processing unit core) can be used to search for the storage table corresponding to the second input data 302B, and the third second processor core 301C (for example, central processing unit core) can be used to search for the storage table corresponding to the third input data 302C.
[0058] In an exemplary embodiment of the present disclosure, parallel search may include: determining corresponding data in the training data to be used in the next training phase (or corresponding to the next training phase); and based on at least a part of the corresponding data, parallelly searching for each storage table in the plurality of storage tables in the memory device. Here, parallel search can be performed based on the training data to be used, so as to implement the search of the storage table.
[0059] When the training time is too short (e.g., less than the average training time) during a training phase, in the present disclosure, a part of the training data to be used in this training phase can be randomly sampled, and subsequent operations of looking up the storage table can be performed to adapt to the training time of the deep learning recommendation model. In some embodiments, this part can be, for example, 10%, but the embodiments are not limited thereto. When the training time is not short (e.g., greater than or equal to the average training time) during a training phase, in the present disclosure, the training data to be used in this training phase can be used to perform subsequent operations of looking up the storage table.
[0060] In an exemplary embodiment of the present disclosure, prefetching the storage table from the memory device based on the multiple storage tables may include: determining a first storage table with a popularity less than a popularity threshold among the multiple storage tables; determining a second storage table with an approximation degree greater than or equal to a first approximation threshold to the first storage table; and prefetching the second storage table to replace the first storage table based on the popularity of the second storage table being greater than or equal to the popularity threshold. Here, a hotter storage table can be used to replace a colder storage table, so that it is no longer necessary to transfer the colder storage table.
[0061] In an exemplary embodiment of the present disclosure, popularity represents the access frequency of the storage table, and the popularity threshold represents the threshold of the access frequency of the storage table. For example, popularity can represent the frequency of the storage table being prefetched or / and used. The higher the access frequency of a storage table, the higher the popularity; the lower the access frequency of a storage table, the lower the popularity. As an example, when the popularity threshold is 5, the storage tables with a popularity less than 5 among the multiple storage tables (e.g., storage table 1 with a popularity of 3, storage table 4 with a popularity of 1, storage table 7 with a popularity of 0, etc.) can be determined as the first storage table. A hotter storage table can refer to a storage table with a higher access frequency, and a colder storage table can refer to a storage table with a lower access frequency.
[0062] In an exemplary embodiment of the present disclosure, the approximation degree refers to the degree of approximation or similarity between the storage table and the first storage table, and the first approximation threshold is a threshold for determining whether an approximation degree meets the condition of being determined as the second storage table. As an example, when the first approximation threshold is 85%, the storage table with an approximation degree greater than 85% to the first storage table (e.g., storage table 3 with an approximation degree of 90% to the first storage table) is determined as the second storage table.
[0063] In the present disclosure, all storage tables can be divided into n (for example, but not limited to, 3) layers according to the access frequency of the storage tables, which can be basic variables for approximate replacement. The Euclidean distance between each storage table in the m-th layer and the (m + 1)-th layer (m = [0, n)) can be calculated, and the storage table closest to the current storage table can be its approximate storage table (for example, approximate storage table). In the present disclosure, the following settings can be made: the Euclidean distance between the current storage table and the approximate storage table can be less than a given threshold; the layer number of the approximate storage table can be less than the layer number of the current storage table; not every storage table can have an approximate storage table. In the present disclosure, any method can be used to calculate the Euclidean distance between two storage tables, and the present disclosure does not limit this.
[0064] In an exemplary embodiment of the present disclosure, prefetching a second storage table to replace a first storage table may include: prefetching the second storage table based on the current storage table replacement rate being less than the replacement rate threshold. In an exemplary embodiment of the present disclosure, prefetching a storage table from the memory device based on the multiple storage tables may further include: determining not to perform a storage table replacement operation and prefetching the first storage table based on the current storage table replacement rate being greater than or equal to the replacement rate threshold. Here, in order not to let approximate replacement overly affect the performance of the deep learning recommendation model, n% of the storage tables are kept from being replaced at a time, so as to retain the generalization ability of the deep learning recommendation model as much as possible.
[0065] In an exemplary embodiment of the present disclosure, prefetching a storage table from the memory device based on the multiple storage tables may include: determining that an approximate storage table is included in the storage tables that have been transmitted to a first processor (for example, a graphics processor), where the approximate storage table has an approximation degree greater than or equal to a second approximation threshold with a third storage table in the multiple storage tables, determining not to prefetch the third storage table, and determining to use the approximate storage table in the first processor to replace the third storage table. Here, when a storage table similar to the current storage table is already in the first processor (for example, a graphics processor) (for example, has been loaded into the first processor or has been stored in the first processor), it can be used to replace the current storage table without transmitting the current storage table to the first processor (for example, a graphics processor) again, so that the number of storage tables transmitted to the first processor (for example, a graphics processor) is significantly reduced, thereby making the training time of the deep learning recommendation model overlap with the transmission time of the storage tables.
[0066] In an exemplary embodiment of the present disclosure, determining not to pre-fetch the third storage table may include: determining not to pre-fetch the third storage table based on that the replacement rate of the current storage table is less than a replacement rate threshold. In an exemplary embodiment of the present disclosure, pre-fetching a storage table from the memory device based on the multiple storage tables also includes: determining not to perform a storage table replacement operation and pre-fetching the third storage table based on that the replacement rate of the current storage table is greater than or equal to the replacement rate threshold. Here, while keeping n% of the storage tables from being replaced at a time, when a storage table similar to the current storage table has been loaded or stored in the first processor (e.g., a graphics processor), it can be used to replace the current storage table, thereby reducing the number of storage tables transmitted to the first processor (e.g., a graphics processor) while retaining the generalization ability of the deep learning recommendation model as much as possible.
[0067] In step S112, in response to pre-fetching the storage table, the pre-fetched storage table is transmitted to the first processor.
[0068] The above has been combined Figures 1A to 3 A method for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure is described. Figure 4A and Figure 5 An apparatus for training a deep learning recommendation model and units thereof, and a system for training a deep learning recommendation model according to exemplary embodiments of the present disclosure are described.
[0069] Figure 4A A block diagram of an apparatus 40 for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure is shown.
[0070] Reference Figure 4A The device 40 for training a deep learning recommendation model includes a storage table unloading unit 401, a model training unit 402, a storage table loading unit 403 and a storage table updating unit 404.
[0071] The storage table unloading unit 401 may be configured to unload the storage table to a memory device. Here, after the deep learning recommendation model initializes the storage table, the storage table may be unloaded from the first processor to the memory device. Here, before training the deep learning recommendation model, the training data may be loaded into the deep learning recommendation model located in the first processor (e.g., a graphics processor). Loading the training data into the deep learning recommendation model located in the first processor (e.g., a graphics processor) may mean that the deep learning recommendation model will be trained based on the training data. Based on this, the storage table unloading unit 401 may unload the storage table to the memory device.
[0072] In an exemplary embodiment of the present disclosure, the memory device may include a dynamic random access memory based on a computing fast link memory module, and the storage table may include an embedded table.
[0073] The model training unit 402 can be configured to train the deep learning recommendation model based on training data. Here, the training data is used for training the deep learning recommendation model.
[0074] The storage table loading unit 403 can be configured to load the storage table prefetched from the memory device according to the training stage during the training of the deep learning recommendation model.
[0075] In an exemplary embodiment of the present disclosure, the storage table loading unit 403 can be configured to: load the prefetched storage table based on determining that the second processor prefetches the storage table from the memory device according to the training stage.
[0076] The storage table updating unit 44 can be configured to store the feature data during training into the loaded storage table. Here, the feature data during the training process is stored into the corresponding storage table respectively.
[0077] Figure 4B The block diagram of the apparatus for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure is shown.
[0078] Refer to Figure 4B , the apparatus 41 for training a deep learning recommendation model includes a storage table prefetching unit 411 and a storage table transferring unit 412.
[0079] The storage table prefetching unit 411 is configured to prefetch a storage table from the memory device according to the training stage of the training of the deep learning recommendation model for storing the feature data during training.
[0080] In an exemplary embodiment of the present disclosure, the storage table prefetching unit 411 can be configured to: prefetch the storage table to be used in the next training stage of the current training stage from the memory device.
[0081] In an exemplary embodiment of the present disclosure, the storage table prefetching unit 411 can be configured to: search for multiple storage tables to be used in the next training stage in parallel in the memory device; and prefetch the storage table from the memory device based on the multiple storage tables.
[0082] In an exemplary embodiment of the present disclosure, the storage table prefetching unit 411 can be configured to: select or determine the corresponding data in the training data to be used in the next training stage; and search for each storage table in the multiple storage tables in parallel in the memory device based on at least a part of the corresponding data.
[0083] In an exemplary embodiment of the present disclosure, the storage table prefetch unit 411 may be configured to: select or determine a first storage table among the multiple storage tables with a popularity lower than a popularity threshold; select or determine a second storage table with a similarity greater than or equal to a first similarity threshold to the first storage table; and based on the popularity of the second storage table being greater than or equal to the popularity threshold, prefetch the second storage table to replace the first storage table. In an exemplary embodiment of the present disclosure, popularity represents the access frequency of the storage table, and the popularity threshold represents the threshold of the access frequency of the storage table. In an exemplary embodiment of the present disclosure, similarity refers to the degree of approximation or similarity between the storage table and the first storage table, and the first similarity threshold is a threshold for determining whether a similarity meets the condition of being determined as the second storage table.
[0084] In an exemplary embodiment of the present disclosure, the storage table prefetch unit 411 may be configured to: based on the current storage table replacement rate being lower than a replacement rate threshold, prefetch the second storage table; and based on the current storage table replacement rate being greater than or equal to the replacement rate threshold, determine not to perform a storage table replacement operation and prefetch the first storage table.
[0085] In an exemplary embodiment of the present disclosure, the storage table prefetch unit 411 may be configured to: based on determining that an approximate storage table is included in the storage tables that have been transmitted to the first processor, where the approximate storage table has a similarity greater than or equal to a second similarity threshold to a third storage table among the multiple storage tables, determine not to prefetch the third storage table and determine to use the approximate storage table in the first processor to replace the third storage table.
[0086] In an exemplary embodiment of the present disclosure, the storage table prefetch unit 411 may be configured to: based on the current storage table replacement rate being lower than the replacement rate threshold, determine not to prefetch the third storage table; based on the current storage table replacement rate being greater than or equal to the replacement rate threshold, determine not to perform a storage table replacement operation and prefetch the third storage table.
[0087] The storage table transfer unit 412 is configured to transfer the prefetched storage table to the first processor based on the prefetched storage table.
[0088] Figure 5 A schematic diagram showing a system for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure.
[0089] As Figure 5 shown, the system for training a deep learning recommendation model includes a first processor 51 (e.g., a graphics processor), a second processor 52 (e.g., a central processing unit), and a memory device 53 (e.g., a dynamic random access memory based on a compute express link memory module).
[0090] The first processor 51 may be configured to unload the storage table to the memory device 53, train the deep learning recommendation model based on the training data, and during the training of the deep learning recommendation model, load the storage table prefetched from the memory device 53 according to the training phase.
[0091] The second processor 52 may be configured to prefetch the storage table from the memory device 53 based on the training phase associated with the training of the deep learning recommendation model, where the storage table may be used to store the feature data during training, and transfer the prefetched storage table to the first processor based on the prefetched storage table.
[0092] The first processor 51 may also be configured to store the feature data during training into the loaded storage table.
[0093] In an exemplary embodiment of the present disclosure, the first processor may be a graphics processor, and the second processor may be a central processing unit.
[0094] In an exemplary embodiment of the present disclosure, the memory device 53 may include a dynamic random access memory based on a compute express link memory module, and wherein the storage table may include an embedding table.
[0095] In addition, the second processor 52 may include a dynamic prefetch module 521, a parallel lookup module 522, and an approximate replacement module 523.
[0096] In an exemplary embodiment of the present disclosure, the dynamic prefetch module 521 may be configured to prefetch the storage table from the memory device 53 according to the training phase, thereby reducing the training time.
[0097] In an exemplary embodiment of the present disclosure, the dynamic prefetch module 521 may be configured to prefetch the storage table to be used in the next training phase of the current training phase from the memory device 53. Here, by prefetching the storage table to be used in the next training phase during the current training phase, almost no waiting occurs for the main process and the prefetch process, so that the training time of the deep learning recommendation model overlaps with the transmission time of the storage table, without slowing down the training progress of the deep learning recommendation model in the graphics processor.
[0098] In an exemplary embodiment of the present disclosure, the parallel lookup module 522 may be configured to parallelly look up multiple storage tables to be used in the next training phase in the memory device 53; and prefetch the storage table from the memory device 53 based on the multiple storage tables. Here, by performing a parallel lookup operation, the time for performing the lookup operation can be saved, so that the training time of the deep learning recommendation model overlaps with the transmission time of the storage table.
[0099] In an exemplary embodiment of the present disclosure, the parallel search module 522 may be configured to select or determine corresponding data in the training data that will be used in the next training phase; and based on at least a portion of the corresponding data, parallel search each of the plurality of storage tables in the memory device 53.
[0100] In an exemplary embodiment of the present disclosure, the approximate replacement module 523 may be configured to select or determine a first storage table among the plurality of storage tables with a popularity less than a popularity threshold; select or determine a second storage table with an approximation degree greater than or equal to a first approximation threshold with respect to the first storage table; and based on the popularity of the second storage table being greater than or equal to the popularity threshold, prefetch the second storage table to replace the first storage table. Here, a hotter storage table can be used to replace a colder storage table, so that it is no longer necessary to transfer the colder storage table. In an exemplary embodiment of the present disclosure, popularity represents the access frequency of a storage table, and the popularity threshold represents a threshold of the access frequency of a storage table. In an exemplary embodiment of the present disclosure, the approximation degree refers to the degree of approximation or similarity between a storage table and the first storage table, and the first approximation threshold is a threshold for determining whether an approximation degree meets the condition of being determined as the second storage table.
[0101] In an exemplary embodiment of the present disclosure, the approximate replacement module 523 may be configured to prefetch the second storage table based on the current storage table replacement rate being less than a replacement rate threshold; and based on the current storage table replacement rate being greater than or equal to the replacement rate threshold, determine not to perform a storage table replacement operation and prefetch the first storage table. Here, in order not to let the approximate replacement overly affect the performance of the deep learning recommendation model, n% of the storage tables can be kept from being replaced at a time, so as to retain the generalization ability of the deep learning recommendation model as much as possible.
[0102] In an exemplary embodiment of the present disclosure, the approximate replacement module 523 may be configured to determine not to prefetch a third storage table based on determining that an approximate storage table is included in the storage tables that have been transmitted to the first processor, where the approximate storage table has an approximation degree greater than or equal to a second approximation threshold with respect to a third storage table among the plurality of storage tables, and determine to use the approximate storage table in the first processor to replace the third storage table. Here, when a storage table similar to the current storage table has been loaded or stored in the first processor, it can be used to replace the current storage table without transmitting the current storage table to the first processor again, so that the number of storage tables transmitted to the first processor is significantly reduced, and thus the training time of the deep learning recommendation model overlaps with the transmission time of the storage tables.
[0103] In an exemplary embodiment of the present disclosure, the approximate replacement module 523 may be configured to determine not to pre-fetch the third storage table based on the replacement rate of the current storage table being less than the replacement rate threshold; and to determine not to perform a storage table replacement operation and pre-fetch the third storage table based on the replacement rate of the current storage table being greater than or equal to the replacement rate threshold. Here, in the case where n% of the storage tables are kept unreplaced at a time, when a storage table similar to the current storage table has been loaded or stored in the graphics processor, it can be used to replace the current storage table, thereby reducing the number of storage tables transmitted to the graphics processor while retaining the generalization ability of the deep learning recommendation model as much as possible.
[0104] In addition, according to an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed, the method of training a deep learning recommendation model according to an exemplary embodiment of the present disclosure is implemented.
[0105] In an exemplary embodiment of the present disclosure, the computer-readable storage medium may carry one or more programs, and when the computer program is executed, the following steps may be implemented: unloading the storage table to the memory device; training the deep learning recommendation model based on the training data, wherein, in the process of training the deep learning recommendation model, the storage table pre-fetched from the memory device is loaded according to the training stage; and the feature data in the training is stored in the loaded storage table, so that the training time and the transmission time overlap by pre-fetching the storage table. Therefore, the training time can be reduced, and the training efficiency can be improved.
[0106] In an exemplary embodiment of the present disclosure, the computer-readable storage medium may carry or store one or more programs, and when the computer program is executed, the following steps may be implemented: pre-fetching a storage table from a memory device according to a training phase of training the deep learning recommendation model by the first processor to store feature data in training; in response to pre-fetching the storage table, transmitting the pre-fetched storage table to the first processor, thereby realizing pre-fetching of the storage table and reducing training time.
[0107] A computer-readable storage medium may, for example, but is not limited to, be a system, apparatus, or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In an embodiment of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a computer program, and the computer program may be used by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable storage medium may be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above. The computer-readable storage medium may be contained in any device; it may also exist separately without being assembled into the device.
[0108] In addition, according to an exemplary embodiment of the present disclosure, a computer program product is further provided, and the instructions in the computer program product can be executed by a processor of a computer device to complete the method for training a deep learning recommendation model according to the exemplary embodiment of the present disclosure.
[0109] The above has been combined with FIGS. 4 to Figure 5 The apparatus and system for training a deep learning recommendation model according to the exemplary embodiment of the present disclosure have been described. Next, in combination with Figure 6 The computing device according to the exemplary embodiment of the present disclosure will be described.
[0110] Figure 6 A schematic diagram of a computing device according to an exemplary embodiment of the present disclosure is shown.
[0111] Referring to Figure 6 , a computing device 6 according to an exemplary embodiment of the present disclosure includes a memory 61 and a processor 62. A computer program is stored on the memory 61, and when the computer program is executed by the processor 62, the method for training a deep learning recommendation model according to the exemplary embodiment of the present disclosure is implemented.
[0112] In an exemplary embodiment of the present disclosure, when the computer program is executed by the processor 62, the following steps can be implemented: unloading a storage table to a memory device; training the deep learning recommendation model based on training data, wherein during the training of the deep learning recommendation model, a storage table prefetched from the memory device is loaded according to the training stage; and storing the feature data during training into the loaded storage table, so that the training time overlaps with the transmission time through the prefetching of the storage table, reducing the training time and thus improving the training efficiency.
[0113] In an exemplary embodiment of the present disclosure, when the computer program is executed by the processor 62, the following steps can be implemented: prefetching a storage table from a memory device according to the training stage of the training of the deep learning recommendation model by a first processor for storing the feature data during training; and in response to the prefetched storage table, transmitting the prefetched storage table to the first processor, thereby realizing the prefetching of the storage table and reducing the training time.
[0114] The computing device in the embodiments of the present disclosure may include, but is not limited to, devices such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, etc. Figure 6 The illustrated computing device is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0115] The above has been described with reference to FIGS. 1 to Figure 6 a method and apparatus for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure. However, it should be understood that: FIGS. 4 to Figure 5 the apparatus and its units for training the deep learning recommendation model shown therein may be respectively configured as software, hardware, firmware, or any combination of the above items for performing specific functions, Figure 6 the computing device shown therein is not limited to including the components shown above, but some components may be added or deleted as needed, and the above components may also be combined.
[0116] A method, apparatus, and system for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure may include unloading a storage table to a memory device, training the deep learning recommendation model based on training data, wherein during the training of the deep learning recommendation model, a storage table prefetched from the memory device is loaded according to the training stage, and storing the feature data during training into the loaded storage table, so that the training time overlaps with the transmission time through the prefetching of the storage table, which can reduce the training time and thus can improve the training efficiency.
[0117] A method, apparatus, and system for training a deep learning recommendation model according to an exemplary embodiment of the present disclosure may include prefetching a storage table from a memory device during a training phase of training the deep learning recommendation model by a first processor for storing feature data during training, and transmitting the prefetched storage table to the first processor in response to the prefetching of the storage table, thereby achieving prefetching of the storage table and reducing the training time.
[0118] Although some embodiments of the present disclosure have been specifically shown and described with reference to their exemplary embodiments, those skilled in the art should understand that various changes in form and details may be made thereto without departing from the spirit and scope of the present disclosure defined by the claims.
Claims
1. A method for training a Deep Learning Recommendation Model (DLRM) executed by a first processor, comprising: Unloading one or more storage tables to a memory device; Training the deep learning recommendation model based on training data, wherein, during the training of the deep learning recommendation model, storage tables prefetched from the memory device are loaded according to the training phase; And Storing the feature data obtained during training into the loaded storage tables.
2. The method according to claim 1, wherein, The memory device includes a Compute Express Link Memory Module - Dynamic Random Access Memory (CMM - D), and wherein the storage tables include embedding tables.
3. The method according to claim 1, wherein, The prefetched storage tables are loaded based on determining that a second processor prefetches storage tables from the memory device according to the training phase.
4. A method for training a Deep Learning Recommendation Model (DLRM) executed by a second processor, comprising: Prefetching storage tables from a memory device according to a training phase associated with the training of the deep learning recommendation model by a first processor, wherein the storage tables are configured to store feature data obtained during training; Based on the prefetched storage tables, transferring the prefetched storage tables to the first processor.
5. The method according to claim 4, wherein, The storage tables correspond to the next training phase after the current training phase.
6. The method according to claim 5, wherein, The prefetched storage tables include: Parallelly searching in the memory device for multiple storage tables corresponding to the next training phase; and Based on the multiple storage tables, prefetching storage tables from the memory device.
7. The method according to claim 6, wherein, The parallel search includes: Selecting corresponding data corresponding to the next training phase in the training data for training the deep learning recommendation model; Based on at least a part of the corresponding data, parallelly searching each of the multiple storage tables in the memory device.
8. The method according to claim 6, wherein The prefetched storage tables include: Selecting a first storage table from the multiple storage tables, wherein the popularity associated with the first storage table is less than a popularity threshold; Selecting a second storage table whose similarity to the first storage table is greater than or equal to a first similarity threshold; Based on the popularity of the second storage table being greater than or equal to the popularity threshold, prefetching the second storage table to replace the first storage table.
9. The method according to claim 8, wherein, Based on the current storage table replacement rate being less than a replacement rate threshold, the second storage table is prefetched, wherein, the prefetched storage tables further include: Based on the current storage table replacement rate being greater than or equal to the replacement rate threshold, determining not to perform a storage table replacement operation and prefetching the first storage table.
10. The method according to claim 6, wherein, The prefetched storage tables include: Based on determining that an approximate storage table has been transferred to the first processor, wherein the approximate storage table has a similarity greater than or equal to a second similarity threshold with a third storage table among the multiple storage tables, determining not to prefetch the third storage table and determining to use the approximate storage table to replace the third storage table.
11. The method according to claim 10, wherein The determination of not prefetching the third storage table includes: Based on the current storage table replacement rate being less than the replacement rate threshold, determining not to prefetch the third storage table, wherein, the prefetched storage tables further include: Based on the current storage table replacement rate being greater than or equal to the replacement rate threshold, it is determined not to perform a storage table replacement operation, and a third storage table is prefetched.
12. An apparatus for training a deep learning recommendation model DLRM, comprising: A storage table unloading unit configured to unload one or more storage tables to a memory device; A model training unit configured to train the deep learning recommendation model based on training data; A storage table loading unit configured to load, during the training of the deep learning recommendation model, the storage tables prefetched from the memory device according to the training phase; And A storage table updating unit configured to store the feature data obtained during training into the loaded storage tables.
13. An apparatus for training a deep learning recommendation model DLRM, comprising: A storage table prefetching unit configured to prefetch storage tables from a memory device according to a training phase associated with the training of the deep learning recommendation model, wherein the storage tables are configured to store the feature data obtained during training; A storage table transferring unit configured to transfer the prefetched storage tables to a first processor based on the prefetched storage tables.
14. A system for training a deep learning recommendation model DLRM, comprising: A memory device; A first processor configured to: unload one or more storage tables to the memory device, train the deep learning recommendation model based on training data, and during the training of the deep learning recommendation model, load the storage tables prefetched from the memory device according to the training phase, A second processor configured to: prefetch storage tables from the memory device according to the training phase, and transfer the prefetched storage tables to the first processor based on the prefetched storage tables, wherein the first processor is further configured to store the feature data obtained during training into the loaded storage tables.
15. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by a processor, the method for training a deep learning recommendation model according to any one of claims 1 to 11 is implemented.
16. A computing device, comprising: At least one processor; At least one memory storing a computer program, which when executed by the at least one processor, implements the method for training a deep learning recommendation model according to any one of claims 1 to 11.