Method, device, system, equipment, medium, and computer program for training deep learning models

The method addresses the high hardware requirements and time costs of deep learning model training by determining target parameters and adjusting network parameters based on available storage slots, enhancing training efficiency and scalability.

JP7682387B2Active Publication Date: 2025-05-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024519091
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-05-19
Filing Date
2022-09-27
Publication Date
2025-05-23
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

Existing deep learning model training methods require high hardware resources and incur significant time costs due to increasing data and model scales, particularly in large-scale sparse parameter training.

Method used

A method for training deep learning models that determines target parameters based on training data and adjusts network parameters by utilizing remaining storage slots in a target memory, optimizing storage and reducing hardware requirements.

Benefits of technology

This approach reduces hardware requirements and improves training efficiency by effectively managing storage slots and optimizing the use of memory resources, enabling large-scale model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007682387000001
    Figure 0007682387000001
  • Figure 0007682387000002
    Figure 0007682387000002
  • Figure 0007682387000003
    Figure 0007682387000003
Patent Text Reader

Abstract

The present disclosure provides a method for training a deep learning model, which relates to the field of artificial intelligence, particularly to the field of deep learning and intelligent recommendation. The specific implementation solution of the method for training a deep learning model includes: determining a first target parameter that needs to be written into a target memory, which is a memory included in a target processor, in a first network parameter that is required for performing embedding processing on the first training data according to a first training data of a current training round; determining remaining storage slots in the target memory according to a first mapping relationship between the storage slots of the target memory and the network parameter; and writing the first target parameter into the target memory in response to the remaining storage slots meeting the storage requirements of the first target parameter, so that a computing core included in the target processor adjusts the first network parameter according to the first training data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application claims priority to Chinese patent application filed on May 19, 2022 and bearing application number 202210559489.0, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to the field of artificial intelligence, specifically to the field of deep learning and intelligent recommendation, and in particular to a method, an apparatus, a system, and an electronic device for training a deep learning model , written storage medium and computer programs Regarding. [Background technology]

[0003] With the development of computer technology, network technology and communication technology, the application of technologies such as deep learning in fields such as intelligent recommendation is becoming more and more widespread. With the push of big data and the development of deep learning technology, the data scale and model scale of deep learning technology are both increasing significantly. Accordingly, model training places high requirements on the hardware environment, and the time cost of normal training is also very high. Summary of the Invention [Problem to be solved by the invention]

[0004] The present invention relates to a deep learning model training method, device, system, and electronic device that are useful for reducing hardware requirements and achieving large-scale model training. , written storage medium and computer programs to provide. [Means for solving the problem]

[0005] According to one aspect of the present disclosure, there is provided a method for training a deep learning model, comprising: determining, based on first training data of a current training round, first target parameters that need to be written to a target memory, which is a memory included in a target processor, in a first network parameter required for performing an embedding process on the first training data; determining remaining storage slots in the target memory based on a first mapping relationship between the storage slots of the target memory and the network parameter; and writing the first target parameters to the target memory in response to the remaining storage slots satisfying the storage requirements of the first target parameters, such that a computing core included in the target processor adjusts the first network parameter based on the first training data.

[0006] According to another aspect of the present disclosure, there is provided a method for training a deep learning model, the method including: a first processor determining, based on first training data of a current training round, a first target parameter that needs to be written to a target memory, which is a memory included in a second processor, in a first network parameter required for performing an embedding process on the first training data; the first processor determining remaining storage slots in the target memory based on a first mapping relationship between the storage slots of the target memory and the network parameter; in response to the remaining storage slots satisfying the storage requirements of the first target parameter, the first processor writing the first target parameter to the target memory and sending training task information based on the first training data to the second processor; and in response to a computing core of the second processor receiving the training task information, adjusting the first network parameter based on the first training data.

[0007] According to another aspect of the present disclosure, there is provided a deep learning model training apparatus, including: a target parameter determination module that determines, based on first training data of a current training round, first target parameters that need to be written to a target memory, which is a memory included in a target processor, in a first network parameter required for performing embedding processing on the first training data; a remaining slot determination module that determines remaining storage slots in the target memory based on a first mapping relationship between the storage slots of the target memory and the network parameters; and a parameter writing module that writes the first target parameters to the target memory in response to the remaining storage slots satisfying the storage requirements of the first target parameters, such that a computing core included in the target processor adjusts the first network parameter based on the first training data.

[0008] According to another aspect of the present disclosure, there is provided a system for training a deep learning model, comprising: a first processor and a second processor, the second processor comprising a target memory and a computation core, the first processor is configured to determine, based on first training data of a current training round, first target parameters to be written to the target memory in a first network parameter required for performing embedding processing on the first training data, determine remaining storage slots in the target memory based on a first mapping relationship between the storage slots of the target memory and the network parameters, and in response to the remaining storage slots satisfying the storage requirements of the first target parameters, write the first target parameters to the target memory and transmit training task information based on the first training data to the second processor, and the second processor is configured to adjust the first network parameters based on the first training data in response to the computation core receiving the training task information.

[0009] According to another aspect of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executable by the at least one processor such that the at least one processor can perform a method for training a deep learning model provided by the present disclosure.

[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium having stored thereon computer instructions is provided, where the computer instructions cause a computer to perform a method for training a deep learning model provided by the present disclosure.

[0011] According to another aspect of the present disclosure , P and a computer program that, when executed by a processor, implements the method for training a deep learning model provided by the present disclosure. M We provide it.

[0012] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, and are not intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following description.

[0013] The drawings are used for a better understanding of the present solution and are not intended to limit the present disclosure. [Brief description of the drawings]

[0014] [Figure 1] FIG. 1 is an application scenario architecture diagram of a deep learning model training method, apparatus and system according to an embodiment of the present disclosure. [Diagram 2] FIG. 2 is a schematic flow chart of a method for training a deep learning model according to an embodiment of the present disclosure. [Diagram 3] FIG. 3 is a schematic flow chart of a method for training a deep learning model according to another embodiment of the present disclosure. [Figure 4] FIG. 4 is a structural schematic diagram of a processor cache according to an embodiment of the present disclosure. [Diagram 5] FIG. 5 is an overall flowchart of a method for training a deep learning model according to an embodiment of the present disclosure. [Figure 6] FIG. 6 is a communication topology diagram of a stand-alone multi-card processor according to an embodiment of the present disclosure. [Figure 7] FIG. 7 is a principle schematic diagram of training a model in an asynchronous pipeline format according to an embodiment of the present disclosure. [Figure 8] FIG. 8 is a structural block diagram of a deep learning model training apparatus according to an embodiment of the present disclosure. [Figure 9] FIG. 9 is a structural block diagram of a deep learning model training system according to an embodiment of the present disclosure. [Figure 10] FIG. 10 is a block diagram of an electronic device for implementing a method for training a deep learning model according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings, which are useful for understanding various details of the embodiments of the present disclosure and should be considered as examples. Therefore, as will be understood by those skilled in the art, various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and simplicity, the following description will omit descriptions of known functions and structures.

[0016] With the development of big data push and deep learning technology, the data scale and model scale in the industrial recommendation scene are both significantly increasing. For example, in order to improve the accuracy of a recommendation model, it is generally necessary to generate training samples based on billions of click data to train the recommendation model. In the recommendation model, embedding technology is generally used to convert the high-dimensional sparse feature vector of an object (such as a user or an item) into a low-dimensional dense feature vector. In this way, the parameters related to the embedding technology always reach the level of hundreds of billions or even tens of billions, and the related parameters have sparse characteristics.

[0017] To achieve training on large-scale sparse parameters, a parameter server architecture based on a CPU or GPU is used to perform distributed training on large-scale sparse parameters, thereby improving training efficiency.

[0018] Parameter server architectures may include, for example, HugeCTR, Paddle-GPUPS, Persia, etc.

[0019] For example, HugeCTR is a framework for accelerating recommendation model training using GPU, which supports multi-machine multi-card acceleration, and supports a mixed training method of model parallel training for embedding layer with parameter sparse distribution and data parallel training for network with parameter dense distribution. HugeCTR divides embedding layer into multiple parts and assigns them to multi-machine multi-cards respectively, and stores a part of global embedding layer in each GPU, and at the same time, each GPU has a complete network with parameter dense distribution. When training recommendation model, global sample data can be randomly shuffled and divided, and different sample data is assigned to each GPU to train in data parallel mode.

[0020] HugeCTR supports two methods of embedding layer storage: one is to cache sparse parameters belonging to the same slot in the video card memory of the same GPU; the other is to distribute the entire amount of sparse parameters and then store them in the video card memory of different GPUs. Both of these methods have the problem that some sparse parameters are repeatedly cached, which will result in a certain amount of waste in the video card memory. In addition, HugeCTR requires multiple CPUs to participate in model training, which results in high training costs.

[0021] For example, the emergence of Paddle-GPUPS solves the problem of high training costs for over a hundred CPU servers. The architecture builds a high bandwidth memory (HBM) hash table in each GPU. Before starting training, the architecture first loads the sparse parameters required for embedding the features of the data in the currently acquired pass from the CPU memory to the video card memory. When loading, the sparse parameters required for the same feature group are distributed and then stored in different video card memories. In this way, when training a model based on one batch of data extracted from one pass, each GPU needs to copy the required sparse parameters from other video card memories based on the feature identifier. In the training process, the architecture has a large communication overhead between GPUs, and an HBM hash table is built and stored in each GPU, so the requirements for the size of the video card memory are high.

[0022] For example, Persia is a recommendation model training framework for large-scale heterogeneous cluster training. The framework optimizes the training algorithm and the training system in two dimensions, so that the maximum number of trainable model parameters is at the level of one million billion. The framework performs asynchronous updates for the embedding layer and synchronous updates for the network with dense parameter distribution, and some communication and calculation processes can be overlapped in time through system optimization. The framework introduces the role of an embedding worker into the traditional framework, and the training update task of the embedding layer is separated from the training task of the whole model and is executed by the embedding worker. The framework requires a lot of CPU because of the introduction of the embedding worker, which increases the cost of training the model.

[0023] In addition, artificial intelligence (AI) chips, such as deep learning processors (DPUs), neural network processors (NPUs) and tensor processing units (TPUs), are being produced to accelerate neural network computing capabilities in order to improve model training efficiency.

[0024] For example, Konron's second generation chip is a general-purpose AI chip that uses GDDR6 video memory. The chip operates according to the XPU-R architecture, which can significantly improve the core computing power of the chip and improve the general computing ability of the chip.

[0025] The application scenario of the method and device provided by the present disclosure will be described below with reference to FIG. FIG. 1 is an application scene diagram of a deep learning model training method, apparatus and system according to an embodiment of the present disclosure.

[0026] 1, the application scenario 100 includes an electronic device, which may be a notebook computer, a desktop computer, a server, etc. The electronic device is provided with a processor CPU 110, an artificial intelligence chip 120, a memory 130, and a hard disk memory 140.

[0027] Memory 130 is an internal memory, and is a space for direct addressing and storage by CPU 110. The memory can temporarily store operating data within the CPU and data exchanged with external memory such as a hard disk. When the computer is running, the CPU calls the data that needs to be calculated from the memory and performs the calculation, and after the calculation is completed, the CPU sends the result. Memory 130 may be, for example, a random memory, and thus the CPU may read data from it and write data to it.

[0028] The hard disk processor 140 may be, for example, a solid state disk (SSD) with an NVMe interface, and the present disclosure is not limited thereto.

[0029] The artificial intelligence chip has data processing capabilities, and can assist the operation of the CPU and improve the overall operation speed. Examples of the artificial intelligence chip include the above-mentioned DPU, NPU, TPU, etc. The artificial intelligence chip 120 can include a calculation core, a video memory, and its related circuits. The video memory is the display memory 150, i.e., the dedicated memory of the artificial intelligence chip, and its function is to store the rendering data that the calculation core has processed or is to extract, similar to the memory 130, and the display memory 150 stores information such as model parameters, training samples, etc. to be processed.

[0030] The computing core in the artificial intelligence chip 120 cannot directly read data in the memory 130, and can only read data from the display memory 150. The CPU can assign a computing task to the computing core, and in the process of the computing core performing the computing task, under the control of the CPU 110, data exchange can be performed between the memory 130 and the display memory 150, so that the computing core copies the data required when performing the computing task from the memory 130 to the display memory 150, or directly transfers the data in the memory 130 to the display memory 150.

[0031] When training a model built based on deep learning technology, the CPU 110 can, for example, assign a training task to the artificial intelligence chip 120, and transfer the model from the memory 130 to the display memory 150. In one embodiment, the model can be stored in the hard disk storage space provided by the hard disk memory 140. A tertiary buffer space composed of the display memory 150, the memory 130 and the hard disk memory 140 is established. In this way, when the model is stored in the hard disk memory 140, during the model training process, according to the training needs, the CPU 110 reads data from the hard disk memory 140 and caches it in the memory, and when the CPU 110 assigns a training task to the artificial intelligence chip 120, the model parameters related to the calculation core performing the current calculation task are transferred from the memory 130 to the display memory 150, and the data processed by the calculation core stored in the display memory 150 is transferred from the display memory 150 to the memory 130, thereby avoiding the shortage of storage space of the display memory.

[0032] In one embodiment, the electronic device can be equipped with, for example, multiple artificial intelligence chips, which can perform model training tasks in parallel based on different training samples, thereby improving the efficiency of model training.

[0033] It can be understood that the deep learning model training method provided by the present disclosure can be executed by an electronic device, specifically by calling corresponding program code by a CPU or an artificial intelligence chip. Accordingly, the deep learning model training device and deep learning model training system provided by the present disclosure can be installed in an electronic device.

[0034] The deep learning model training method provided by the present disclosure will be described in detail below with reference to FIGS.

[0035] FIG. 2 is a flowchart of a method for training a deep learning model according to an embodiment of the present disclosure.

[0036] 2, the deep learning model training method 200 of the embodiment may include operations S210 to S230. The method 200 may be executed by, for example, a CPU in the above-mentioned electronic device.

[0037] In operation S210, based on the first training data of the current training round, a first target parameter that needs to be written into the target memory in the first network parameter required for performing embedding processing on the first training data is determined.

[0038] In operation S220, the remaining storage slots in the target memory are determined based on a first mapping relationship between the storage slots of the target memory and the network parameters.

[0039] In operation S230, in response to the remaining storage slots satisfying the storage requirements of the first target parameters, the first target parameters are written to the target memory, whereby a computing core included in the target processor adjusts the first network parameters based on the first training data.

[0040] According to the embodiment of the present disclosure, the target memory may be, for example, a memory included in the target processor. The target processor may be, for example, the above-mentioned artificial intelligence chip, or may be a graphics processor GPU, etc. The target processor may receive a computation task assigned by the CPU, and execute the assigned computation task based on the data stored in the target memory. The computation task may include, for example, a model training task, thereby training a deep learning model. The deep learning model may include, for example, an image processing model, a voice processing model, or a text processing model, etc. In a specific scenario, the deep learning model may be a recommendation model, and the embodiment may train the recommendation model by a method such as gradient descent based on a large number of user's dialogue act information for the recommended object. After the model parameters of the recommendation model are aggregated, personalized recommendations can be made to the user.

[0041] In one embodiment, the deep learning model may include, for example, an embedding layer and a prediction network. The embedding layer is used to project input data from a high-dimensional sparse space to a low-dimensional dense feature space by performing an embedding process on the data input to the deep learning model. The first network parameters required to perform the embedding process on the first training data in the embodiment of the present disclosure are the network parameters in the embedding layer. The first network parameters may be determined, for example, by calling a core function.

[0042] In one embodiment, the determined first network parameters can be compared with the network parameters stored in the target memory, and the network parameters in the first network parameters that are not stored in the target memory can be determined and set as the first target parameters that need to be written to the target memory. Or, in the embodiment, the first network parameters can be further compared with the network parameters stored in the memory and / or the hard disk processor, and the network parameters in the first network parameters that are stored in the memory and / or the hard disk processor can be set as the first target parameters. As can be understood, when comparing the network parameters, the comparison can be made based on the feature identifier Feature sign (abbreviated as Feasign) of the data that is embedded based on the network parameters.

[0043] For example, a training sample may include feature data of multiple objects, each object may include multiple feature data, and each feature data may correspond to a feature identifier. For each feature data, certain network parameters need to be adopted to perform embedding processing. For example, the embodiment of the present disclosure may store network parameters of an embedding layer based on the correspondence between the network parameters and the feature data, and add the feature identifier of the feature data corresponding to the network parameters.

[0044] In one embodiment, the CPU can maintain in a cache or memory a mapping relationship table between the feature identifiers of the feature data and the network parameters stored in the target memory, the mapping relationship table taking the feature identifiers as Key and the identifier information of the network parameters having a mapping relationship as Value. The embodiment queries the mapping relationship table based on the feature identifiers of the feature data included in the first training data, determines the feature identifiers not present in the mapping relationship table, and sets the network parameters for performing embedding processing on the feature data identified by the not-present feature identifiers in the first network parameters as first target parameters.

[0045] As can be understood, the network parameters stored in the target memory can be stored according to slots, for example, and the network parameters stored in each slot are all the network parameters corresponding to one feature data. That is, the network parameters can be stored in groups, and all the network parameters corresponding to one feature data constitute one network parameter group. In this way, the target memory can be divided into multiple storage slots, and each storage slot stores one network parameter group.

[0046] After determining the first target parameter, the embodiment first determines whether the storage space in the target memory is sufficient, and can write the first target parameter to the target memory only if the storage space is sufficient.

[0047] For example, the embodiment can maintain a first mapping relationship between storage slots and network parameters in a CPU cache or memory. The first mapping relationship can be stored in the form of a mapping table, and since the network parameters correspond one-to-one with the feature data, the embodiment can employ the feature identifier of the feature data to represent the network parameters and number the storage slots in the target memory. Thus, the first mapping relationship can be represented as a mapping table between Feasign and FId, with the feature identifier of the feature data as Key and the serial number of the storage slot (set to FId) as Value. Thus, the embodiment can determine the remaining storage slots in the target memory based on the first mapping relationship.

[0048] For example, if the total storage space of the target memory is divided into 100 storage slots, and the serial numbers of the 100 storage slots are set to be integers from 0 to 99, and the first mapping relationship only includes mapping information with serial numbers from 0 to 49, it can be determined that there are 50 remaining storage slots.

[0049] After determining the remaining storage slots, the embodiment compares the remaining storage slots with the number of groups of network parameters in the first target parameter, and if the number of groups of network parameters in the first target parameter is smaller than the remaining storage slots, the remaining storage slots are determined to satisfy the storage requirements of the first target parameter, and the CPU can transfer the first target parameter from the memory to the target memory. In one embodiment, when writing the network parameters to the target memory, the above-described group-by-group writing method can be adopted.

[0050] The embodiment of the present disclosure maintains a first mapping relationship in the CPU, determines the remaining storage space of the target memory based on the first mapping relationship, and controls the writing of network parameters based on the first mapping relationship, thereby realizing management of the storage space of the video card memory, avoiding the huge pressure on the video card memory caused by too many network parameters required for embedding processing in the model training process, helping to reduce the high requirements for hardware conditions of large-scale model training, and helping to realize the training of large-scale models. Furthermore, since the first mapping relationship in the embodiment is maintained in a memory or cache accessible by the CPU, compared with the technical solution of storing a hash table indicating the mapping relationship in the video card memory in the related art, the video card memory can be fully utilized to perform model training, which further reduces the pressure on the video card memory and helps to reduce the communication overhead between the CPU and the target processor.

[0051] As can be understood, when it is determined that the remaining storage slots meet the storage requirements of the first target parameter, the embodiment can further first assign a storage slot among the remaining storage slots to the first target parameter, and write the first target parameter to the assigned storage slot. For example, if the first target parameter includes network parameters corresponding to 10 feature data, and the storage slots with serial numbers 0 to 49 in the target memory have already stored the network parameters, the first target parameter can be assigned to the storage slots with serial numbers 50 to 49.

[0052] After the storage slot is assigned to the first target parameter, the first mapping relationship can be updated according to the serial number of the assigned storage slot (i.e., the identifier information of the storage slot) and the identifier information of the first target parameter (i.e., the identifier information of the feature data corresponding to the first target parameter). In this way, the correspondence relationship between the storage slot and the network parameter in the first mapping relationship can be maintained.

[0053] It can be understood that, in each round of training, the embodiment also writes the third network parameters required for performing prediction processing on the training data into the target memory, and the computing core of the target processor can call, and the third network parameters can be adjusted according to the call result.This is because the network parameters required for general prediction processing are densely distributed parameters, and the parameters are small, so writing the entire amount of network parameters required for prediction processing into the target memory generally does not cause obvious pressure.In the recommendation model, the third network parameters can be, for example, the network parameters included in the prediction network, and the prediction network can include, for example, a multilayer perceptron (MLP).

[0054] As can be seen, the training process of a deep learning model generally includes three parts. The first part is to obtain the loss of the deep learning model through calculation in the forward calculation process, the second part is to obtain the gradient through calculation in the backward calculation process, and the third part is to update the network parameters of the deep learning model according to the gradient. The calculation core specifically adjusts the first network parameters and the third network parameters according to the gradient obtained by the backward calculation, so that the network parameters of the deep learning model gradually converge.

[0055] In one embodiment, if the remaining storage slots do not meet the storage requirements of the first target parameters, the CPU can, for example, transfer temporarily unnecessary network parameters in the target memory, thereby leaving enough space for the first target parameters and providing conditions for subsequent training of the deep learning model. This embodiment can dynamically adjust the size of the cache space in the target memory, combine the maintenance of the first mapping relationship in the memory, and effectively reduce the communication overhead between the CPU and the target processor.

[0056] Illustratively, the CPU further maintains in a cache or memory a second mapping relationship between the storage slots and the parameter states of the network parameters stored in the storage slots, thereby providing a basis for determining the transferable network parameters.

[0057] For example, the parameter state of a network parameter may include a reference state. If the network parameter is a network parameter required for the current training round, the reference state is set to a reference state, and if the current training round does not require the network parameter, the reference state is set to an unreferenced state. For example, the reference state may be represented by a reference count (abbreviated as Count), where a value of 1 in the reference count represents a reference state, and a value of 0 in the reference count represents an unreferenced state.

[0058] In this embodiment, the second mapping relationship is represented by a mapping table composed of the above-mentioned correspondence relationship between FId, FeaSign and Count, and each FeaSign corresponds to each RefCount, and is used to indicate whether the network parameters required when performing embedding processing in the feature data of the FeaSign identifier are referenced. In this embodiment, the network parameters corresponding to the FeaSign whose RefCount value in the second mapping relationship is 0 can be regarded as transferable network parameters.

[0059] For example, the parameter state of a network parameter may include a usage count. When a network parameter is called in one training round, the usage count is incremented by 1, and the initial value of the usage count may be 0. For example, the usage count may be represented by a frequency count (abbreviated as FreqCount).

[0060] In this embodiment, the second mapping relationship is represented by a mapping table composed of the above-mentioned correspondence relationship between FId, FeaSign, and FreqCount, and each FeaSign corresponds to a respective FreqCount, which is used to indicate the number of times the network parameter is used when performing embedding processing in the feature data of the FeaSign identifier. In this embodiment, the network parameters corresponding to the FeaSigns whose FreqCount values ​​in the second mapping relationship are smaller than a threshold value can be regarded as transferable network parameters.

[0061] For example, the parameter state of the network parameter includes not only the citation state but also the number of uses. The second mapping relationship is represented by a mapping table configured with the above-mentioned correspondence relationship between FId, FeaSign, RefCount and FreqCount, and each FeaSign corresponds to each RefCount and FreqCount. In this embodiment, the network parameter corresponding to the FeaSign whose citation state is not quoted and whose number of uses is less than a threshold can be regarded as a transferable network parameter.

[0062] The method of the above embodiment can timely transmit required network parameters according to demand, leaving sufficient memory slots for training the deep learning model, which is helpful in improving the training efficiency of the deep learning model.

[0063] Illustratively, the embodiment may further compare the first network parameter with the network parameters stored in the target memory, and the network parameters that do not belong to the first network parameter and have a citation status of not citation may be set as transferable network parameters. For example, when determining the transferable network parameters, the number of groups of network parameters that need to be transferred is determined based on, for example, the number of feature data corresponding to the first target parameter, and set as the target group number. Next, the network parameters that are in a non-citation status and have a target group number that is used infrequently are set as transferable network parameters.

[0064] After determining the transferable network parameters, the transferable network parameters can be transferred from the target memory to the memory. And after the transferable network parameters are transferred, the first target parameters are written to the target memory. As can be understood, similar to the above, when writing the first target parameters to the target memory, the first target parameters can be assigned the remaining storage slots in the target memory. As can be understood, the remaining storage slots here include the storage slots where the transferable network parameters are located. Then, the first target parameters are written to the assigned storage slots. After assigning the storage slots, the embodiment can further update the first mapping relationship and the second mapping relationship according to the identifier information of the feature data corresponding to the first target parameters and the serial number of the storage slots assigned to the first target parameters.

[0065] For example, when updating the second mapping relationship, in addition to the need to update the mapping relationship between FId and FeaSign, it is also necessary to update the parameter state of the first network parameter. For example, the citation state of the first network parameter can be changed to a cited state, i.e., the RefCount of FeaSign corresponding to the first network parameter can be changed from 0 to 1. For example, the number of times the first network parameter is used can be incremented by 1, i.e., the FreqCount of FeaSign corresponding to the first network parameter can be incremented by 1.

[0066] In one embodiment, after the computing core completes the adjustment to the first network parameter, the embodiment may further update the reference state of the first network parameter by updating the second mapping relationship, in particular, changing the RefCount of FeaSign corresponding to the first network parameter from 1 to 0.

[0067] In one embodiment, a tertiary buffer structure composed of a target memory, a memory and a hard disk processor is adopted to reduce the storage pressure of the memory and the target memory. When the above-mentioned first target parameter is written to the target memory, the first target parameter can be read from the memory or the hard disk memory. The memory can be a cache of the hard disk memory, and when the memory occupancy rate is high, the CPU can write the data cached in the memory to the hard disk memory. The setting of the tertiary buffer structure can accelerate the search and retrieval of network parameters in the model training process, which is helpful in realizing the training of a large-scale deep learning model, for example, the model parameters of the supported deep learning model can reach the T level.

[0068] For example, when the CPU transfers the transferable network parameters from the target memory to the memory, it can first determine whether the remaining storage space of the memory is less than the space threshold. If it is less than the space threshold, the memory is used as a cache, and the transferable network parameters are written to the hard disk memory through the memory. That is, the transferable network parameters are cached in the memory, and the transferable network parameters cached in the memory are written to the hard disk memory.

[0069] In one embodiment, when determining the first target parameters in the first network parameters, the first network parameters necessary for performing the embedding process in the first training data can be determined, for example, by adopting the method described above. Specifically, first, feature data included in the first training data can be determined, and all network parameters corresponding to the feature data can be set as the first network parameters. Then, a deduplication process is performed on the first network parameters, and the deduplication network parameters are obtained. For example, deduplication can be performed on the first network parameters based on the identifier information of the feature data. Then, based on the first mapping relationship and the identifier information of the deduplication network parameters, the network parameters not stored in the target memory in the deduplication network parameters are determined, and the determined network parameters are set as the first target parameters.

[0070] For example, first, deduplication can be performed on the feature data included in the first training data based on the identifier information of the feature data, and then the network parameters required for performing embedding processing on the deduplication-depleted feature data are set as the deduplication-depleted network parameters.

[0071] The first training data generally includes multiple training data, and different training data may include the same feature data. When all the determined first network parameters are written into the target memory, there exists a situation where the same network parameters are written into multiple slots of the target memory. The embodiment of the present disclosure can avoid the above situation by performing deduplication on the first network parameters, thus reducing the waste of storage space of the target memory, improving the utilization rate of the storage space of the target memory, reducing the pressure on the target memory during large-scale model training, and helping to realize the training of large-scale models.

[0072] It can be understood that after the CPU writes the first target parameters to the target memory, for example, by sending training task information based on the first training data to the target processor, the computing core of the target processor can process the first training data according to the first network parameters stored in the target memory, and adjust the first network parameters according to the processing result. Based on this, the present disclosure provides yet another model processing method. The other model training method is described in detail below with reference to FIG. 3.

[0073] FIG. 3 is a schematic flow chart of a method for training a deep learning model according to another embodiment of the present disclosure.

[0074] 3, the deep learning model training method 300 of the embodiment can include operations S310 to S340. The model training method 300 can be performed by the electronic device described above.

[0075] In operation S310, the first processor determines, based on the first training data of the current training round, first target parameters that need to be written to the target memory in the first network parameters required to perform embedding processing on the first training data.

[0076] According to an embodiment of the present disclosure, the first processor may be the above-mentioned CPU, and the target memory is a memory included in the second processor. The second processor is similar to the above-mentioned target processor, and the implementation manner of the operation S310 is similar to the above-mentioned operation S210, and the description is omitted here.

[0077] In operation S320, the first processor determines the remaining storage slots in the target memory according to the first mapping relationship between the storage slots of the target memory and the network parameters. This operation S320 is similar to the above operation S220, and the description is omitted here.

[0078] In operation S330, the first processor responds by writing the first target parameter to the target memory in response to the remaining storage slots satisfying the storage requirements of the first target parameter, and sends training task information based on the first training data to the second processor.

[0079] In operation S340, the computational core of the second processor is responsive to receiving the training task information to adjust the first network parameters based on the first training data.

[0080] According to the embodiment of the present disclosure, the implementation manner of writing the first target parameters into the target memory is similar to the implementation manner of the above-mentioned operation S230, and thus the description is omitted here.

[0081] According to an embodiment of the present disclosure, after the first processor writes the first target parameters into the target memory, the first processor can further send training task information based on the first training data to the second processor. In this way, after the computing core of the second processor receives the training task information, it can directly call the first network parameters stored in the target memory to process the first training data, and perform backward calculation according to the processing result to obtain gradient data for the first training data, and adjust the first network parameters according to the gradient data.

[0082] According to an embodiment of the present disclosure, the first processor can further send training task information based on the first training data to the second processor in the process of writing the first target parameters to the target memory. In this way, after receiving the training task information, the computing core of the second processor can incrementally call the network parameters stored in the target memory, and if the required network parameters have not yet been written to the target memory, it can provisionally execute the execution of the training task until it first reads the required network parameters from the target memory.

[0083] According to an embodiment of the present disclosure, the first processor can also write the first training data into the cache of the second processor at the same time as writing the first target parameters into the target memory. The training task information can include, for example, forward calculation task information, backward calculation task information, and parameter update task information. The forward calculation task information can include, for example, first training data calling information, network parameter calling information, and loss calculation information. The network parameter calling information can include identifier information of the network parameters that need to be called and calling order information of the network parameters. The backward calculation task information can include, for example, information such as a learning rate, and the parameter update task information can include, for example, adjusting a step size.

[0084] The embodiment of the present disclosure maintains a first mapping relationship in the CPU, determines the remaining storage space of the target memory based on the first mapping relationship, and controls the writing of network parameters based on the first mapping relationship, thereby realizing management of the storage space of the video card memory, avoiding the huge pressure on the video card memory caused by too many network parameters required for embedding processing in the model training process, helping to reduce the high requirements for hardware conditions of large-scale deep learning model training, and helping to realize the training of large-scale deep learning models. Furthermore, since the first mapping relationship in the embodiment is maintained in a memory or cache accessible by the CPU, compared with the technical solution of storing a hash table indicating the mapping relationship in the video card memory in the related art, the video card memory can be fully utilized to perform model training, which further helps to reduce the pressure on the video card memory and save the communication overhead between the CPU and the target processor.

[0085] As can be seen, as described above, in each round of training, the embodiment may further write a third network parameter required for performing prediction processing on the training data to the target memory, call the third network parameter by the computing core of the target processor, and adjust the third network parameter based on the call result.

[0086] In order to better understand the present disclosure, the structure of a processor cache for implementing the model training method provided by the present disclosure will be described in detail below with reference to FIG. 4.

[0087] FIG. 4 is a structural schematic diagram of a processor cache according to an embodiment of the present disclosure. As shown in Fig. 4, in an embodiment 400, in order to realize the method for training a deep learning model provided by the present disclosure, a processor cache structure can include a memory 410 and a target memory 420. This embodiment takes the target memory 420 as an example of a video card memory. It can be understood that the target memory 420 can be any high bandwidth memory (HBM).

[0088] A first hash table 411 and a second hash table 412 may be maintained in the memory 410. The first hash table 411 is used to indicate the first mapping relationship, and the second hash table 412 is used to indicate the second mapping relationship. Specifically, a key in the first hash table is identifier information FeaSign of feature data, and a value in the first hash table is a serial number of a slot stored in the video card memory 420. A key in the second hash table is a serial number of a slot stored in the video card memory 420, and a value is tag information of feature data (Feature Meta, abbreviated as FeaMeta), and the tag information includes identifier information FeaSign of feature data, and is a reference status RefCount and a usage count FreqCount of network parameters required when performing embedding processing in feature data.

[0089] For example, if the video card memory 420 in this embodiment is configured to allow storing 100 sets of network parameters for performing embedding processing on a maximum of 100 pieces of feature data, then the video card memory 420 will include 100 storage slots, and the 100 storage slots will be numbered 0, 1, 2, ..., 98, and 99, respectively. The data cached in each storage slot may include a set of embedding layer network parameters and hyperparameters required when tuning the set of embedding layer network parameters.

[0090] When the processor CPU 430 executes the corresponding operations of the above-mentioned deep learning model training method, it can search the first mapping table to determine the number of available memory slots in the video card memory 420, and allocate the memory slots to the target parameters to be written to the video card memory 420 required when performing embedding processing on the training data, and can perform operations such as querying, adding, and deleting on the information in the first hash table 411 and the second hash table 412 stored in the memory 410 according to needs.

[0091] CPU 430 further copies data that needs to be cached in video card memory 420 to the allocated storage slots when performing corresponding operations in the deep learning model training method, and copies relevant network parameters from video card memory 420 when a target processor such as a GPU completes adjustments to network parameters and needs to release storage slots. During the model training process, CPU 430 essentially plays the role of a cache manager.

[0092] In one embodiment, the video card memory 420 can be a memory in an artificial intelligence chip, specifically a memory in a Konron second generation chip. In this way, when the deep learning model training method is performed in this embodiment, the computing power of the Konron second generation chip can be fully utilized, which is helpful to realize the training of a large-scale recommendation model.

[0093] In one embodiment, by installing multiple target processors in one electronic device, the target processors can perform parallel training on deep learning models based on different training data, thereby improving model training efficiency.

[0094] For example, the target processor includes multiple processors, and in one round of training, multiple batches of data can be acquired, and the multiple batches of data constitute the first training data. This embodiment writes only the network parameters required for embedding in the data of each batch into the target memory of the processor corresponding to each batch, thereby reducing the buffer pressure of the target memory in the target processor.

[0095] For example, the embodiment further includes the steps of: when writing the first target parameter to the target memory, firstly, determining a parameter required for embedding processing in a batch of data corresponding to each processor in the first target parameter, and setting it as a designated parameter for each processor. Then, using the predetermined parameter to replace other parameters other than the designated parameter in the first target parameter, thereby obtaining a parameter to be written for each processor. The number of parameters in the parameter to be written is the same as the number of parameters in the first target parameter. Then, based on the storage slots allocated to the first target parameter, the parameters to be written are written to the target memory included in each processor. This method can make the number of network parameters stored in the multiple target memories included in the multiple target processors and the distribution of the network parameters the same. The predetermined parameter may be a null value. In this way, in addition to reducing the buffer pressure of the target memory in the target processor, it is also useful for synchronizing the network parameters through communication between the multiple target processors.

[0096] For example, multiple target processors can synchronize computed network parameter gradient data based on the network parameters stored in their target memories and the slots the network parameters are located in. In this way, communication overhead between the target processors and the CPU can be reduced.

[0097] Specifically, the computing core of each processor can perform forward calculation and backward calculation based on the training data and network parameters of a batch corresponding to each processor, and obtain gradient data for the first network parameter. For example, the computing core obtains network parameters for embedding and predicting the feature data from the target memory based on the feature data in the training data of the corresponding batch, and processes the feature data based on the network parameters to obtain a processing result. Then, the deep learning model determines a loss for the data of the batch based on the processing result, thereby completing the task of forward calculation. Then, based on the loss and the network parameters for embedding and predicting the feature data, a backward propagation algorithm is adopted to calculate and obtain gradient data for the first network parameter, thereby completing the task of backward calculation. Finally, based on the communication between the storage slot in which the first network parameter is located and the other target processor, gradient data for the first network parameter obtained by the other target processor is obtained. At the same time, gradient data for the third network parameter used in the prediction process obtained by the other target processor can be obtained through communication with the other target processor. Finally, all the gradient data is compiled, and the first network parameter and the third network parameter are adjusted based on the compilation result, thereby completing the task of parameter update.

[0098] The overall flow of the deep learning model training method will be described in detail below with reference to Figure 5.

[0099] FIG. 5 is an overall flowchart of a method for training a deep learning model according to an embodiment of the present disclosure.

[0100] 5, the method 500 for training a deep learning model of the embodiment can include operations S501 to S518. Operations S509 to S512 are performed by the target processor, and the other operations are all performed by the CPU.

[0101] In operation S501, a batch of data is obtained, specifically, a predetermined number of sample data is obtained from a hard disk memory or an external database, so that a deep learning model can be trained.

[0102] In operation S502, the entire data is perturbed to improve the randomness of the training data obtained for each batch.

[0103] In operation S503, data of the current training round is obtained. For example, training data of batch_size*cards can be randomly obtained from the batch data and used as the first training data. The number of cards is the number of target processors provided in the electronic device. The batch_size can be set according to actual needs. For example, the batch_size can be determined based on the storage capacity of the target memory in the target processor. For example, the number of network parameters required to perform embedding processing on the batch_size training data can be related to the storage capacity of the target memory. Specifically, the number of slots stored in the target memory can be twice the number of groups of network parameters required for embedding processing.

[0104] In operation S504, it is determined whether the remaining storage slots of the target memory are sufficient. If the remaining storage slots are sufficient, execute operations S505 to S513; otherwise, execute operations S514 to S516. As can be understood, it can be set that the multiple target processors are the same type of processor, and the sizes of the storage capacities of the multiple target memories included in the multiple target processors are equal.

[0105] In operation S505, a de-duplication process is performed on the network parameters required for embedding the first training data based on the FeaSign of the feature data included in the first training data, and the de-duplication network parameters are obtained.

[0106] In operation S506, determine an increment for the cache parameter in the target memory, i.e., compare the network parameter after deduplication with the network parameter stored in the target storage device according to the first mapping relationship, determine the network parameter that needs to be written to the target memory, and obtain the first target parameter.

[0107] In operation S507, a storage slot is allocated to the network parameters that need to be written into the target memory, and the first mapping relationship and the second mapping relationship are updated according to the allocation result, specifically, the mapping relationship between FId and FeaSign is added to the first mapping relationship, the mapping relationship between FId and FeaMeta is added to the second mapping relationship, and the FeaMeta data of the feature data corresponding to each set of network parameters in the first network parameters is updated, specifically, RefCount and FreqCount are both added to 1.

[0108] In operation S508, the network parameters added to the target memory are copied (pulled), and specifically, the parameters to be written for each target memory are determined based on the predetermined parameters as described above, and the parameters to be written can be written to the assigned storage slots. In this manner, each target processor can call up the network parameters in the target memory and execute operations S509 to S512 based on the training samples of one batch. As can be understood, a third network parameter of the prediction network can further be copied to the target memory included in each target processor among the multiple target processors.

[0109] In operation S509, a forward calculation task is performed to obtain a training sample loss for the batch of the deep learning model.

[0110] In operation S510, a backward calculation task is performed to obtain gradient data for the training samples of a batch based on the loss, which should include gradient data of a first network parameter and gradient data of a third network parameter.

[0111] In operation S511, adopt an All Reduce algorithm to aggregate the gradient data obtained from multiple target processors. It can be understood that, when aggregating the gradient data of the first network parameter, the storage slot where the first network parameter is located should be referred to, because there exists a difference between the values ​​of the first network parameter stored in different target memories.

[0112] In operation S512, values ​​of the network parameters stored in the target memory are updated based on the aggregation results, which may include, for example, calculating an average value for all gradient data for each network parameter, obtaining a final gradient, and updating the value of each network parameter based on the final gradient.

[0113] In operation S513, the RefCount value of the feature data corresponding to the network parameters used in the current batch is decremented by 1. Up to this point, the target processor completes the tuning of the network parameters based on the first training data.

[0114] In operation S514, the transferable network parameters having a RefCount of 0 and a low FreqCount are filtered out. The RefCount of the feature data corresponding to the transferable network parameter is 0, and the value of the FreqCount is lower than the frequency threshold.

[0115] In operation S515, copy the transferable network parameters from the target memory and cache the copied transferable network parameters in the memory.

[0116] In operation S516, the mapping relationship between FeaSign and FId of the feature data corresponding to the transferable network parameters in the first mapping relationship is deleted. After performing operation S516, the operation S504 can be returned to and re-determined whether the remaining storage slots are sufficient.

[0117] According to an embodiment of the present disclosure, after the target processor completes the adjustment to the network parameters based on the first training data, the CPU can, for example, execute operation S517 to determine whether any of the acquired batch data has been trained. That is, determine whether any of the acquired batch data has been used as training data to train the deep learning model. If so, execute operation S518 to copy and write the updated network parameters stored in the target memory (e.g., HMB) to memory or hard disk memory. If not, return to and execute operation S503 to start training the next training round.

[0118] In order to better understand the deep learning model training method provided by the present disclosure, the communication topology of a standalone multi-card processor is described in detail below with reference to FIG. 6.

[0119] FIG. 6 is a communication topology diagram of a stand-alone multi-card processor according to an embodiment of the present disclosure.

[0120] As shown in FIG. 6, in the embodiment 600, the electronic device of the stand-alone multi-card structure may include one CPU and four XPUs, for example, XPU#0 to XPU#3. The CPU may be communicatively connected to the four XPUs via, for example, a PCIe (Peripheral Component Interconnect Express) interface. A network interface controller (NIC) is used to connect the electronic device to a local area network. The NIC is connected to an access switch (TOR Switch) by, for example, Ethernet, so that the electronic device accesses the local area network. Here, the XPU refers to a Konlon chip, and specifically, for example, may refer to a Konlon second generation chip.

[0121] In the four XPUs, XPU#0 and XPU#1, XPU#0 and XPU#3, XPU#1 and XPU#2, and XPU#2 and XPU#3 can be connected via a cache coherence interconnect protocol (CCIX) to form a processor ring. CCIX allows two or more devices to share data inter-sheet interconnect in a cache coherence manner. The inter-sheet interconnect structure provides the basis for the use of the All Reduce algorithm. As can be seen, the topology structure shown in FIG. 6 can be the communication topology of the Konron 2nd generation chip, and the topology structure can achieve AllReduce communication supporting partial sparse parameters (network parameters for embedding processing). As can be seen, the embodiment can adopt a method of using each XPU to broadcast all gradient data to other XPUs and receive all gradient data of other XPUs to adjust network parameters. In this manner, the gradient data of XPU#0 broadcast can be transferred to XPU#2 via XPU#3, XPU#1, or CPU#1, for example.

[0122] In one embodiment, as shown in FIG. 6, when training a deep learning model, two more electronic devices or more electronic devices may be employed, and the multiple electronic devices may be connected via a local area network, and the CPUs in the multiple electronic devices may be communicatively connected via a Common System Interface (QPI), which is an architecture that realizes interconnection between chips.

[0123] Based on the network architecture provided by the present disclosure, AllReduce communication of sparse parameters can be realized, thereby realizing synchronous training of deep learning models on multiple target processors, and further realizing training of large-scale deep learning models, thereby reducing communication overhead.

[0124] According to the embodiments of the present disclosure, an asynchronous pipeline method can further be adopted to train the deep learning model, thereby improving the efficiency of model training.

[0125] FIG. 7 is a principle schematic diagram of training a model in an asynchronous pipeline format according to an embodiment of the present disclosure.

[0126] As shown in FIG. 7, in the embodiment 700, when training a deep learning model, an asynchronous pipeline design can be performed. For example, when the computing core of the target processor executes the training task 730 of the current training round, the CPU performs pre-processing 710 on the training data of the next training round, and after completing the pre-processing, it allocates slots for the target parameters that need to be written to the target memory, and copies the target parameters to the target memory, i.e., executes the task 720 of allocating slots and copying data. In this way, after the computing core executes the training task 730 of the current training round, it can directly execute the training task of the next training round. This method effectively improves the model training efficiency, reduces the interval between the repeated training of two adjacent rounds, and improves the utilization rate of the target processor.

[0127] Specifically, in the embodiment 700, the CPU can respond to the computing core training the first network parameters based on the first training data, and based on the second training data of the next training round, determine the second target parameters that need to be written into the target memory in the second network parameters required for performing embedding processing on the second training data. Then, based on the first mapping relationship between the storage slots of the target memory and the network parameters, determine the remaining storage slots in the target memory. Then, if the remaining storage slots meet the storage requirements of the second target parameters, allocate the storage slots to the second target parameters and write the second target parameters into the target memory.

[0128] Based on the deep learning model training method provided by the present disclosure, the present disclosure further provides a deep learning model training apparatus, which is described in detail below with reference to FIG. 8.

[0129] FIG. 8 is a structural block diagram of a deep learning model training apparatus according to an embodiment of the present disclosure.

[0130] As shown in FIG. 8 , the deep learning model training apparatus 800 of this embodiment may include a target parameter determination module 810, a remaining slot determination module 820 and a parameter writing module 830.

[0131] The target parameter determination module 810 is used to determine, based on the first training data of the current training round, a first target parameter that needs to be written to the target memory in the first network parameter required for performing embedding processing on the first training data. The target memory is a memory included in the target processor. In one embodiment, the target parameter determination module 810 is used to perform the above operation S210, and the description is omitted here.

[0132] The remaining slot determination module 820 determines the remaining storage slots in the target memory according to a first mapping relationship between the storage slots of the target memory and the network parameters. In one embodiment, the remaining slot determination module 820 performs the above operation S220, and the description is omitted here.

[0133] In response to the remaining storage slots satisfying the storage requirements of the first target parameters, the parameter writing module 830 writes the first target parameters to the target memory, so that the computing cores included in the target processor adjust the first network parameters based on the first training data. In one embodiment, the parameter writing module 830 performs the above operation S230, and the description is omitted here.

[0134] According to an embodiment of the present disclosure, the above-mentioned apparatus 800 may further include: a slot allocation module for allocating a storage slot among the remaining storage slots to the first target parameter in response to the remaining storage slots satisfying the storage requirement of the first target parameter, and a first relationship update module for updating the first mapping relationship based on the identifier information of the storage slot allocated to the first target parameter and the identifier information of the first target parameter. The parameter writing module 830 writes the first target parameter to the storage slot allocated to the first target parameter.

[0135] According to an embodiment of the present disclosure, the target parameter determination module 810 may include: a necessary parameter determination submodule for determining first network parameters required for performing embedding processing on the first training data; a deduplication submodule for performing a deduplication process on the first network parameters to obtain deduplication network parameters; and a target parameter determination submodule for determining a network parameter not stored in the target memory in the deduplication network parameters according to the first mapping relationship and the identifier information of the deduplication network parameters, and setting the network parameter as the first target parameter. According to an embodiment of the present disclosure, the above-mentioned device 800 may further include a transfer parameter determination module for determining a transferable network parameter in the network parameters stored in the target memory in response to the remaining storage slots not satisfying the storage requirements of the first target parameter, and a parameter transfer module for transferring the transferable network parameter from the target memory to the memory. The parameter writing module 830 further writes the first target parameter to the target memory in response to the transferable network parameter being transferred to the memory.

[0136] According to an embodiment of the present disclosure, the transfer parameter determination module determines that the network parameter whose parameter state is the target state is a transferable network parameter based on a second mapping relationship between the storage slot of the target memory and the parameter state of the network parameter stored in the storage slot. The parameter state includes at least one of a quoted state and a usage count. The target state includes at least one of the quoted state being a non-quoted state and the usage count being less than a usage count threshold. The above-mentioned apparatus 800 may further include a slot allocation module that allocates the remaining storage slots in the target memory to the first target parameter in response to the transferable network parameter being transferred to the memory, and a second relationship update module that updates the parameter state of the first network parameter by updating the second mapping relationship based on the storage slots allocated to the first target parameter and the storage slots in which other parameters other than the first target parameter are located in the first network parameter.

[0137] According to an embodiment of the present disclosure, the second relationship update module further updates the reference state of the first network parameter by updating the second mapping relationship in response to the computing core completing an adjustment to the first network parameter.

[0138] According to an embodiment of the present disclosure, the parameter transfer module specifically writes the transferable network parameters to the hard disk memory via the memory in response to the remaining storage space of the memory being less than a space threshold.

[0139] According to an embodiment of the present disclosure, the target parameter determination module 810 further determines, in response to the computing core training the first network parameters based on the first training data, a second target parameter that needs to be written to the target memory in the second network parameters required for performing embedding processing on the second training data based on the second training data of the next training round. The remaining slot determination module 820 further determines the remaining storage slots in the target memory based on a first mapping relationship between the storage slots of the target memory and the network parameters. The parameter writing module 830 further writes the second target parameter to the target memory in response to the remaining storage slots satisfying the storage requirements of the second target parameters.

[0140] According to an embodiment of the present disclosure, the target processor includes a plurality of processors, and the first training data includes multiple batches of data corresponding to the plurality of processors, respectively. The parameter writing module 830 may include: a designated parameter determining submodule for determining, for each processor among the plurality of processors, a designated parameter required for performing embedding processing on a batch of data corresponding to each processor in the first target parameter, a parameter replacing submodule for replacing parameters other than the designated parameter in the first target parameter with a predetermined parameter value to obtain parameters to be written for each processor, and a writing submodule for writing the parameters to be written to a target memory included in each processor, so that a computing core included in each processor trains the designated parameter based on the batch of data corresponding to each processor.

[0141] According to an embodiment of the present disclosure, for each batch of data in the multi-batches of data, the number of network parameters required to perform embedding processing on each batch of data is related to the storage capacity of the target memory in the processor corresponding to each batch of data.

[0142] According to an embodiment of the present disclosure, the parameter writing module 830 further writes third network parameters required for performing prediction processing on multiple batches of data to the target memory in each processor, so that the computing cores included in each processor adjust the third network parameters based on one batch of data corresponding to each processor.

[0143] Based on the deep learning model training method provided by another embodiment of the present disclosure, the present disclosure further provides a deep learning model training system, which is described in detail below with reference to FIG. 9 .

[0144] FIG. 9 is a structural block diagram of a deep learning model training system according to an embodiment of the present disclosure.

[0145] As shown in FIG. 9, the deep learning model training system 900 of this embodiment includes a first processor 910 and a second processor 920, where the second processor includes a target memory and a computing core.

[0146] The first processor 910 is configured to: determine, based on the first training data of the current training round, a first target parameter that needs to be written into the target memory in the first network parameter required for performing embedding processing on the first training data; determine remaining storage slots in the target memory based on a first mapping relationship between the storage slots of the target memory and the network parameters; and in response to the remaining storage slots satisfying the storage requirements of the first target parameters, write the first target parameters into the target memory, and send training task information based on the first training data to the second processor. As can be understood, the first processor may be configured to perform the above operations S310 to S330, and the description thereof will be omitted here.

[0147] The second processor 920 is configured to: adjust a first network parameter based on the first training data in response to receiving the training task information;

[0148] According to an embodiment of the present disclosure, the second processor includes multiple processors, and the first training data includes multiple batches of data corresponding to the multiple processors respectively. The above-mentioned first processor 910 is configured to write the first target parameters to the target memory in the following manner: for each processor in the multiple processors, determine a designated parameter required for performing embedding processing on one batch of data corresponding to each processor in the first target parameters, replace other parameters other than the designated parameter in the first target parameters with the predetermined parameter, obtain the parameters to be written for each processor, and write the parameters to be written to the target memory included in each processor.

[0149] According to an embodiment of the present disclosure, the multiple processors are connected to form a processor ring through a cache coherence interconnection protocol, and each processor in the multiple processors is configured to adjust a first network parameter in the following manner: a calculation core performs forward calculation and backward calculation according to a batch of data and a specified parameter corresponding to each processor, and obtains gradient data for the first network parameter; and according to a storage slot where the first network parameter is located, adopts an All reduce algorithm to adjust the first network parameter according to the gradient data for the first network parameter and the gradient data obtained by other processors in the multiple processors.

[0150] According to an embodiment of the present disclosure, the second processor includes an artificial intelligence chip, and the artificial intelligence chip includes a Konloncore second generation chip.

[0151] It should be explained that in the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure, application and other processes of user personal information shall all comply with the provisions of relevant laws and regulations, adopt necessary security measures, and not violate public order and morals. In the technical solutions disclosed herein, before obtaining or collecting user personal information, the permission or consent of the user shall all be obtained.

[0152] According to an embodiment of the present disclosure, the present disclosure further includes an electronic device, a readable storage medium, and a computer program. M provide.

[0153] FIG. 10 illustrates an exemplary block diagram of an exemplary electronic device 1000 for training a deep learning model according to an embodiment of the present disclosure. The electronic device is intended to represent various types of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The electronic device may also represent various types of mobile devices, such as personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure as described and / or claimed herein.

[0154] As shown in Fig. 10, the device 1000 includes a computing unit 1001, which can perform various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 1002 or loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 can further store various programs and data required for the operation of the device 1000. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0155] A number of components in the electronic device 1000 are connected to an I / O interface 1005, including an input unit 1006, such as a keyboard, a mouse, etc., an output unit 1007, such as various types of displays, speakers, etc., a storage unit 1008, such as a magnetic disk, an optical disk, etc., and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 enables the electronic device 1000 to exchange information / data with other devices via a computer network, such as the Internet, and / or various types of electric communication networks.

[0156] The computing unit 1001 may be various general-purpose and / or dedicated processing modules having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various machine learning model algorithm computing units, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the above-described methods and processes, such as the deep learning model training method. For example, in some embodiments, the deep learning model training method may be realized as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, some or all of the computer program may be loaded and / or installed in the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, it may execute one or more steps of the above-described deep learning model training method. Alternatively, in another embodiment, the computing unit 1001 may be configured to perform the method for training a deep learning model in any other suitable form (e.g., via firmware).

[0157] Various embodiments of the systems and techniques described herein may be realized in digital electronic circuitry systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be embodied in one or more computer programs that may be executed and / or interpreted by a programmable system that includes at least one programmable processor, which may be a special purpose or general purpose programmable processor, and may include a processor that is capable of receiving data and instructions from, and transmitting data and instructions to, a storage system, at least one input device, and at least one output device.

[0158] The program codes for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer or other programmable data processing apparatus, so that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are performed. The program codes may be fully executed on the device, partially executed on the device, partially executed on the device as a separate software package and partially executed on a remote device or fully executed on a remote device or server.

[0159] In the context of this disclosure, a machine-readable medium may be a tangible medium, and may contain or store a program for use in or in combination with an instruction execution system, device, or appliance. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or appliance, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections through one or more wires, portable computer disks, hard disks, random access memories (RAMs), read-only memories (ROMs), erasable programmable read-only memories (EPROMs or flash memories), optical fibers, compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0160] A computer may implement the systems and techniques described herein to provide interaction with a user, and may include a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user may provide input to the computer. Other types of devices may also provide interaction with a user, for example, the feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback) and may receive input from the user in any form (including voice input, speech input, or tactile input).

[0161] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser, through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such background components, middleware components, or front-end components. The components of the system can be connected to each other by digital data communication in any form or medium (e.g., a communication network). Exemplary communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0162] A computer system may include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. The relationship between a client and a server is generated by a computer program running on the corresponding computer and having a client-server relationship. A server may be a cloud server, also referred to as a cloud computing server or a cloud host, which is a host product in a cloud computing service system, thereby solving the defects of high management difficulty and weak service scalability existing in conventional physical hosts and VPS services (abbreviated as "Virtual Private Server" or "VPS"). A server may also be a server in a distributed system or a server combined with a blockchain.

[0163] It should be understood that various types of flows shown above may be used, and steps may be rearranged, added, or deleted. For example, each step described in the present invention may be performed in parallel, sequentially, or in a different order, and the present specification is not limited thereto as long as the desired results of the technical solution of the present disclosure can be achieved.

[0164] The specific embodiments described above do not limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, subcombinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principle of the present disclosure should be included within the scope of protection of the present disclosure.

Claims

1. determining, based on first training data of a current training round, a first target parameter that needs to be written to a target memory, which is a memory included in a target processor, in a first network parameter required for performing an embedding process on the first training data; determining remaining storage slots in the target memory based on a first mapping relationship between storage slots of the target memory and network parameters; and writing the first target parameters to the target memory in response to the remaining storage slots satisfying storage requirements for the first target parameters such that a computing core included in the target processor adjusts the first network parameters based on the first training data. How to train a deep learning model.

2. assigning a storage slot among the remaining storage slots to the first target parameter in response to the remaining storage slots satisfying the storage requirements of the first target parameter; updating the first mapping relationship based on identifier information of a storage slot assigned to the first target parameter and identifier information of the first target parameter; wherein writing the first target parameter to the target memory includes writing the first target parameter to a storage slot assigned to the first target parameter. The method of claim 1.

3. Determining a first target parameter that needs to be written to a target memory among the first network parameters required for performing an embedding process on the first training data, determining first network parameters necessary to perform an embedding process on the first training data; and performing a deduplication process on the first network parameters to obtain deduplication network parameters; determining a network parameter that is not stored in the target memory in the de-duplication network parameters based on the first mapping relationship and the identifier information of the de-duplication network parameters, and setting the determined network parameter as the first target parameter. The method of claim 1.

4. determining transferable network parameters among the network parameters stored in the target memory in response to the remaining storage slots not satisfying the storage requirements of the first target parameter; transferring the transferable network parameters from the target memory to a memory; and writing the first target parameters to the target memory in response to the transferable network parameters being transferred to the memory. The method of claim 1.

5. Determining transferable network parameters in the network parameters stored in the target memory includes: determining, based on a second mapping relationship between the storage slots of the target memory and the parameter states of the network parameters stored in the storage slots, that the network parameters whose parameter states are the target state are the transferable network parameters; The parameter status includes at least one of a citation status and a usage count; The target state includes at least one of a citation state being an uncited state and a usage count being less than a usage count threshold; The method comprises: allocating a remaining storage slot in the target memory to the first target parameter in response to the transferable network parameters being transferred to the memory; and updating the parameter state of the first network parameter by updating the second mapping relationship based on a storage slot assigned to the first target parameter and a storage slot in which other parameters other than the first target parameter are located in the first network parameter. The method according to claim 4.

6. and updating a reference state of the first network parameter by updating the second mapping relationship in response to the computation core completing an adjustment to the first network parameter. The method according to claim 5.

7. Transferring the transferable network parameters from the target memory to a memory includes: and in response to the remaining storage space of the memory being less than a space threshold, writing the transferable network parameters via the memory to a hard disk memory. The method according to claim 4.

8. In response to the computing core training the first network parameters based on the first training data, determining, based on second training data of a next training round, second target parameters that need to be written to a target memory in second network parameters required for performing embedding processing on the second training data; determining remaining storage slots in the target memory based on a first mapping relationship between storage slots of the target memory and network parameters; and writing the second target parameter to the target memory in response to the remaining storage slots satisfying the storage request for the second target parameter. The method of claim 1.

9. The target processor includes a plurality of processors, and the first training data includes multiple batches of data corresponding to the plurality of processors, respectively, and writing the first target parameters to the target memory includes: determining, for each processor in the plurality of processors, designated parameters necessary for performing embedding processing on one batch of data corresponding to each processor in the first target parameters; replacing parameters other than the designated parameters in the first target parameters with a predetermined parameter value to obtain parameters to be written to each of the processors; writing the parameters to be written to a target memory included in each of the processors such that a computing core included in each of the processors trains the designated parameters based on a batch of data corresponding to each of the processors. The method of claim 1.

10. For each batch of data in the multi-batch data, the number of network parameters required to perform embedding processing on each batch of data is related to the storage capacity of a target memory in a processor corresponding to each batch of data.

10. The method of claim 9.

11. The method further includes writing the third network parameters to a target memory in each of the processors so that a calculation core included in each of the processors adjusts a third network parameter required for performing a prediction process on the multiple batches of data based on one batch of data corresponding to each of the processors.

10. The method of claim 9.

12. A first processor determines a first target parameter that needs to be written to a target memory, which is a memory included in a second processor, in a first network parameter required for performing an embedding process on the first training data based on the first training data of a current training round; a first processor determining remaining storage slots in the target memory based on a first mapping relationship between storage slots of the target memory and network parameters; a first processor, in response to the remaining storage slots satisfying the storage requirements of the first target parameters, writing the first target parameters to the target memory and sending training task information based on the first training data to the second processor; and adjusting the first network parameters based on the first training data in response to the second processor's computational core receiving the training task information. How to train a deep learning model.

13. The second processor includes a plurality of processors, and the first training data includes multiple batches of data corresponding to the plurality of processors, respectively. Writing the first target parameters to the target memory includes: determining, for each processor in the plurality of processors, a designated parameter required for performing an embedding process on one batch of data corresponding to each processor in the first target parameters; replacing parameters other than the designated parameters in the first target parameters with predetermined parameters to obtain parameters to be written to each of the processors; writing the parameters to be written to a target memory included in each of the processors. The method of claim 12.

14. The plurality of processors are connected to form a processor ring via a cache coherence interconnect protocol, and adjusting the first network parameter based on the first training data includes: A computing core of each processor in the plurality of processors performs forward calculation and backward calculation based on a batch of data corresponding to each processor and the specified parameters to obtain gradient data for the first network parameters; Each of the processors adjusts the first network parameter based on a storage slot in which the first network parameter is located, using an All Reduce algorithm to adjust the first network parameter based on gradient data for the first network parameter and gradient data acquired by other processors in the plurality of processors. The method of claim 13.

15. The second processor includes an artificial intelligence chip, and the artificial intelligence chip includes a Konroncore second generation chip. The method according to any one of claims 12 to 14.

16. a target parameter determination module for determining, based on first training data of a current training round, a first target parameter that needs to be written to a target memory, which is a memory included in a target processor, in a first network parameter required for performing an embedding process on the first training data; a remaining slot determination module that determines remaining storage slots in the target memory based on a first mapping relationship between storage slots of the target memory and network parameters; and a parameter writing module configured to write the first target parameters to the target memory in response to the remaining storage slots satisfying storage requirements for the first target parameters such that a computing core included in the target processor adjusts the first network parameters based on the first training data. Training equipment for deep learning models.

17. a slot allocation module that allocates a storage slot among the remaining storage slots to the first target parameter in response to the remaining storage slots satisfying the storage requirements of the first target parameter; A first relationship update module updates the first mapping relationship according to identifier information of a storage slot assigned to the first target parameter and identifier information of the first target parameter; The parameter writing module writes the first target parameter to a storage slot assigned to the first target parameter.

17. The apparatus of claim 16.

18. The target parameter determination module includes: a necessary parameter determination submodule for determining a first network parameter required for performing an embedding process on the first training data; a deduplication submodule for performing a deduplication process on the first network parameters and obtaining deduplication network parameters; a target parameter determination submodule for determining a network parameter that is not stored in the target memory in the deduplication network parameters based on the first mapping relationship and the identifier information of the deduplication network parameters, and setting the determined network parameter as the first target parameter.

17. The apparatus of claim 16.

19. a transfer parameter determination module that determines a transferable network parameter among the network parameters stored in the target memory in response to the remaining storage slots not satisfying the storage requirements of the first target parameter; a parameter transfer module for transferring the transferable network parameters from the target memory to the memory; The parameter writing module further writes the first target parameters to the target memory in response to the transferable network parameters being transferred to the memory.

17. The apparatus of claim 16.

20. The transfer parameter determination module includes: Determine that the network parameters whose parameter states are the target states are the transferable network parameters based on a second mapping relationship between the storage slots of the target memory and the parameter states of the network parameters stored in the storage slots; The parameter status includes at least one of a citation status and a usage count; The target state includes at least one of a citation state being an uncited state and a usage count being less than a usage count threshold; The apparatus comprises: a slot allocation module that allocates remaining storage slots in the target memory to the first target parameters in response to the transferable network parameters being transferred to the memory; and a second relationship update module for updating a parameter state of the first network parameter by updating the second mapping relationship based on a storage slot assigned to the first target parameter and a storage slot in which other parameters than the first target parameter are located in the first network parameter.

20. The apparatus of claim 19.

21. The second relationship updating module further comprises: updating a reference state of the first network parameter by updating the second mapping relationship in response to the computation core completing an adjustment to the first network parameter.

21. The apparatus of claim 20.

22. The parameter transfer module includes: and writing the transferable network parameters via the memory to a hard disk memory in response to the remaining storage space of the memory being less than a space threshold.

20. The apparatus of claim 19.

23. The target parameter determination module further determines, in response to the computing core training the first network parameters based on the first training data, second target parameters that need to be written to the target memory in the second network parameters required for performing embedding processing on the second training data based on second training data of a next training round; The remaining slot determination module further determines remaining storage slots in the target memory based on a first mapping relationship between the storage slots of the target memory and network parameters; The parameter writing module further writes the second target parameter to the target memory in response to the remaining storage slots satisfying the storage request for the second target parameter.

17. The apparatus of claim 16.

24. The target processor includes a plurality of processors, and the first training data includes multiple batches of data corresponding to the plurality of processors respectively, and the parameter writing module: a designated parameter determination submodule that determines, for each processor in the plurality of processors, designated parameters required for performing embedding processing on one batch of data corresponding to each processor in the first target parameters; a parameter replacement submodule for replacing parameters other than the designated parameters in the first target parameters with a predetermined parameter value to obtain parameters to be written to each of the processors; a write submodule for writing the parameters to be written to a target memory included in each of the processors so that a computing core included in each of the processors trains the designated parameters based on a batch of data corresponding to each of the processors.

17. The apparatus of claim 16.

25. For each batch of data in the multi-batch data, the number of network parameters required to perform embedding processing on each batch of data is related to the storage capacity of a target memory in a processor corresponding to each batch of data.

25. The apparatus of claim 24.

26. The parameter writing module further includes: The third network parameters are written to a target memory in each of the processors so that a calculation core included in each of the processors adjusts a third network parameter required for performing a prediction process on the multi-batch data based on one batch of data corresponding to each of the processors.

25. The apparatus of claim 24.

27. a first processor and a second processor, the second processor including a target memory and a computation core; The first processor, Determine, based on first training data of a current training round, first target parameters that need to be written to the target memory in a first network parameter required for performing an embedding process on the first training data; determining remaining storage slots in the target memory based on a first mapping relationship between the storage slots of the target memory and network parameters; configured to, in response to the remaining storage slots fulfilling the storage request for the first target parameter, write the first target parameter to the target memory and send training task information based on the first training data to the second processor; The second processor is configured to adjust the first network parameters based on the first training data in response to the computing core receiving the training task information. A system for training deep learning models.

28. the second processor includes a plurality of processors, the first training data includes multiple batches of data respectively corresponding to the plurality of processors, and the first processor is configured to write the first target parameters to the target memory in the following manner: determining, for each processor in the plurality of processors, designated parameters necessary for performing an embedding process on one batch of data corresponding to each processor in the first target parameters; Substituting a predetermined parameter for a parameter other than the designated parameter in the first target parameter to obtain a parameter to be written to each of the processors; The parameters to be written are written to a target memory included in each of the processors.

28. The system of claim 27.

29. the plurality of processors are connected to form a processor ring via a cache coherence interconnect protocol, and each of the processors is configured to adjust the first network parameter in the following manner: A calculation core performs forward calculation and backward calculation according to a batch of data corresponding to each processor and the specified parameters to obtain gradient data for the first network parameters; According to a storage slot in which the first network parameter is located, an All Reduce algorithm is adopted to adjust the first network parameter according to gradient data for the first network parameter and gradient data obtained by other processors in the plurality of processors.

30. The system of claim 28.

30. The second processor includes an artificial intelligence chip, and the artificial intelligence chip includes a Konroncore second generation chip. A system according to any one of claims 27 to 29.

31. At least one processor; a memory in communication with the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform the method according to any one of claims 1 to 11. electronic equipment.

32. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause a computer to carry out the method according to any one of claims 1 to 11. A non-transitory computer-readable storage medium.

33. When executed by a processor, the method implements the steps of the method according to any one of claims 1 to 11. Computer program.

Citation Information

Patent Citations

  • Storage space allocation method and device

    CN110532198A

  • Prediction model parameter updating method and device

    CN111898740A

  • Model training method and device thereof

    CN113159284A

  • Learning platform for patient journey mapping

    US20210027896A1

  • Accelerated embedding layer computations

    WO2021067376A1