Method, apparatus, device, and storage medium for training a deep learning model
By dividing the training data into multiple data sets and exchanging them within the computing nodes, the problem of low communication bandwidth between different computing nodes is solved, efficient deep learning model training is realized, and training time is shortened.
Patent Information
- Application Number
- CN202210275033.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-03-18
AI Technical Summary
During the training of deep learning models, due to the low communication bandwidth between different computing nodes, communication efficiency is low and the overall training time is longer.
The training data is divided into multiple data sets, and data exchange is performed within the same computing node. Using the internal communication characteristics of high bandwidth, data exchange is performed in the computing node cluster through all-to-all operations to obtain the first exchange result, and then data exchange is performed within the computing node to obtain the second exchange result, and finally model training is performed using the second exchange result.
It improves the communication efficiency between processing units, shortens the overall training completion time, and optimizes the training process of deep learning models.
Smart Images

Figure CN114626523B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular to technologies such as artificial intelligence and deep learning. Background Art
[0002] In the field of deep learning, MoE (Mixure-of-Experts) is one of the technical paths to implement the training of ultra-large-scale models. In MoE, an all-to-all communication method can be adopted. The all-to-all operation is a communication operation. For example, in a deep learning task, data can be exchanged between processes through the all-to-all operation, and the exchanged data can be used for subsequent calculations. Summary of the Invention
[0003] The present disclosure provides a method, an apparatus, a device, and a storage medium for training a deep learning model.
[0004] According to one aspect of the present disclosure, a method for training a deep learning model is provided, including: dividing training data into N first data sets, where N is an integer greater than 1; according to the N first data sets, performing data exchange with a target computing node in a computing node cluster where the current computing node is located to obtain a first exchange result; according to the first exchange result, performing data exchange with a target processing unit in the current computing node to obtain a second exchange result; and training the deep learning model by using the second exchange result.
[0005] According to another aspect of the present disclosure, an apparatus for training a deep learning model is provided, including: a dividing module configured to divide training data into N first data sets, where N is an integer greater than 1; a first exchange module configured to perform data exchange with a target computing node in a computing node cluster where the current computing node is located according to the N first data sets to obtain a first exchange result; a second exchange module configured to perform data exchange with a target processing unit in the current computing node according to the first exchange result to obtain a second exchange result; and a training module configured to train the deep learning model by using the second exchange result.
[0006] Another aspect of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method shown in the embodiments of the present disclosure.
[0007] In accordance with another aspect of the embodiments of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method shown in the embodiments of the present disclosure.
[0008] In accordance with another aspect of the embodiments of the present disclosure, there is provided a computer program product including computer programs / instructions, characterized in that when the computer programs / instructions are executed by a processor, the steps of the method shown in the embodiments of the present disclosure are implemented.
[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0011] Figure 1 is a schematic diagram of a system architecture to which a method, apparatus, electronic device, and storage medium for training a deep learning model according to the embodiments of the present disclosure can be applied;
[0012] Figure 2 Schematically shows a flowchart of a method for training a deep learning model according to an embodiment of the present disclosure;
[0013] Figure 3 Schematically shows a flowchart of a method for data exchange with a target computing node according to an embodiment of the present disclosure;
[0014] Figure 4 Schematically shows a flowchart of a method for data exchange with a target processing unit according to an embodiment of the present disclosure;
[0015] Figure 5 Schematically shows a flowchart of a method for training a deep learning model according to an embodiment of the present disclosure;
[0016] Figure 6A Schematically shows a schematic diagram of training a deep learning model according to another embodiment of the present disclosure;
[0017] Figure 6B Schematically shows a schematic diagram of training a deep learning model according to another embodiment of the present disclosure;
[0018] Figure 6C Schematically shows a schematic diagram of training a deep learning model according to another embodiment of the present disclosure;
[0019] Figure 7A block diagram schematically showing an apparatus for training a deep learning model according to an embodiment of the present disclosure; and
[0020] Figure 8 A block diagram schematically showing an example electronic device that can be used to implement the embodiments of the present disclosure. Detailed implementation manners
[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] The following will be combined with Figure 1 to describe the system architecture of the method, apparatus, electronic device, and storage medium that can apply the method for training a deep learning model provided by the present disclosure.
[0023] Figure 1 is a schematic diagram of the system architecture of the method, apparatus, electronic device, and storage medium that can apply the method for training a deep learning model according to an embodiment of the present disclosure. It should be noted that Figure 1 The illustration shown is only an example of the system architecture that can apply the embodiments of the present disclosure to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.
[0024] As Figure 1 shown, the system architecture 100 may include computing nodes 110, 120, and a network 130. The computing nodes 110 and 120 may each include a plurality of processing units. Exemplarily, in this embodiment, the computing node 110 may include, for example, processing units 111, 112. The computing node 120 may include, for example, processing units 121, 122.
[0025] According to an embodiment of the present disclosure, the network 130 may be a medium for providing a communication link between the processing units 111, 112, 121, and 122. The network 130 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0026] According to an embodiment of the present disclosure, the computing nodes 110 and 120 can be, for example, servers. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system or a server combined with a blockchain.
[0027] According to an embodiment of the present disclosure, the processing units 111, 112, 121, and 122 can include, for example, a GPU (Graphics Processing Unit), a CPU (Central Processing Unit), an NPU (Natural-network Processing Unit), and so on.
[0028] According to an embodiment of the present disclosure, the communication bandwidth between the processing units within the same computing node is relatively high. While the communication bandwidth between the processing units in different computing nodes is relatively low. Exemplarily, in this embodiment, the communication bandwidth between the processing units 111 and 112 is high, and the communication bandwidth between the processing units 121 and 122 is high. The communication bandwidth between the processing unit 111 and the processing unit 121, between the processing unit 111 and the processing unit 122, between the processing unit 112 and the processing unit 121, and between the processing unit 112 and the processing unit 122 is low.
[0029] Exemplarily, when the processing units 111, 112, 121, and 122 need to perform all-to-all communication, the processing unit 111 needs to send the same amount of data to the processing units 112, 121, and 122. Since the communication bandwidth between the processing unit 111 and the processing unit 112 is high, the time required for the processing unit 111 to send data to the processing unit 112 is relatively short. Since the communication bandwidth between the processing unit 111 and the processing units 121 and 122 is low, the time for the processing unit 111 to send data to the processing units 121 and 122 is relatively long.
[0030] It can be seen that when training a deep learning model, communication is required between the computing nodes to exchange training data. However, the communication efficiency between the computing nodes based on different hardware is low, resulting in a relatively long overall training time.
[0031] According to an embodiment of the present disclosure, the training data in each computing node can be divided into N first data sets, where N is an integer greater than 1. According to the N first data sets, data exchange is performed with a target computing node in the computing node cluster where the current computing node is located to obtain a first exchange result. Then, according to the first exchange result, data exchange is performed with a target processing unit in the current computing node to obtain a second exchange result. Next, the deep learning model is trained using the second exchange result. During the training process, the characteristic that the communication bandwidth between processing units within the same computing node is relatively high while the communication bandwidth between different computing nodes is relatively low is utilized, thereby improving the communication efficiency between processing units and shortening the overall training completion time.
[0032] In the technical solution of the present disclosure, the processing of data such as the training data involved in collection, storage, use, processing, transmission, provision, disclosure, and application all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.
[0033] The following will be combined with Figure 2 to describe the method for training a deep learning model provided by the present disclosure.
[0034] Figure 2 A flowchart of a method for training a deep learning model according to an embodiment of the present disclosure is schematically shown.
[0035] As Figure 2 shown, the method 200 for training a deep learning model includes operations S210 to S240. This method can be executed by the processing unit shown above, for example. Exemplarily, in this embodiment, the processing unit as the execution subject is referred to as the current processing unit, and the computing node where the current processing unit is located is referred to as the current computing node.
[0036] In operation S210, the training data is divided into N first data sets. Where N is an integer greater than 1.
[0037] According to an embodiment of the present disclosure, the training data can be, for example, data for training a deep learning model. The training data in the computing node can include any number of data. The present disclosure does not make a specific limitation on the data volume of the training data.
[0038] Exemplarily, in this embodiment, the training data can be evenly divided into N parts, and each part constitutes a first data set.
[0039] According to an embodiment of the present disclosure, for example, the value of N can be determined according to the number of computing nodes in the computing node cluster where the current computing node is located. For example, if the number of computing nodes in the computing node cluster is 2, then the value of N can be determined to be 2, which is the same as the number of computing nodes.
[0040] Then, in operation S220, data exchange is performed with a target computing node within the computing node cluster where the current computing node is located according to the N first data sets, and a first exchange result is obtained.
[0041] According to an embodiment of the present disclosure, the target computing node may include, for example, other computing nodes in the computing node cluster except the current computing node.
[0042] According to an embodiment of the present disclosure, data exchange between the current computing node and the target computing node may be performed, for example, through an all-to-all operation.
[0043] In operation S230, data exchange is performed with a target processing unit in the current computing node according to the first exchange result, and a second exchange result is obtained.
[0044] According to an embodiment of the present disclosure, the target processing unit may include, for example, other processing units in the current computing node except the current processing unit.
[0045] According to an embodiment of the present disclosure, data exchange between the current processing unit and the target processing unit may be performed, for example, through an all-to-all operation.
[0046] In operation S240, the deep learning model is trained using the second exchange result.
[0047] According to an embodiment of the present disclosure, for each processing unit in the computing node, the training data of each processing unit is divided into multiple data sets. Then, data exchange is performed with other computing nodes in the computing node cluster to obtain a first exchange result. Next, data exchange is performed again within the computing node according to the first exchange result of each processing unit to obtain a second exchange result. Then, the deep learning model is trained using the second exchange result. By utilizing the characteristic that the communication bandwidth between processing units within the same computing node is relatively high while the communication bandwidth between different computing nodes is relatively low, the communication efficiency is improved and the overall training completion time is shortened.
[0048] According to an embodiment of the present disclosure, the deep learning model may include, for example, MoE (Mixure-of-Experts).
[0049] The following will be combined with Figure 3 A method for performing data exchange with a target computing node within the computing node cluster where the current computing node is located provided by the present disclosure will be described.
[0050] Figure 3 A flowchart of a method for performing data exchange with a target computing node according to an embodiment of the present disclosure is schematically shown.
[0051] As Figure 3As shown, the method 320 for data exchange with a target computing node includes, in operation S321, determining corresponding processing units in each target computing node.
[0052] According to an embodiment of the present disclosure, for example, the first unit number of the current processing unit in the current computing node can be obtained. Then, the processing unit in the target computing node with a unit number matching the first unit number is determined as the corresponding processing unit.
[0053] Exemplarily, in this embodiment, each computing node in the computing node cluster can generate unit numbers for each processing unit in the computing node according to the same number generation rule. If the numbers of two processing units are the same in their respective nodes, it means that the unit numbers of the two processing units match.
[0054] For example, the computing node cluster may include computing nodes Node a1 and Node a2. The current computing node can be Node a1, and the current processing unit in the current computing node can be processing unit Unit a1_2, and the number of processing unit Unit a1_2 is "2". Based on this, it can be determined that the processing unit Unit a2_2 with the same number "2" in computing node Node a2 is the corresponding processing unit corresponding to Unit a1_2.
[0055] In operation S322, the target first data sets corresponding to each corresponding processing unit in the N first data sets are sent to the corresponding processing units.
[0056] According to an embodiment of the present disclosure, numbers can be set for each computing node in the computing node cluster in advance from 1 to N. Based on this, in operation S322, for example, for each target computing node, the node number i of the target computing node can be obtained, where i is an integer and 1 ≤ i ≤ N. Then, the i-th first data set in the N first data sets is determined as the target first data set, and the target first data set is sent to the corresponding processing unit in the target computing node.
[0057] For example, the computing node cluster may include computing nodes Node b1, Node b2, and Node b3. The node number of Node b1 can be set to "1", the node number of Node b2 can be set to "2", and the node number of Node b3 can be set to "3". For each processing unit in Node b1, the training data in each processing unit can be divided into 3 parts. The second first data set among the 3 first data sets is used as the target first data set corresponding to Node b2 and sent to the corresponding processing unit in Node b2. The third first data set among the 3 first data sets is used as the target first data set corresponding to Node b3 and sent to the corresponding processing unit in Node b3.
[0058] In operation S323, receive the second data sets from each corresponding processing unit.
[0059] According to an embodiment of the present disclosure, each corresponding processing unit may also send the corresponding data set, that is, the second data set, to the current processing unit in the same manner as the above operation S322.
[0060] In operation S324, determine each second data set and the other first data sets in the N first data sets except the target first data set as the first exchange result.
[0061] The following will be combined with Figure 4 Describe the method for data exchange with the target processing unit in the current computing node provided by the present disclosure.
[0062] Figure 4 Schematically shows a flowchart of a method for data exchange with a target processing unit according to an embodiment of the present disclosure.
[0063] As Figure 4 shown, the method 430 for data exchange with the target processing unit includes, in operation S431, dividing the first exchange result into M first data according to the total number M of processing units in the computing node cluster. Wherein, M is an integer greater than 1.
[0064] In operation S432, send the target first data corresponding to each target processing unit in the M first data to the target processing unit.
[0065] According to an embodiment of the present disclosure, for example, the total number M' of processing units in the current computing node can be obtained. Then, for the j-th first data among the M first data, where j is an integer and 1 ≤ j ≤ M', determine the intermediate parameter y according to the following formula, where % represents the modulo operation:
[0066] y = (j - 1) % M' + 1,
[0067] Next, send the j-th first data to the y-th processing unit in the current computing node.
[0068] For example, the current computing node may include processing units Unit c1 and Unit c2. For the current processing unit Unit c1, the first exchange result can be divided into 4 first data. For the first of the 4 first data, the intermediate parameter y=(1 - 1)%2 + 1 = 1 can be determined, that is, the first of the 4 first data should be sent to the first processing unit in the current computing node. However, since the current processing unit Unit c1 is the first processing unit, Unit c1 does not need to send the first first data. For the second of the 4 first data, the intermediate parameter y=(2 - 1)%2 + 1 = 2 can be determined, that is, the second first data should be sent to the second processing unit Unit c2. Similarly, it can be determined that the third first data does not need to be sent, and the fourth first data is sent to the processing unit Unit c2.
[0069] In operation S433, receive the second data from each target processing unit.
[0070] According to an embodiment of the present disclosure, each target processing unit can also send corresponding data, that is, the second data, to the current processing unit in the same manner as the above operation S432.
[0071] In operation S434, determine each second data and the other first data among the M first data except the target first data as the second exchange result.
[0072] According to an embodiment of the present disclosure, by dividing the training data of each processing unit into multiple data sets. Then exchange with other computing nodes in the computing node cluster to obtain the first exchange result. Next, perform an exchange again within the computing node according to the first exchange result of each processing unit to obtain the second exchange result. The all-to-all communication between each processing unit is realized, and the advantage of high communication bandwidth between processing units within the same computing node is utilized to improve the communication efficiency.
[0073] The following will be combined with Figure 5 Describe the method for training a deep learning model using the second exchange result provided by the present disclosure.
[0074] Figure 5 Schematically shows a flowchart of a method for training a deep learning model according to an embodiment of the present disclosure.
[0075] As Figure 5As shown, the method 550 for training the deep learning model includes, in operation S551, inputting the second exchange result into the deep learning model to obtain an output result.
[0076] In operation S552, adjust the parameters of the deep learning model according to the output result.
[0077] According to an embodiment of the present disclosure, for example, a loss function can be used to determine a loss value corresponding to the output result, and then the parameters of the deep learning model are adjusted according to the loss value. Among them, the loss function can be selected according to actual needs, and the present disclosure does not make specific limitations thereon.
[0078] Next, with reference to Figures 6A - 6C , the method for training the deep learning model shown above will be further described in conjunction with specific embodiments. Those skilled in the art can understand that the following exemplary embodiments are only for understanding the present disclosure, and the present disclosure is not limited thereto.
[0079] Figure 6A FIG. schematically shows a schematic diagram of training a deep learning model according to another embodiment of the present disclosure.
[0080] In Figure 6A it is shown that, by way of example, in this embodiment, there are 2 computing nodes Node d1 and Node d2 in the computing node cluster, and the number of processing units included in each computing node is 2. Node d1 includes processing units G0 and G1, and Node d2 includes processing units G2 and G3. Therefore, there are a total of 4 processing units in the computing node cluster.
[0081] According to an embodiment of the present disclosure, a process corresponds to each processing unit, and this process needs to communicate with the processes in other processing units. Based on this, for each processing unit, the data to be communicated in each processing unit can be evenly divided into 2 data sets. For example, the data 1_1 and 2_1 in the processing unit G0 can be divided into a data set S01, and the data 3_1 and 4_1 can be divided into a data set S02. The data 1_2 and 2_2 in the processing unit G1 can be divided into a data set S11, and the data 3_2 and 4_2 can be divided into a data set S12. The data 1_3 and 2_3 in the processing unit G2 can be divided into a data set S21, and the data 3_3 and 4_3 can be divided into a data set S22. The data 1_4 and 2_4 in the processing unit G3 can be divided into a data set S32, and the data 3_4 and 4_4 can be divided into a data set S33.
[0082] Then, as shown in FIG. 6B, for example, each process can obtain the node number i of the computing node where the other process is located, where i is an integer and 1 ≤ i ≤ 2. Then, the i-th data set in the data set is sent to the corresponding processing unit in the i-th computing node.
[0083] Exemplarily, in this embodiment, the Kth (1 ≤ K ≤ 2) (i.e., numbered K) processing units of each computing node can be used as corresponding processing units for each other. Based on this, the Kth processes of each computing node can form a virtual group, and all-to-all communication is performed using the Kth data of each processing unit within the virtual group to exchange data sets.
[0084] For example, the process corresponding to the first processing unit G0 in Node d1 and the process corresponding to the first processing unit G2 in Node d2 can form a virtual group, and the process corresponding to the second processing unit G1 in Node d1 and the process corresponding to the second processing unit G3 in Node d2 can form a virtual group.
[0085] The first process in each virtual group retains the first data set within the corresponding processing unit and sends the second data set within the corresponding processing unit to the second process. The second process in each virtual group retains the second data set within the corresponding processing unit and sends the first data set within the corresponding processing unit to the first process.
[0086] For example, for the process corresponding to the processing node G0 in Node d1, it can retain the first data set S01 within G0 and send the second data set S02 within G0 to the corresponding processing unit G2 in Node d2. For the process corresponding to the processing node G1 in Node d1, it can retain the first data set S11 within G1 and send the second data set S12 within G1 to the corresponding processing unit G3 in Node d2. For the process corresponding to the processing node G2 in Node d2, it can retain the second data set S22 within G2 and send the first data set S21 within G2 to the corresponding processing unit G0 in Node d1. For the process corresponding to the processing node G3 in Node d2, it can retain the second data set S32 within G3 and send the first data set S31 within G3 to the corresponding processing unit G1 in Node d1.
[0087] Next, as shown in 6C for example, each process further interacts with the first exchange result obtained after exchanging the data sets within the same computing node. In this embodiment, each process can split the corresponding first exchange result into 4 parts, and for the jth part, it is sent to the yth processing unit within the same computing node, where y = (j - 1) % 2 + 1, 1 ≤ k ≤ 2, and % represents the modulo operation.
[0088] For example, G0 can divide the first exchange result into data 1_1, 2_1, 1_3, and 2_3. For the first data 1_1, determine y = (1 - 1) % 2 + 1 = 1, so the first data 1_1 is retained. For the second data 2_1, determine y = (2 - 1) % 2 + 1 = 2, so the second data 2_1 is sent to the second processing unit G1. For the third data 1_3, determine y = (3 - 1) % 2 + 1 = 1, so the fourth data 1_3 is retained. For the fourth data 2_3, determine y = (4 - 1) % 2 + 1 = 2, so the fourth data 2_3 is sent to the second processing unit G1.
[0089] Similarly, G1 can divide the first exchange result into data 1_2, 2_2, 1_4, and 2_4, send the first data 1_2 and the third data 1_4 to the first processing unit G0, and retain the second data 2_2 and the fourth data 2_4. G2 can divide the first exchange result into data 3_1, 4_1, 3_3, and 4_3, send the second data 4_1 and the fourth data 4_3 to the fourth processing unit G3, and retain the first data 3_1 and the third data 3_3. G3 can divide the first exchange result into data 3_2, 4_2, 3_4, and 4_4, send the first data 3_2 and the third data 3_4 to the third processing unit G2, and retain the second data 4_2 and the fourth data 4_4.
[0090] Next, each processing unit can perform calculation operations in the deep learning model training based on the second exchange result after the two exchanges.
[0091] According to an embodiment of the present disclosure, the calculation operation may include, for example, calculation operations such as addition, subtraction, multiplication, cos, and sin.
[0092] According to some other embodiments of the present disclosure, after the calculation operation is completed, the above two data exchange processes may be performed again on the calculation result as needed until the deep model training is completed.
[0093] According to an embodiment of the present disclosure, for each processing unit in the computing node, the training data of each processing unit is divided into multiple data sets. Then, it is exchanged with other computing nodes in the computing node cluster to obtain the first exchange result. Next, according to the first exchange result of each processing unit, an internal exchange is performed within the computing node to obtain the second exchange result. Then, the second exchange result is used to train the deep learning model. By utilizing the characteristic that the communication bandwidth between processing units within the same computing node is relatively high while the communication bandwidth between different computing nodes is relatively low, the communication efficiency is improved, and the overall training completion time is shortened.
[0094] The following will be combined with Figure 7A device for training a deep learning model provided by the present disclosure is described.
[0095] Figure 7 A block diagram of a device for training a deep learning model according to an embodiment of the present disclosure is schematically shown.
[0096] As Figure 7 shown, the device 700 includes a partitioning module 710, a first switching module 720, a second switching module 730, and a training module 740.
[0097] The partitioning module 710 is configured to partition training data into N first data sets, where N is an integer greater than 1.
[0098] The first switching module 720 is configured to perform data exchange with a target computing node in a computing node cluster where the current computing node is located according to the N first data sets to obtain a first switching result.
[0099] The second switching module 730 is configured to perform data exchange with a target processing unit in the current computing node according to the first switching result to obtain a second switching result.
[0100] The training module 740 is configured to train the deep learning model by using the second switching result.
[0101] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.
[0102] Figure 8 A block diagram of an exemplary electronic device 800 that can be used to implement the embodiments of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0103] As Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0104] Multiple components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0105] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the method of training a deep learning model. For example, in some embodiments, the method of training a deep learning model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the method of training a deep learning model described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the method of training a deep learning model in any other appropriate way (e.g., by means of firmware).
[0106] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0107] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0108] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0110] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0111] A computer system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other.
[0112] The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.
[0113] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.
[0114] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for training a deep learning model, which is applied to a current processing unit deployed in a computing node cluster. The computing node cluster includes multiple computing nodes, and each computing node includes multiple processing units. The computing node where the current processing unit is located is the current computing node. The method includes: Dividing the training data in the current processing unit into N first data sets, where N is an integer greater than 1, and N is the number of the multiple computing nodes; Performing data exchange with a target computing node according to the N first data sets to obtain a first exchange result, including: sending a target first data set among the N first data sets to the corresponding processing unit in the target computing node; receiving a second data set from each corresponding processing unit; the target computing node includes other computing nodes except the current computing node among the multiple computing nodes; Performing data exchange with a target processing unit in the current computing node according to the first exchange result to obtain a second exchange result, including: dividing the first exchange result into M first data; sending the target first data corresponding to each target processing unit among the M first data to the target processing unit; receiving second data from each target processing unit; and determining each second data and the other first data except the target first data among the M first data as the second exchange result; where M is an integer greater than 1, and M is the total number of processing units in the computing node cluster; and Training the deep learning model by using the second exchange result.
2. The method according to claim 1, wherein, The performing data exchange with a target computing node in the computing node cluster where the current computing node is located according to the N first data sets to obtain a first exchange result further includes: Determining the corresponding processing unit in each target computing node; and Determining each second data set and the other first data sets except the target first data set among the N first data sets as the first exchange result.
3. The method according to claim 2, wherein The determining the corresponding processing unit in each target computing node includes: Obtaining a first unit number of the current processing unit in the current computing node; and Determining the processing unit in the target computing node whose unit number matches the first unit number as the corresponding processing unit.
4. The method according to claim 2 or 3, wherein, The sending the target first data set corresponding to each corresponding processing unit among the N first data sets to the corresponding processing unit includes: For each target computing node, Obtaining a node number i of the target computing node, where i is an integer and 1 ≤ i ≤ N; and Determining the i-th first data set among the N first data sets as the target first data set, and sending the target first data set to the corresponding processing unit in the target computing node.
5. The method according to claim 1, wherein, The sending the target first data corresponding to each target processing unit among the M first data to the target processing unit includes: Obtaining the total number M' of processing units in the current computing node; For the j-th first data among the M first data, where j is an integer and 1 ≤ j ≤ M, determine an intermediate parameter y according to the following formula: y = (j - 1) % M' + 1, Send the j-th first data to the y-th processing unit in the current computing node.
6. The method according to claim 1, wherein, The training of the deep learning model using the second exchange result includes: Input the second exchange result into the deep learning model to obtain an output result; and Adjust the parameters of the deep learning model according to the output result.
7. An apparatus for training a deep learning model, which is applied to a current processing unit deployed in a computing node cluster. The computing node cluster includes a plurality of computing nodes, and each computing node includes a plurality of processing units. The computing node where the current processing unit is located is the current computing node; the apparatus includes: A partitioning module, configured to partition the training data in the current processing unit into N first data sets, where N is an integer greater than 1, and N is the number of the plurality of computing nodes; A first exchange module, configured to perform data exchange with a target computing node according to the N first data sets to obtain a first exchange result, including: sending a target first data set among the N first data sets to a corresponding processing unit in the target computing node; receiving second data sets from each corresponding processing unit; the target computing node includes other computing nodes except the current computing node among the plurality of computing nodes; A second exchange module, configured to perform data exchange with a target processing unit in the current computing node according to the first exchange result to obtain a second exchange result, including: dividing the first exchange result into M first data; sending the target first data corresponding to each target processing unit among the M first data to the target processing unit; receiving second data from each target processing unit; and determining each second data and the other first data among the M first data except the target first data as the second exchange result; where M is an integer greater than 1, and M is the total number of processing units in the computing node cluster; and A training module, configured to train a deep learning model using the second exchange result.
8. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; where The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-6.
9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, the steps of the method according to any one of claims 1-6 are implemented.