Neural network training method and device
During the training process of the deep neural network, the control circuit determines the first processing layer combination and instructs the training circuit to save only the feature values of the layer, which solves the problem of insufficient GPU video memory, reduces the computational amount of reverse training, and improves training efficiency.
Patent Information
- Application Number
- CN202010191111.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-03-18
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-03-18
AI Technical Summary
During the training process of deep neural networks, due to insufficient GPU video memory, the feature value storage requirements cannot be met, resulting in the training task being unable to be completed, and the time-consuming reverse training is increased and training efficiency is reduced.
The first processing layer combination is determined by the control circuit and the instruction information is sent to the training circuit, so that the training circuit can save only the feature value of the first processing layer during the forward training process, and recalculate the feature value of the other layers through the feature value to complete the reverse training.
It effectively utilizes GPU video memory, avoids idle and waste of video memory, reduces the amount of computation during reverse training, and improves the training efficiency of neural networks.
Smart Images

Figure CN113496267B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of neural networks, and in particular to a neural network training method and device. Background Art
[0002] When a neural network contains a large number of layers, it can also be called a deep neural network (DNN). Usually, multiple graphics processing units (GPUs) are used to train the DNN model in parallel to improve training efficiency. The training process of the DNN model can be divided into forward training and reverse training. Specifically, forward training is performed first, and the eigenvalues of each layer are calculated layer by layer from the first layer to the last layer and saved in the GPU memory. Then, using the eigenvalues of each layer, reverse training is performed layer by layer from the last layer to the first layer to obtain the updated weights of each layer. When the reverse training of a layer is completed, the eigenvalue of the layer is released until the reverse training is completed, thereby completing a training process of the DNN model.
[0003] When a GPU is responsible for calculating the number of layers of a DNN model, the more eigenvalues it needs to save during the forward training process. When the GPU's video memory size cannot meet the storage requirements of the eigenvalues, the training task cannot be completed. To this end, a GPU video memory optimization scheme is proposed to reduce the amount of GPU video memory occupied during the entire DNN model training process. Specifically, during the forward training process, only the eigenvalues of the first layer that the GPU is responsible for are saved, and during the reverse training process, the eigenvalues of the first layer are used to recalculate the eigenvalues of the other layers that the GPU is responsible for, such as the second to the last layer, and the eigenvalues of the first layer and the recalculated eigenvalues of the other layers are used to complete the reverse training, and the eigenvalues of the first layer are released after the reverse training is completed.
[0004] However, the above solution of only saving the eigenvalues of the first layer that the GPU is responsible for, and using the eigenvalues of the first layer to recalculate the eigenvalues of other layers that the GPU is responsible for to complete the reverse training requires recalculating the eigenvalues of all layers except the first layer, which increases the time consumption of reverse training, thereby reducing the training efficiency of the DNN model, and may cause a large amount of GPU video memory to be idle and wasteful. Summary of the invention
[0005] The present application provides a neural network training method and device, which can solve the problem of a large amount of GPU video memory being idle and wasted, and can reduce the amount of calculation required to recalculate eigenvalues during reverse training, thereby improving the training efficiency of the neural network.
[0006] In a first aspect, a training method for a neural network is provided. The training method for the neural network includes: a control circuit determines a first processing layer combination. The first processing layer combination includes multiple processing layers, and the multiple processing layers are selected from the processing layers allocated to the training circuit, and the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process. Then, the control circuit sends first information to the training circuit. The first information is used to instruct the training circuit to store the characteristic values corresponding to the first processing layer combination during the forward training process.
[0007] Based on the training method of the neural network provided in the first aspect, the control circuit selects a first processing layer combination from the processing layers allocated to the training circuit according to the storage capacity of the training circuit and the storage requirements of the eigenvalues corresponding to each processing layer, so that the storage capacity of the training circuit meets the storage requirements of the eigenvalues corresponding to the first processing layer combination during the training process, and then sends first information to the training circuit to instruct the training circuit to store the eigenvalues corresponding to the first processing layer combination during the forward training process. This can solve the problem of waste caused by a large amount of idle video memory of the training circuit while ensuring the normal execution of the training task of the neural network, and can reduce the time spent on reverse calculations, thereby improving the training efficiency of the neural network.
[0008] In a possible design scheme, the control circuit determines the first processing layer combination, including: the control circuit determines multiple processing layer combinations. Wherein, the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to any one of the multiple processing layer combinations during the training process. Then, the control circuit determines one or more processing layer combinations with the largest number of processing layers contained in the multiple processing layer combinations as the second processing layer combination. Afterwards, the control circuit selects the first processing layer combination from the second processing layer combination. For example, the control circuit may be a central processing unit (CPU), and the training circuit may be a GPU.
[0009] In this way, in the process of completing the training task normally, the more eigenvalues of the processing layers that are saved, the more effectively the storage capacity of the GPU can be utilized. Therefore, selecting the processing layer combination with the largest number of processing layers as the first processing layer combination can solve the problem of waste caused by a large amount of idle storage space of the GPU, and can reduce the amount of calculation of recalculating eigenvalues during reverse training, thereby further improving training efficiency.
[0010] In a possible design, the number of processing layers included in the first processing layer combination is: Where N is the number of processing layers in the first processing layer combination, C MAX is the storage capacity of the training circuit, C1 is the first reserved capacity, and C is the characteristic value storage requirement of any processing layer allocated to the training circuit. To round down.
[0011] If the storage requirements for the characteristic values of each processing layer in the processing layers allocated to the training circuit are equal, the number of processing layers included in the first processing layer combination can be calculated according to the above formula, and then the first processing layer combination can be determined from the processing layer combinations with N processing layer numbers, thereby eliminating the tedious steps of determining multiple processing layer combinations and determining one or more processing layer combinations with the largest number of processing layer layers among the multiple processing layer combinations as the second processing layer combination. The calculation is simpler, which can improve the efficiency of determining the optimal first processing layer combination, thereby further improving the training efficiency.
[0012] In a possible design scheme, the control circuit selects the first processing layer combination from the second processing layer combination, which may include: the control circuit determines the processing layer combination with the shortest first training time in the second processing layer combination as the first processing layer combination. The first training time includes the eigenvalue recalculation time, or the first training time includes the forward training time, the reverse training time and the eigenvalue recalculation time. In this way, the processing layer combination with the shortest first training time is selected from the processing layer combination including the largest number of processing layers, which not only reduces the waste caused by the idleness of a large amount of GPU video memory during the training process of the neural network, but also can minimize the time consumption of eigenvalue recalculation, further improving the training efficiency of the neural network.
[0013] In a possible design scheme, if the training duration of each processing layer allocated to the training circuit is equal, the training method provided in the first aspect may further include: the control circuit calculates the first step length and the second step length according to the following formula: Wherein, step1 is the first step length, step2 is the second step length, M is the total number of processing layers allocated to the training circuit, N is the number of processing layers in the first processing layer combination, Then, the control circuit selects the processing layer combination with the shortest first training time from the processing layer combinations corresponding to the first step length and the second step length as the first processing layer combination. Alternatively, the control circuit selects the processing layer combination with the smallest number of eigenvalue recalculations from the processing layer combinations corresponding to the first step length and the second step length as the first processing layer combination.
[0014] In this way, when the training duration of each processing layer allocated to the training circuit is equal, the above formula can be used to calculate the first step length and the second step length, and then only the first training duration or the number of eigenvalue recalculations of the processing layer combination corresponding to the first step length and the second step length need to be calculated to determine the optimal first processing layer combination. There is no need to traverse all second processing layer combinations, which improves the efficiency of determining the optimal first processing layer combination, thereby further improving training efficiency.
[0015] Optionally, the storage capacity of the above-mentioned training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process, and may include: the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process and the first reserved capacity; wherein, the first reserved capacity is used to store the intermediate data generated during the training process.
[0016] It should be noted that, in addition to the characteristic value of the first processing layer combination and the first reserved capacity, the storage capacity of the training circuit also needs to meet the storage requirements of the read-in training samples. Therefore, optionally, a reserved storage space can also be set for the above-mentioned read-in training samples. Among them, the amount of reserved storage space can be set according to actual needs or actual experience, and the embodiments of the present application do not specifically limit this.
[0017] In a second aspect, a neural network training method is provided. The neural network training method includes: a training circuit receives first information sent by a control circuit. The first information is used to instruct the training circuit to store a feature value corresponding to a first processing layer combination during forward training, the first processing layer combination includes multiple processing layers, and the multiple processing layers are selected from the processing layers allocated to the training circuit, and the storage capacity of the training circuit meets the storage requirements of the feature value corresponding to the first processing layer combination during training. Then, the training circuit performs reverse training according to the feature value corresponding to the first processing layer combination.
[0018] In a possible design scheme, the training circuit performs reverse training according to the characteristic value corresponding to the first processing layer combination, including: if the first reverse training layer belongs to the first processing layer combination, the training circuit performs reverse training of the first reverse training layer based on the characteristic value of the first reverse training layer. If the first reverse training layer does not belong to the first processing layer combination, the training circuit determines the first reference layer of the first reverse training layer, and performs reverse training of the first reverse training layer based on the characteristic value of the first reference layer. Among them, the first reverse training layer is any processing layer among the processing layers assigned to the training circuit; the first reference layer is a processing layer in the first processing layer combination whose forward training order is before the first reverse training layer and whose number of processing layers separated from the first reverse training layer is the smallest.
[0019] In addition, the technical effects of the training method provided in the second aspect can refer to the technical effects of the neural network training method described in any implementation of the first aspect, and will not be repeated here.
[0020] In a third aspect, a control circuit of a neural network is provided. The control circuit includes: a processor and a transmitter. The processor is used to determine a first processing layer combination. The first processing layer combination includes multiple processing layers, and the multiple processing layers are selected from the processing layers allocated to the training circuit. The storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process. The transmitter is used to send first information to the training circuit. The first information is used to instruct the training circuit to store the characteristic values corresponding to the first processing layer combination during the forward training process.
[0021] In a possible design, the processor is further used to determine a plurality of processing layer combinations. The storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to any of the plurality of processing layer combinations during the training process. The processor is further used to determine one or more processing layer combinations with the largest number of processing layer layers included in the plurality of processing layer combinations as the second processing layer combination. The processor is further used to select the first processing layer combination from the second processing layer combination.
[0022] In a possible design, the number of processing layers included in the first processing layer combination is: Where N is the number of processing layers in the first processing layer combination, C MAX is the storage capacity of the training circuit, C1 is the first reserved capacity, and C is the characteristic value storage requirement of any processing layer allocated to the training circuit. To round down.
[0023] In a possible design, the processor is further configured to determine the processing layer combination with the shortest first training duration in the second processing layer combination as the first processing layer combination, wherein the first training duration includes the eigenvalue recalculation time, or the first training duration includes the forward training time, the reverse training time, and the eigenvalue recalculation time.
[0024] Optionally, if the training duration of each processing layer allocated to the training circuit is equal, the processor is further configured to calculate the first step length and the second step length according to the following formula: Wherein, step1 is the first step length, step2 is the second step length, M is the total number of processing layers allocated to the training circuit, N is the number of processing layers in the first processing layer combination, The processor is further used to select, from the processing layer combinations corresponding to the first step length and the second step length, a processing layer combination with the shortest first training time as the first processing layer combination, or the processor is further used to select, from the processing layer combinations corresponding to the first step length and the second step length, a processing layer combination with the smallest number of eigenvalue recalculations as the first processing layer combination.
[0025] Optionally, the storage capacity of the above-mentioned training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process, and may include: the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process and the first reserved capacity; wherein, the first reserved capacity is used to store the intermediate data generated during the training process.
[0026] It should be noted that, in addition to the characteristic value of the first processing layer combination and the first reserved capacity, the storage capacity of the training circuit also needs to meet the storage requirements of the read-in training samples. Therefore, optionally, a reserved storage space can also be set for the above-mentioned read-in training samples. Among them, the amount of reserved storage space can be set according to actual needs or actual experience, and the embodiments of the present application do not specifically limit this.
[0027] Optionally, the control circuit described in the third aspect may further include a receiver. The receiver is used to receive data sent by the training circuit. Further, the receiver and the transmitter may be provided separately or integrated together, i.e., a transceiver. The present application does not specifically limit the specific implementation of the receiver and the transmitter.
[0028] Optionally, the control circuit described in the third aspect may further include a memory storing a program or an instruction. When the processor executes the program or the instruction, the control circuit described in the third aspect may execute the method described in the first aspect.
[0029] It should be noted that the control circuit described in the third aspect may be a CPU, or may be or include a device of a CPU, and this application does not limit this.
[0030] In addition, the technical effects of the control circuit of the neural network provided in the third aspect can refer to the technical effects of the training method of the neural network described in any implementation method of the first aspect, and will not be repeated here.
[0031] In a fourth aspect, a training circuit is provided. The training circuit includes: a processor, a receiver and a memory. The receiver is used to receive the first information sent by the control circuit. The first information is used to instruct the processor to store the characteristic value corresponding to the first processing layer combination during the forward training process, the first processing layer combination includes multiple processing layers, and the multiple processing layers are selected from the processing layers allocated to the training circuit, and the storage capacity of the memory meets the storage requirements of the characteristic value corresponding to the first processing layer combination during the training process. The processor is used to perform reverse training according to the characteristic value corresponding to the first processing layer combination.
[0032] In a possible design, the processor is further configured to perform reverse training of the first reverse training layer based on the characteristic value of the first reverse training layer if the first reverse training layer belongs to the first processing layer combination. The processor is further configured to determine a first reference layer of the first reverse training layer if the first reverse training layer does not belong to the first processing layer combination, and perform reverse training of the first reverse training layer based on the characteristic value of the first reference layer. The first reverse training layer is any processing layer among the processing layers allocated to the training circuit, and the first reference layer is a processing layer in the first processing layer combination whose forward training order precedes the first reverse training layer and whose number of processing layers separated from the first reverse training layer is the smallest.
[0033] Optionally, the training circuit described in the fourth aspect may further include a transmitter. The transmitter is used to send data to the control circuit. Further, the receiver and the transmitter may be provided separately or integrated together, i.e., a transceiver. The present application does not specifically limit the specific implementation of the receiver and the transmitter.
[0034] Optionally, the memory of the training circuit described in the fourth aspect may also store a program or instruction. When the processor executes the program or instruction, the training circuit described in the fourth aspect may execute the method described in the second aspect.
[0035] It should be noted that the training circuit described in the fourth aspect may be a GPU, or may be or include a device of a GPU, or other acceleration device, and this application does not limit this.
[0036] In addition, the technical effects of the training circuit provided in the fourth aspect can refer to the technical effects of the neural network training method described in any implementation method of the first aspect, and will not be repeated here.
[0037] In a fifth aspect, a training device is provided. The training device includes: a processor coupled to a memory. The memory is used to store a computer program. The processor is used to execute the computer program stored in the memory so that the training device performs the neural network training method described in any possible implementation of the first aspect.
[0038] In a possible design solution, the training device described in the fifth aspect may further include a transceiver. The transceiver may be a transceiver circuit or an input / output interface. The transceiver may be used for the training device to communicate with other training devices.
[0039] In the present application, the training device described in the fifth aspect may be a control circuit, such as a CPU.
[0040] The technical effects of the training device described in the fifth aspect can refer to the technical effects of the neural network training method described in any one of the implementations of the first aspect, and will not be repeated here.
[0041] In a sixth aspect, a training device is provided. The training device includes: a processor coupled to a memory. The memory is used to store a computer program. The processor is used to execute the computer program stored in the memory so that the training device performs the neural network training method described in any possible implementation of the second aspect.
[0042] In a possible design solution, the training device described in the sixth aspect may further include a transceiver. The transceiver may be a transceiver circuit or an input / output interface. The transceiver may be used for the training device to communicate with other training devices.
[0043] In the present application, the training device described in the sixth aspect may be a training circuit, such as a GPU.
[0044] The technical effects of the training device described in the sixth aspect can refer to the technical effects of the neural network training method described in any one of the implementations of the first aspect, and will not be repeated here.
[0045] In a seventh aspect, a training device is provided, which includes the control circuit described in any possible implementation of the third aspect, and one or more training circuits described in any possible implementation of the fourth aspect.
[0046] In a possible design solution, the control circuit may be a central processing unit (CPU), and the training circuit may be a graphics processing unit (GPU).
[0047] The technical effects of the training device described in the seventh aspect can refer to the technical effects of the neural network training method described in any one of the implementations of the first aspect, and will not be repeated here.
[0048] In an eighth aspect, a computer-readable storage medium is provided, comprising: the computer-readable storage medium includes a program or an instruction, and when the program or the instruction is executed on a computer, the computer executes the training method of a neural network described in any possible implementation of the first aspect to the second aspect. It should be noted that the computer-readable storage medium may be a non-transitory computer-readable storage medium.
[0049] In a ninth aspect, a computer program product is provided, the computer program product comprising: a computer program code, when the computer program code is run on a computer, the computer executes the neural network training method described in any possible implementation of the first aspect to the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 A schematic diagram of the architecture of a training system provided in an embodiment of the present application;
[0051] Figure 2 A schematic diagram of the structure of a neural network training device provided in an embodiment of the present application Figure 1 ;
[0052] Figure 3 A flowchart of a neural network training method provided in an embodiment of the present application;
[0053] Figure 4 A schematic diagram of parallel training of a neural network model provided in an embodiment of the present application;
[0054] Figure 5 A schematic diagram of data parallel training of a neural network provided in an embodiment of the present application;
[0055] Figure 6 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 1 ;
[0056] Figure 7 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 2 ;
[0057] Figure 8 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 3 ;
[0058] Fig. 9 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 4 ;
[0059] Fig.10 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 5 ;
[0060] Fig.11 A schematic diagram of the structure of a control circuit of a neural network provided in an embodiment of the present application;
[0061] Fig.12 A schematic diagram of the structure of a neural network training circuit provided in an embodiment of the present application;
[0062] Fig.13 A schematic diagram of the structure of a neural network training device provided in an embodiment of the present application Figure 2 . DETAILED DESCRIPTION
[0063] The technical solution in this application will be described below in conjunction with the accompanying drawings.
[0064] The technical solutions of the embodiments of the present application can be applied to the distributed training system of deep neural networks, but are not limited thereto. The present application will present various aspects, embodiments or features around a system that may include multiple devices, components, modules, etc. It should be understood and appreciated that each system may include additional devices, components, modules, etc., and / or may not include all devices, components, modules, etc. discussed in conjunction with the accompanying drawings. In addition, a combination of these solutions may also be used.
[0065] The system architecture described in the embodiments of the present application is intended to more clearly illustrate the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided in the embodiments of the present application. A person of ordinary skill in the art can appreciate that, with the evolution of the system architecture, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0066] Figure 1 A schematic diagram of the architecture of a training system applicable to the neural network training method provided in the embodiment of the present application. Figure 1 The training system applicable to the embodiment of the present application is described in detail with the training system architecture diagram shown in FIG. Figure 1 As shown, the training system includes a control circuit and at least one training circuit. The training circuit is interconnected with the control circuit, and the neural network training software can be deployed in the control circuit. The control circuit generates and sends a calculation graph and sample data to the training circuit. The training circuit performs calculations according to the sample data corresponding to the calculation graph and feeds back the calculation results to the control circuit.
[0067] It should be noted that the control circuit and the training circuit may be different hardware devices. For example, the control circuit may be a central processing unit (CPU) or a device including a CPU, and the training circuit may be a GPU or other acceleration device or a device including a GPU.
[0068] In addition, the control circuit and the training circuit may also be in the same hardware device, for example, a server including a CPU and a GPU.
[0069] In the embodiment of the present application, a neural network is an algorithmic network capable of learning, summarizing and generalizing, and can be built into a computing node in the form of neural network software or hardware, such as a neural network training program, an executable script, etc. The neural network can learn and summarize through the experimental application of training data to improve the recognition ability of the neural network. Generally speaking, a DNN model consists of multiple layers of neurons (operators), each layer has multiple inputs and outputs, and the input or output is a multidimensional array, also known as a tensor. Each layer has one or more weighted values, called weights. The output result of a certain layer, also called an eigenvalue, is equal to the result of mathematical operations such as the addition or multiplication of the input and weight of the layer, usually involving matrix multiplication operations, that is, involving a large number of multiplication and accumulation operations.
[0070] In the embodiment of the present application, the process of creating a model by learning the weights of each layer of the DNN model through training samples is called a training process. After executing a training, the process of correcting the weights of each layer of the DNN model is called completing a training iteration. The training method of the neural network provided in the embodiment of the present application is applicable to a training process and a training iteration process. For the convenience of description, the present application uniformly describes the training process and a training iteration process as a training process.
[0071] The above training process may include forward training and reverse training. During the forward training process, the accelerator processes the input training data and the weight matrix of each layer layer by layer, and finally obtains the eigenvalue corresponding to the training data and saves the eigenvalue. For example, the eigenvalue of the first intermediate layer is calculated using the training data provided by the input layer and the weight matrix of the first intermediate layer, and the eigenvalue of the first intermediate layer is saved. Then, the eigenvalue of the first intermediate layer is used as the training data of the second intermediate layer, and the eigenvalue of the second intermediate layer is calculated and saved using the weight matrix of the second intermediate layer as the training data of the third intermediate layer. Similarly, the eigenvalue of the last intermediate layer is used as the training data, and the eigenvalue of the output layer is calculated using the weight matrix of the output layer, that is, the recognition result of the entire DNN model. Then, the deviation between the recognition result and the artificial label is used as the feedback error of the output layer, that is, the feedback error of the entire DNN model, and reverse training is started.
[0072] In the reverse training process, reverse training is performed layer by layer from the output layer to the input layer. The eigenvalues of each layer in the forward training are used to reverse train each layer, and the weight matrix of each layer is updated. After each reverse training of a layer is completed, the GPU memory occupied by the eigenvalues of the corresponding layer is released. For example, the eigenvalues of the output layer are used to calculate the new weight matrix of the output layer, and the eigenvalues of the output layer are released. Then, the eigenvalues of the penultimate intermediate layer are used to calculate the new weight matrix of the penultimate intermediate layer, and the eigenvalues of the penultimate intermediate layer are released. This process is repeated until a new weight matrix for each layer of the DNN model is obtained.
[0073] As can be seen from the above process, the storage of eigenvalues is usually involved in the training process of the DNN model. As the scale of the neural network becomes larger and more complex, more and more eigenvalues need to be stored during the training process, which will cause the GPU's video memory capacity to be unable to meet the storage requirements of the eigenvalues. In order to reduce the amount of GPU video memory occupied during the DNN model training process, a solution is proposed that only saves the eigenvalues of the first layer that the accelerator is responsible for, and uses the eigenvalues of the first layer to recalculate the eigenvalues of the other layers that the accelerator is responsible for to complete the reverse training. However, this solution requires recalculating the eigenvalues of all layers except the first layer, which increases the time consumption of reverse training, thereby reducing the training efficiency of the DNN model, and may still cause a large amount of accelerator video memory to be idle and wasteful.
[0074] In the embodiments of the present application, "example" and "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "example" or "for example" in the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the word "example" is used to present concepts in a concrete way.
[0075] In the embodiments of the present application, "of", "corresponding, relevant" and "corresponding" may sometimes be used interchangeably. It should be noted that when the distinction between them is not emphasized, the meanings they intend to express are similar or consistent.
[0076] It should be understood that Figure 1 The simplified schematic diagram is only for the sake of understanding. The training system may also include other devices. Figure 1 Not drawn in.
[0077] Figure 2 1 is a schematic diagram of a training device 200 that can be used to perform the neural network training method provided in an embodiment of the present application. The training device 200 can be a control circuit or a training circuit. The training device 200 can also be a training device including a control circuit and one or more training circuits. Figure 2As shown, the training device 200 includes one or more processors, such as processor 201 and / or processor 207, at least one communication interface, such as communication interface 204, and communication line 202. Optionally, the training device 200 may also include a memory 203. The processor 201 is used as an example for description below.
[0078] Processor 201 can be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), an FPGA (Field Programmable Gate Array), or one or more integrated circuits that integrate multiple processing circuit functions (such as CPU+ASIC).
[0079] The communication link 202 may include one or more pathways for connecting different components.
[0080] The communication interface 204 may be a transceiver circuit for communicating with other devices or communication networks, such as a cloud computing network, Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. For example, the transceiver circuit may be a device such as a transceiver or a transceiver. Optionally, the communication interface 204 may also be an input / output (I / O) circuit of the processor 201, for implementing signal input and signal output of the processor 201.
[0081] The memory 203 may be a device with a storage function. For example, it may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 203 may exist independently and be connected to the processor 201 via the communication line 202. Of course, the memory 203 may also be integrated with the processor 201.
[0082] The memory 203 is used to store computer execution instructions for executing the solution of the present application, and the execution is controlled by the processor 201. The processor 201 is used to read and execute computer instructions (such as for a CPU) or configuration files (such as for an FPGA) stored in the memory 203, thereby implementing the neural network training method provided in the embodiment of the present application.
[0083] Alternatively, optionally, in an embodiment of the present application, the processor 201 may also execute relevant processing functions in the neural network training method provided in the following embodiments of the present application, and the communication interface 204 is responsible for communicating with other devices or communication networks, which is not specifically limited in the embodiment of the present application.
[0084] Optionally, the computer-executable instructions in the embodiments of the present application may also be referred to as application code, which is not specifically limited in the embodiments of the present application.
[0085] In a specific implementation, as an embodiment, the processor 201 may include one or more CPUs, such as Figure 2 CPU0 and CPU1 in.
[0086] In a specific implementation, as an embodiment, the training device 200 may also include multiple processors, such as Figure 2201 and processor 207 in the embodiment. Each of these processors may be a single-CPU processor, a multi-CPU processor, or a graphics (GPU) processor. The processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).
[0087] In a specific implementation, as an embodiment, the training device 200 may also include an output device 205 and an input device 206. The output device 205 communicates with the processor 201 and may output information in a variety of ways. For example, the output device 205 may be a touch screen, a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, a projector, or a printer. The input device 206 communicates with the processor 201 and may receive user input in a variety of ways. For example, the input device 206 may be a mouse, a keyboard, a touch screen device, or a sensor device.
[0088] The training device 200 may also be referred to as a training device, which may be a general device or a dedicated device. For example, the training device may be a client, a desktop computer, a portable computer, a network server, a personal digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, an embedded device, or a computer having Figure 2 Of course, the training device 200 may also be a software and / or hardware entity disposed inside each of the above-mentioned single devices, such as a chip or chip system for executing the training task provided in the embodiment of the present application. The embodiment of the present application does not limit the type of the training device 200.
[0089] It should be understood that Figure 2 This is a simplified schematic diagram for ease of understanding only. The neural network training device may also include other components, circuits or devices. Figure 2 They are not drawn in the picture.
[0090] The following will be combined Figure 3-Figure 10 The training method of the neural network provided in the embodiment of the present application is specifically described. The method can be applied to Figure 1 In the neural network training system shown in , the communication between the control circuit and the training circuit is carried out to complete the training task.
[0091] like Figure 3As shown, the training method of the neural network includes:
[0092] S301, the control circuit determines a first processing layer combination.
[0093] Exemplarily, the control circuit can be a processor with powerful computing power such as a CPU that can deploy neural network training software, generate and send calculation graphs and sample data, and the following training circuit can be an accelerator, or a GPU in an accelerator. For the sake of simplicity, the embodiments of the present application are described using CPU and GPU as examples.
[0094] Optionally, the neural network training method provided in the embodiment of the present application can be applied to training methods such as model parallelism, data parallelism, and hybrid parallelism. The CPU can divide the processing layer of the DNN model that each GPU is responsible for according to different training modes. The following is a detailed introduction to model parallelism, data parallelism, and hybrid parallelism.
[0095] Model parallel training method, that is, multiple GPUs jointly complete the training calculation of a DNN model. Each GPU is responsible for the training calculation of a part of the processing layer of the DNN model and is responsible for the training calculation of all samples.
[0096] For example, Figure 4 A schematic diagram of parallel training of a neural network model provided in an embodiment of the present application. Figure 4 As shown in the figure, the DNN model includes 32 processing layers. The size and computing time of each processing layer are equal. The number of samples is 1024. Four GPUs are responsible for training. Each GPU is responsible for the training calculation of 1024 samples in 8 processing layers on average.
[0097] Data parallel training method, that is, multiple GPUs jointly complete the training calculations of all samples in a DNN model. Each GPU is responsible for the training calculations of all processing layers of the DNN model and is responsible for the training calculations of some samples.
[0098] Figure 5 A schematic diagram of data parallel training of a neural network provided in an embodiment of the present application. Figure 5 As shown in the figure, the DNN model includes 32 processing layers. The size and computing time of each processing layer are equal. The number of samples is 1024. Four GPUs are responsible for training. Each GPU is responsible for the training calculation of 256 samples in 32 processing layers on average.
[0099] Hybrid parallel training mode, that is, a mixture of model parallelism and data parallelism. For example, the DNN model includes 32 processing layers, the number of samples is 1024, and 4 GPUs are responsible for training. Each GPU is responsible for running the training calculation of 256 samples in 8 processing layers, or two of the GPUs are responsible for the training calculation of 512 samples, and the other two GPUs are responsible for the training calculation of 16 processing layers.
[0100] Of course, the basic principle of allocating GPUs is to make the computing time of each GPU equal or close as much as possible. If the training time of each processing layer of the DNN model is different, the number of processing layers responsible for each GPU can also be different. This application does not limit how to allocate processing layers to GPUs.
[0101] The first processing layer combination includes multiple processing layers, which are selected from the processing layers allocated to the training circuit, and the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process.
[0102] It should be noted that the first processing layer combination is a combination that meets the above-mentioned condition of "the storage capacity of the training circuit meets the storage requirements of the eigenvalues corresponding to the first processing layer combination during training", that is, the storage requirements of the eigenvalues corresponding to the first processing layer combination are smaller than the storage capacity of the training circuit, thereby ensuring that the storage capacity of the training circuit can meet the storage requirements of the eigenvalues corresponding to the first processing layer combination and complete the training task of the neural network.
[0103] Optionally, the storage requirement of the characteristic value corresponding to the first processing layer combination includes the storage requirement of the characteristic value of each processing layer in the first processing layer combination.
[0104] The embodiment provided in the present application is described by taking an example in which a GPU is responsible for training calculations of 8 processing layers and the storage capacity of the GPU that can be used for storing forward training feature values is 19 gigabytes (GB).
[0105] Optionally, the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process, which may include: the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process and the first reserved capacity.
[0106] The first reserved capacity is used to store intermediate data generated during the training process.
[0107] Figure 6 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 1 .like Figure 6As shown, the storage capacity of the training circuit is 19GB, the storage requirement of the eigenvalues of (layer, L)1 to L6 is 4GB, the storage requirement of the eigenvalues of the L7 layer is 5GB, the storage requirement of the eigenvalues of the L8 layer is 5GB, and the first processing layer combination is {L1, L4, L7}.
[0108] Taking the reverse training of L8 layer as an example, the eigenvalues of L8 layer are first calculated using the saved eigenvalues of L7 layer, and then the eigenvalues of L8 layer are used to reversely train L8 layer. In this training process, the generated eigenvalues of L8 layer occupy 5BG of storage space. Therefore, the storage capacity of the training circuit must not only meet the storage requirements of the eigenvalues of L1, L4 and L7 layers, but also meet the storage requirements of the eigenvalues of L8 layer to ensure the smooth progress of the training task.
[0109] Taking the reverse training of the L6 layer as an example, the eigenvalues of the L5 layer are first calculated using the saved eigenvalues of the L4 layer. In this process, the eigenvalues of the L5 layer are intermediate data generated during the training process, and the eigenvalues of the L5 layer occupy 4BG of storage space. Therefore, the storage capacity of the training circuit must not only meet the storage requirements of the eigenvalues of the L1, L4, and L7 layers, but also meet the storage requirements of the eigenvalues of the L5 layer. Then, the eigenvalues of the L6 layer are calculated using the eigenvalues of the L5 layer, and the L6 layer is reversely trained using the eigenvalues of the L6 layer. In this process, the eigenvalues of the L6 layer are intermediate data generated during the training process, and the eigenvalues of the L6 layer occupy 4BG of storage space. The storage capacity of the training circuit must not only meet the storage requirements of the eigenvalues of the L1, L4, and L7 layers, but also meet the storage requirements of the eigenvalues of the L6 layer to ensure the smooth progress of the training task.
[0110] It should be noted that, in addition to the storage requirements of the characteristic values of the first processing layer combination and the first reserved capacity, the storage capacity of the training circuit also needs to meet the storage requirements of the read-in training samples. Therefore, optionally, a reserved storage space can also be set for the above-mentioned read-in training samples. Among them, the amount of reserved storage space can be set according to actual needs or actual experience, and the embodiments of the present application do not specifically limit this.
[0111] In a possible design method, the above S301, the control circuit determines the first processing layer combination, which may include:
[0112] In step 1, the control circuit determines a plurality of processing layer combinations.
[0113] The storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to any one of the multiple processing layer combinations during the training process.
[0114] For the convenience of description, the embodiment of the present application refers to "the storage capacity of the training circuit meeting the storage requirements of the characteristic values corresponding to any processing layer combination in the multiple processing layer combinations during the training process" as "capacity condition".
[0115] Combination Figure 6 , the CPU traverses the L1 layer to the L8 layer, and determines multiple processing layer combinations according to the storage capacity of the training circuit and the storage requirements of the characteristic values of each processing layer. For example, a processing layer combination including 2 processing layers: {L1, L2}, {L1, L3}, {L1, L4}, {L1, L8}, etc., a processing layer combination including 3 processing layers: {L1, L2, L3}, {L1, L2, L4}, {L1, L3, L5}, {L1, L3, L8}, {L1, L4, L7}, etc., which are not listed one by one in the embodiments of the present application.
[0116] Furthermore, combined with Figure 6 , assuming that the first reserved capacity is 5GB, the multiple processing layer combinations determined from the L1 layer to the L8 layer do not include a processing layer combination including 4 processing layers. Specifically, the storage capacity of the training circuit is 19GB, the storage requirement of the processing layer combination {L1, L2, L3} including 3 processing layers is 4GB+4GB+4GB=12GB, the first reserved capacity is 5GB, and the sum of the characteristic value storage requirement of the processing layer combination {L1, L2, L3} and the first reserved capacity is 12GB+5GB=17GB, which is less than the storage capacity of the training circuit of 19GB, which can ensure the normal progress of the training task. If another processing layer, such as L5, is added to the processing layer combination {L1, L2, L3}, the storage requirement of the processing layer combination {L1, L2, L3, L5} is 16GB, the first reserved capacity is 5GB, and the sum of the characteristic value storage requirement and the first reserved capacity of the processing layer combination {L1, L2, L3, L5} is 16GB+5GB=21GB. 21GB is greater than the storage capacity of the training circuit 19GB, and the training task cannot be carried out normally. Similarly, after traversing various combinations of L1 to L8, it is concluded that the number of processing layers contained in multiple layer combinations is at most 3 layers.
[0117] Step 2: The control circuit determines one or more processing layer combinations with the largest number of processing layers among the multiple processing layer combinations as the second processing layer combination.
[0118] That is to say, in the process of completing the training task normally, the more characteristic values of the processing layers are saved, the more effectively the GPU memory can be used, so the processing layer combination containing the largest number of processing layers is selected as the first processing layer combination to solve the problem of idle GPU capacity and waste. At this time, the processing layer combination that meets the conditions can be one or more.
[0119] For example, the second processing layer combination is the processing layer combination of 3 processing layers in the above step one: {L1, L2, L3}, {L1, L2, L4}, {L1, L3, L5}, {L1, L3, L8}, {L1, L4, L7}, etc., which are not listed one by one here.
[0120] Step three: the control circuit selects the first processing layer combination from the second processing layer combination.
[0121] Specifically, any one of the above-mentioned processing layer combinations including three processing layers can be selected as the first processing layer combination, such as the processing layer combination {L1, L4, L7}. When the second processing layer combination includes only one processing layer combination, the processing layer combination can be directly determined as the first processing layer combination.
[0122] Further, if the storage requirements of the characteristic values of each processing layer in the processing layers allocated to the training circuit are equal, the CPU can obtain the number of processing layers included in the first processing layer combination according to the following formula.
[0123]
[0124] In the above formula (1), N is the number of treatment layers in the first treatment layer combination, C MAX is the storage capacity of the training circuit, C1 is the first reserved capacity, C is the storage requirement of the feature value of any processing layer allocated to the training circuit, To round down.
[0125] Figure 7 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 2 . Combined Figure 7 , the storage capacity of the training circuit is 19GB. If the storage space occupied by the feature values of each layer is 4GB, the number of processing layers is
[0126] Then, the CPU may generate a second processing layer combination with N layers, and then select the first processing layer combination from the second processing layer combination.
[0127] Combination Figure 7 , the processing layer combinations with a number of layers of 3 include: {L1, L2, L3}, {L1, L2, L4}, {L1, L3, L5}, {L1, L3, L8}, {L1, L4, L7}, etc., which are not listed here one by one, and then any processing layer combination is selected as the first processing layer combination, for example, {L1, L4, L7}.
[0128] In one possible design, the control circuit selects the first processing layer combination from the second processing layer combination, which may include: the control circuit determines the processing layer combination with the shortest first training duration in the second processing layer combination as the first processing layer combination.
[0129] The first training duration includes the feature value recalculation time, or the first training duration includes the forward training time, the reverse training time and the feature value recalculation time.
[0130] Figure 8 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 3 The first training duration includes the time for recalculating the eigenvalues, combined with Figure 8 , the eigenvalue recalculation time for the processing layer combination {L1, L2, L3} is 18ms, and the specific calculation process is as follows:
[0131] L8 reverse training requires recalculating L8 eigenvalues from L3 eigenvalues, and the number of recalculation layers is 5, which takes 3*1ms+2*2ms=7ms;
[0132] L7 reverse calculation training requires recalculating L7 eigenvalues from L3 eigenvalues, and the number of recalculation layers is 4, which takes 3*1ms+2ms=5ms;
[0133] L6 reverse training requires recalculating L6 eigenvalues from L3 eigenvalues, and the number of recalculation layers is 3, which takes 3*1ms=3ms;
[0134] L5 reverse training requires recalculating L5 eigenvalues from L3 eigenvalues, and the number of recalculation layers is 2, which takes 2*1ms=2ms;
[0135] L4 reverse training requires recalculating L4 eigenvalues from L3 eigenvalues, and the number of recalculation layers is 1, which takes 1*1ms=1ms;
[0136] That is to say, the eigenvalue recalculation time for processing the layer combination {L1, L2, L3} is 7ms+5ms+3ms+2ms+1ms=18ms, that is, the first training duration is 18ms.
[0137] Similarly, all second processing layer combinations are traversed, the eigenvalue recalculation time of each processing layer combination is calculated, and the processing layer combination with the shortest eigenvalue recalculation time is determined as the first processing layer combination. The calculation process of the eigenvalue recalculation time of other second processing layer combinations is similar to the calculation process of the eigenvalue recalculation time of the above-mentioned processing layer combination {L1, L2, L3}, and will not be repeated in detail in the embodiments of the present application.
[0138] Alternatively, the first training duration includes forward training time, reverse training time, and eigenvalue recalculation time. All second processing layer combinations are traversed, and the total training time of each processing layer combination is calculated, that is, the sum of the forward training time, the reverse training time, and the eigenvalue recalculation time, and the processing layer combination with the shortest total training time is determined as the first processing layer combination.
[0139] Combination Figure 8 Taking the processing layer combination {L1, L2, L3} as an example, the forward training time from L1 to L8 is 6*1ms+2*2ms=10ms, the reverse training time from L1 to L8 is 6*1ms+2*2ms=10ms, and the eigenvalue recalculation time is 18ms. The total training time is 10ms+10ms+18ms=38ms, that is, the first training duration is 38ms.
[0140] Similarly, the calculation of the total training time of other second processing layer combinations can refer to the calculation process of the total training time of the above-mentioned processing layer combination {L1, L2, L3}, which will not be repeated in the present embodiment. Then, the total training time of all second processing layer combinations can be compared, and the processing layer combination with the shortest total training time can be determined as the first processing layer combination.
[0141] In this way, from the processing layer combinations that meet the above-mentioned capacity conditions and include the largest number of processing layers, the processing layer combination with the shortest first training time is selected, that is, the first processing layer combination. The selected first processing layer combination is the optimal processing layer combination, which can not only solve the problem of waste caused by a large amount of GPU video memory being idle during DNN model training, but also can minimize the time spent on recalculating the eigenvalues of the DNN model, thereby further improving the training efficiency of the neural network.
[0142] In one possible design approach, if the training time allocated to each processing layer of the training circuit is equal, Figure 3 The training method of the neural network shown may also include the following steps:
[0143] Step 4: The control circuit can calculate the first step length and the second step length according to the following formula.
[0144]
[0145]
[0146] In the above formula (2) and formula (3), step1 is the first step length, step2 is the second step length, M is the total number of processing layers allocated to the training circuit, N is the number of processing layers in the first processing layer combination, To round down.
[0147] Fig. 9 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 4 . Combined Fig. 9 , assuming that the forward training and reverse training time of each processing layer allocated to the training circuit is 1ms, then according to the above formula (2), the first step length is calculated as According to the above formula (3), the second step length is calculated as
[0148] Step 5: The control circuit selects the processing layer combination with the shortest first training time from the processing layer combinations corresponding to the first step length and the second step length as the first processing layer combination. Alternatively, the control circuit selects the processing layer combination with the shortest number of eigenvalue recalculations from the processing layer combinations corresponding to the first step length and the second step length as the first processing layer combination.
[0149] Exemplarily, the control circuit selects the processing layer combination with the shortest first training time from the processing layer combinations corresponding to the first step length and the second step length as the first processing layer combination, which can be specifically implemented as the following steps:
[0150] First, the processing layer combinations corresponding to the first step length and the second step length are determined.
[0151] Combination Fig. 9 , the processing layer combination corresponding to the first step length is {L1,L(1+step1),L(1+2*step1)}, that is, {L1,L3,L5}, and the processing layer combination corresponding to the second step length is {L1,L(1+step2),L(1+2*step2)}, that is, {L1,L4,L7}.
[0152] Then, the first training duration of the processing layer combination corresponding to the first step length and the second step length is calculated.
[0153] Taking the first training duration as the feature value recalculation duration as an example, calculate the first training duration of the processing layer combination corresponding to the first step and the second step. Fig. 9 , it is calculated that the first training duration of the processing layer combination {L1, L3, L5} is 8ms, and the first training duration of the processing layer combination {L1, L4, L7} is 7ms.
[0154] Finally, the processing layer combination with the shortest first training time is selected as the first processing layer combination.
[0155] Combination Fig. 9, the first training duration 7ms of the processing layer combination {L1, L4, L7} corresponding to step 3 is less than the first training duration 8ms of the processing layer combination {L1, L3, L5} corresponding to step 2, so the processing layer combination {L1, L4, L7} is the first processing layer combination.
[0156] Exemplarily, the control circuit selects the processing layer combination with the minimum number of eigenvalue recalculation times from the processing layer combinations corresponding to the first step length and the second step length as the first processing layer combination, which can be specifically implemented as the following steps:
[0157] First, the processing layer combinations corresponding to the first step length and the second step length are determined.
[0158] Combination Fig. 9 For details, please refer to the above description of determining the processing layer combinations corresponding to the first step length and the second step length, which will not be repeated here.
[0159] Then, the number of eigenvalue recalculations for the combination of processing layers corresponding to the first step length and the second step length is calculated.
[0160] Combination Fig. 9 , the eigenvalue recalculation times of the processing layer combination {L1, L3, L5} is 8 times, and the specific calculation process is as follows:
[0161] L8 reverse training requires recalculating L8 eigenvalues from L5 eigenvalues, and the number of recalculations is 3;
[0162] L7 reverse training requires recalculating L7 eigenvalues from L5 eigenvalues, and the number of recalculations is 2;
[0163] L6 reverse training requires recalculating L6 eigenvalues from L5 eigenvalues, and the number of recalculations is 1;
[0164] L4 reverse training requires recalculating L4 eigenvalues from L3 eigenvalues, and the number of recalculations is 1;
[0165] L2 reverse training requires recalculating L2 eigenvalues from L1 eigenvalues, and the number of recalculations is 1;
[0166] That is, the eigenvalue recalculation time for processing the layer combination {L1, L3, L5} is 3+2+1+1+1+1=8.
[0167] Similarly, the number of eigenvalue recalculations for the processing layer combination {L1, L4, L7} is 7.
[0168] Finally, the processing layer combination with the smallest number of eigenvalue recalculation times is selected as the first processing layer combination.
[0169] Specifically, if the number of eigenvalue recalculations of the processing layer combination {L1, L4, L7} is less than the number of eigenvalue recalculations of the processing layer combination {L1, L3, L5}, then the processing layer combination {L1, L4, L7} is the first processing layer combination.
[0170] In this way, when the training duration of each processing layer allocated to the training circuit is equal, the above formula can be used to calculate the first step length and the second step length, and then only the first training duration or the number of eigenvalue recalculations of the processing layer combination corresponding to the first step length and the second step length need to be calculated to determine the optimal first processing layer combination. There is no need to traverse all second processing layer combinations, which improves the efficiency of determining the optimal first processing layer combination and thereby improves training efficiency.
[0171] S302: The control circuit sends first information to the training circuit. Correspondingly, the training circuit receives the first information from the control circuit.
[0172] The first information is used to instruct the training circuit to store the feature value corresponding to the first processing layer combination during the forward training process.
[0173] Optionally, the first information may also instruct the training circuit not to store feature values corresponding to processing layers other than the first processing layer combination during forward training. Exemplarily, the first information may include a script file or a configuration file.
[0174] Optionally, the control circuit may send indication information of whether the feature value of each processing layer is stored to the training circuit.
[0175] Taking the first processing layer combination {L1, L4, L7} as an example, the control circuit sends instructions or information to the training circuit to store the characteristic value of L1, not store the characteristic value of L2, not store the characteristic value of L3, store the characteristic value of L4, not store the characteristic value of L5, not store the characteristic value of L6, store the characteristic value of L7, and not store the characteristic value of L8. For example, if the binary number 1 represents the storage of the characteristic value and the binary number 0 represents the non-storage of the characteristic value, then the indication information corresponding to the first processing layer combination {L1, L4, L7} can be in the order of the processing layer number from small to large: 10010010.
[0176] Optionally, the control circuit may also send only the information that needs to be stored by the processing layers in the first processing layer combination to the training circuit.
[0177] Taking the first processing layer combination {L1, L4, L7} as an example, the control circuit sends instructions to the training circuit to store the eigenvalues of L1, L4, and L7. For other processing layers that do not instruct the training circuit to store eigenvalues, the default is not to store the eigenvalues of the processing layer.
[0178] S303: The training circuit performs forward training.
[0179] Fig.10 Schematic diagram of the training scenario of the neural network provided in the embodiment of the present application Figure 5 Taking the first processing layer combination {L1, L4, L7} as an example, Fig.10 As shown, the training circuit calculates the input data of the L1 layer and the weight matrix of the L1 layer to obtain the eigenvalue of the L1 layer and saves the eigenvalue. Then, the eigenvalue of the L1 layer is used as the input data of the L2 layer, and the weight matrix of the L2 layer is used for calculation to obtain the eigenvalue of the L2 layer without saving the eigenvalue of the L2 layer. Similarly, the eigenvalue of the L3 layer is calculated but the eigenvalue of the L3 layer is not saved. The eigenvalue of the L4 layer is calculated and saved. The eigenvalue of the L5 layer is calculated but the eigenvalue of the L5 layer is not saved. The eigenvalue of the L6 layer is calculated but the eigenvalue of the L6 layer is not saved. The eigenvalue of the L7 layer is calculated and saved. The eigenvalue of the L8 layer is calculated but the eigenvalue of the L8 layer is not saved.
[0180] S304: The training circuit performs reverse training.
[0181] If the first reverse training layer belongs to the first processing layer combination, the training circuit performs reverse training of the first reverse training layer based on the feature value of the first reverse training layer.
[0182] If the first reverse training layer does not belong to the first processing layer combination, the training circuit determines a first reference layer of the first reverse training layer, and performs reverse training of the first reverse training layer based on the feature value of the first reference layer.
[0183] Among them, the first reverse training layer is any processing layer among the processing layers allocated to the training circuit mentioned above, and the first reference layer is the processing layer in the first processing layer combination whose forward training order is before the first reverse training layer and whose number of processing layers separated from the first reverse training layer is the smallest.
[0184] Taking the first processing layer combination {L1, L4, L7} as an example, combined with Fig.10 The above S304, the training circuit performs reverse training, which can be implemented as follows:
[0185] Step 6: The training circuit performs reverse training on the L8 layer.
[0186] Specifically, the L8 layer is the first reverse training layer, the L8 layer does not belong to the first processing layer combination {L1, L4, L7}, the forward training of the L7 layer is before the L8 layer and the number of processing layers separated from the L8 layer is the smallest, then the L7 layer is the first reference layer of the L8 layer.
[0187] Then, the training circuit recalculates the eigenvalue of the L8 layer. Specifically, the training circuit performs calculation based on the eigenvalue of the L7 layer and the weight matrix of the L8 layer to obtain the eigenvalue of the L8 layer.
[0188] Finally, the training circuit uses the eigenvalues of the L8 layer to perform reverse training on the L8 layer, updates the weight matrix of the L8 layer, and releases the eigenvalues of the L8 layer.
[0189] Step 7: The training circuit performs reverse training on the L7 layer.
[0190] Specifically, the L7 layer is the first reverse training layer. The L7 layer belongs to the first processing layer combination {L1, L4, L7}. The eigenvalues of the L7 layer are stored in the GPU. The eigenvalues of the L7 layer are directly used to reversely train the L7 layer, update the weight matrix of the L7 layer, and release the eigenvalues of the L7 layer.
[0191] Step 8: The training circuit performs reverse training on the L6 layer. Similar to step 6, the eigenvalues of the L5 layer are first calculated using the eigenvalues of the L4 layer, and then the eigenvalues of the L6 layer are calculated using the eigenvalues of the L5 layer. Finally, the eigenvalues of the L6 layer are used to perform reverse training on the L6 layer, update the weight matrix of the L6 layer, and release the eigenvalues of the L6 layer.
[0192] Similarly, the training circuit performs reverse training on the L5, L4, L3, L2, and L1 layers in turn, updates the weight matrix of each processing layer, and completes the reverse training. The process of the training circuit performing reverse training on the L5, L3, and L2 layers can refer to the above step 6, and the process of the training circuit performing reverse training on the L4 and L1 layers can refer to the above step 7, which will not be repeated here.
[0193] based on Figure 3 In the training method of the neural network shown, the control circuit selects a first processing layer combination from the processing layers allocated to the training circuit according to the storage capacity of the training circuit and the storage requirements of the eigenvalues corresponding to each processing layer, so that the storage capacity of the training circuit meets the storage requirements of the eigenvalues corresponding to the first processing layer combination during the training process, and then sends first information to the training circuit to instruct the training circuit to store the eigenvalues corresponding to the first processing layer combination during the forward training process. This can solve the problem of waste caused by a large amount of idle video memory of the training circuit while ensuring the normal execution of the training task of the DNN model, and can reduce the time spent on reverse calculations, thereby improving the training efficiency of the neural network.
[0194] Combination of the above Figure 3-Figure 10 The training method of the neural network in the embodiment of the present application is described in detail. Figure 11-13 The neural network device provided in the embodiments of the present application is described in detail.
[0195] Fig.11This is a schematic diagram of the structure of the control circuit of the neural network provided in the embodiment of the present application. The control circuit can be applied to Figure 1 In the training system shown, execution Figure 3 The function of the control circuit in the training method shown. For the sake of explanation, Fig.11 Only the main components of the control circuit are shown.
[0196] like Fig.11 As shown, the control circuit 1100 includes: a processor 1102 and a transmitter 1101 .
[0197] The processor 1102 is used to determine a first processing layer combination. The first processing layer combination includes multiple processing layers, and the multiple processing layers are selected from the processing layers allocated to the training circuit. The storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process. The transmitter 1101 is also used to send the first information to the training circuit. The first information is used to instruct the training circuit to store the characteristic values corresponding to the first processing layer combination during the forward training process.
[0198] In a possible design, the processor 1102 is further configured to determine a plurality of processing layer combinations. The storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to any of the plurality of processing layer combinations during the training process. The processor 1102 is further configured to determine one or more processing layer combinations with the largest number of processing layer layers included in the plurality of processing layer combinations as the second processing layer combination. The processor 1102 is further configured to select the first processing layer combination from the second processing layer combination.
[0199] In a possible design, the number of processing layers included in the first processing layer combination is: Where N is the number of processing layers in the first processing layer combination, C MAX is the storage capacity of the training circuit, C1 is the first reserved capacity, and C is the characteristic value storage requirement of any processing layer allocated to the training circuit. To round down.
[0200] In a possible design, the processor 1102 is further configured to determine the processing layer combination with the shortest first training duration in the second processing layer combination as the first processing layer combination, wherein the first training duration includes the eigenvalue recalculation time, or the first training duration includes the forward training time, the reverse training time, and the eigenvalue recalculation time.
[0201] Optionally, if the training duration of each processing layer allocated to the training circuit is equal, the processor 1102 is further configured to calculate the first step length and the second step length according to the following formula: Wherein, step1 is the first step length, step2 is the second step length, M is the total number of processing layers allocated to the training circuit, N is the number of processing layers in the first processing layer combination, The processor 1102 is further used to select, from the processing layer combinations corresponding to the first step length and the second step length, a processing layer combination with the shortest first training time as the first processing layer combination, or the processor 1102 is further used to select, from the processing layer combinations corresponding to the first step length and the second step length, a processing layer combination with the smallest number of eigenvalue recalculations as the first processing layer combination.
[0202] Optionally, the storage capacity of the above-mentioned training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process, and may include: the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to the first processing layer combination during the training process and the first reserved capacity; wherein, the first reserved capacity is used to store the intermediate data generated during the training process.
[0203] It should be noted that, in addition to the characteristic value of the first processing layer combination and the first reserved capacity, the storage capacity of the training circuit also needs to meet the storage requirements of the read-in training samples. Therefore, optionally, a reserved storage space can also be set for the above-mentioned read-in training samples. Among them, the amount of reserved storage space can be set according to actual needs or actual experience, and the embodiments of the present application do not specifically limit this.
[0204] Optionally, the control circuit 1100 may further include a receiver ( Fig.11 1101). The receiver is used to receive data sent by the training circuit. Further, the receiver and the transmitter 1101 can be set separately or integrated together, that is, a transceiver. The present application does not specifically limit the specific implementation of the receiver and the transmitter 1101.
[0205] Optionally, the control circuit 1100 may further include a memory ( Fig.11 The memory stores a program or instruction. When the processor 1102 executes the program or instruction, the control circuit 1100 can perform the function of the control circuit in the training method described in the above method embodiment.
[0206] It should be noted that the control circuit 1100 may be Figure 1 The control circuit shown, or Figure 2 The training device 200 is shown, but the present application does not limit this.
[0207] also, Fig.11 The technical effects of the control circuit 1100 shown can refer to the technical effects of the neural network training method described in the above method embodiment, and will not be repeated here.
[0208] Fig.12 A schematic diagram of the structure of a neural network training circuit provided in an embodiment of the present application. The training circuit can be applied to Figure 1 In the training system shown, execution Figure 3 The function of the training circuit in the training method shown. For the sake of explanation, Fig.12 Only the main components of the training circuit are shown.
[0209] like Fig.12 As shown, the training circuit 1200 includes: a receiver 1201, a processor 1202 and a memory 1203. The receiver 1201 is used to receive the first information sent by the control circuit. The first information is used to instruct the processor 1202 to store the characteristic value corresponding to the first processing layer combination during the forward training process. The first processing layer combination includes multiple processing layers, and the multiple processing layers are selected from the processing layers allocated to the training circuit. The storage capacity of the memory 1203 meets the storage requirements of the characteristic value corresponding to the first processing layer combination during the training process. The processor 1202 is used to perform reverse training according to the characteristic value corresponding to the first processing layer combination.
[0210] In a possible design, the processor 1202 is further used to perform reverse training of the first reverse training layer based on the characteristic value of the first reverse training layer if the first reverse training layer belongs to the first processing layer combination. The processor 1202 is further used to determine the first reference layer of the first reverse training layer if the first reverse training layer does not belong to the first processing layer combination, and perform reverse training of the first reverse training layer based on the characteristic value of the first reference layer. The first reverse training layer is any processing layer among the processing layers allocated to the training circuit. The first reference layer is a processing layer in the first processing layer combination whose forward training order is before the first reverse training layer and whose number of processing layers separated from the first reverse training layer is the smallest.
[0211] Optionally, the training circuit 1200 may further include a transmitter ( Fig.12 1201 and the transmitter may be provided separately or integrated together, i.e., a transceiver. The present application does not specifically limit the specific implementation of the receiver and the transmitter.
[0212] Optionally, the memory 1203 of the training circuit 1200 may also store programs or instructions. When the processor 1202 executes the programs or instructions, the training circuit 1200 may perform the functions of the training circuit in the training method described in the above method embodiment.
[0213] It should be noted that the training circuit 1200 may be Figure 1 The training circuit shown, or Figure 2The training device 200 is shown, but the present application does not limit this.
[0214] also, Fig.12 The technical effects of the training circuit 1200 shown can refer to the technical effects of the neural network training method described in the above method embodiment, and will not be repeated here.
[0215] Fig.13 A schematic diagram of the structure of a neural network training device provided in an embodiment of the present application Figure 2 The training device can be used for Figure 1 In the training system shown, execution Figure 3 The functions of the control circuit and the training circuit in the training method shown. For the convenience of explanation, Fig.13 Only the main components of the training device are shown.
[0216] like Fig.13 As shown, the training device 1300 includes: a control circuit 1301 and one or more training circuits 1302 .
[0217] Among them, the control circuit 1301 is the control circuit described in any possible implementation of the training method described in the above method embodiment, and the training circuit 1302 is the training circuit described in any possible implementation of the training method described in the above method embodiment.
[0218] In a possible design, the control circuit 1301 may be a central processing unit CPU, and the training circuit 1302 may be a graphics processing unit GPU.
[0219] also, Fig.13 The technical effects of the training device 1300 shown can refer to the technical effects of the neural network training method described in the above method embodiment, and will not be repeated here.
[0220] The embodiment of the present application provides a computer-readable storage medium, including: the computer-readable storage medium includes a program or instruction, when the program or instruction is run on a computer, the computer executes the function of the control circuit described in the above method embodiment. It should be noted that the computer-readable storage medium can be a non-transitory computer-readable storage medium.
[0221] The embodiment of the present application provides another computer-readable storage medium, including: the computer-readable storage medium includes a program or instruction, and when the program or instruction is executed on a computer, the computer performs the function of the training circuit described in the above method embodiment. It should be noted that the computer-readable storage medium can be a non-transitory computer-readable storage medium.
[0222] An embodiment of the present application provides a computer program product, which includes: a computer program code, when the computer program code is run on a computer, the computer executes the function of the control circuit described in the above method embodiment.
[0223] An embodiment of the present application provides another computer program product, which includes: a computer program code, when the computer program code is executed on a computer, the computer executes the function of the training circuit described in the above method embodiment.
[0224] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM).
[0225] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website site, computer, server or data center to another website site, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0226] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0227] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean one of the following items: a; b; c; a and b; a and c; b and c; a, b, and c, where a, b, and c can be single or multiple.
[0228] It should be understood that in various embodiments of the present application, the size of the sequence number of each process does not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application. Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but this implementation should not be considered to be beyond the scope of the present application.
[0229] If the functions described in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., and other media that can store program codes.
[0230] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A neural network training method, characterized in that: include: The control circuit determines a first processing layer combination; wherein the first processing layer combination includes a plurality of processing layers, the plurality of processing layers are selected from the processing layers allocated for the training circuit, and the storage capacity of the training circuit meets the storage requirements of the feature values corresponding to the first processing layer combination during the training process; The control circuit sends first information to the training circuit; wherein the first information is used to instruct the training circuit to store the feature value corresponding to the first processing layer combination during the forward training process.
2. The neural network training method according to claim 1, characterized in that: The control circuit determines a first processing layer combination, including: The control circuit determines a plurality of processing layer combinations; wherein the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to any one of the plurality of processing layer combinations in the training process; The control circuit determines one or more processing layer combinations with the largest number of processing layers included in the plurality of processing layer combinations as the second processing layer combination; The control circuit selects the first processing layer combination from the second processing layer combination.
3. The neural network training method according to claim 2, characterized in that: The number of processing layers included in the first processing layer combination is: Wherein, N is the number of processing layers of the first processing layer combination, C MAX is the storage capacity of the training circuit, C1 is the first reserved capacity, C is the characteristic value storage requirement of any processing layer allocated to the training circuit, To round down.
4. The neural network training method according to claim 2 or 3, characterized in that: The control circuit selects the first processing layer combination from the second processing layer combination, comprising: The control circuit determines the processing layer combination with the shortest first training time in the second processing layer combination as the first processing layer combination; Wherein, the first training duration includes the feature value recalculation time; or, The first training duration includes forward training time, reverse training time and feature value recalculation time.
5. The neural network training method according to claim 4, characterized in that: The training duration of each processing layer allocated to the training circuit is equal; The neural network training method also includes: The control circuit calculates the first step length and the second step length according to the following formula: Wherein, step1 is the first step length, step2 is the second step length, M is the total number of processing layers allocated to the training circuit, N is the number of processing layers in the first processing layer combination, To round down; The control circuit selects, from the processing layer combinations corresponding to the first step length and the second step length, the processing layer combination with the shortest first training duration as the first processing layer combination; or, The control circuit selects a processing layer combination with the smallest number of eigenvalue recalculations from the processing layer combinations corresponding to the first step length and the second step length as the first processing layer combination.
6. The neural network training method according to any one of claims 1 to 3, characterized in that: The storage capacity of the training circuit meets the storage requirements of the feature values corresponding to the first processing layer combination during the training process, including: The storage capacity of the training circuit meets the storage requirements and the first reserved capacity of the characteristic values corresponding to the first processing layer combination during the training process; wherein the first reserved capacity is used to store the intermediate data generated during the training process.
7. The neural network training method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the first reverse training layer belongs to the first processing layer combination, the training circuit performs reverse training of the first reverse training layer based on the feature value of the first reverse training layer; If the first reverse training layer does not belong to the first processing layer combination, the training circuit determines a first reference layer of the first reverse training layer, and performs reverse training of the first reverse training layer based on a feature value of the first reference layer; Among them, the first reverse training layer is any processing layer among the processing layers allocated to the training circuit; the first reference layer is a processing layer in the first processing layer combination whose forward training order is before the first reverse training layer and whose number of processing layers separated from the first reverse training layer is the smallest.
8. A control circuit of a neural network, characterized in that: include: processor and transmitter; The processor is used to determine a first processing layer combination; wherein the first processing layer combination includes a plurality of processing layers, the plurality of processing layers are selected from processing layers allocated for a training circuit, and the sum of a storage requirement of a feature value corresponding to the first processing layer combination and a storage requirement of a parameter of the processing layer allocated for the training circuit is less than a storage capacity of the training circuit; The transmitter is further used to send first information to the training circuit; wherein the first information is used to instruct the training circuit to store the feature value corresponding to the first processing layer combination during the forward training process.
9. The control circuit of the neural network according to claim 8, characterized in that: The processor is further configured to determine a plurality of processing layer combinations; wherein the storage capacity of the training circuit meets the storage requirements of the characteristic values corresponding to any one of the plurality of processing layer combinations in the training process; The processor is further configured to determine one or more processing layer combinations with the largest number of processing layers among the multiple processing layer combinations as the second processing layer combination; The processor is further configured to select the first processing layer combination from the second processing layer combination.
10. The control circuit of the neural network according to claim 9, characterized in that: The number of processing layers included in the first processing layer combination is: Wherein, N is the number of processing layers of the first processing layer combination, C MAX is the storage capacity of the training circuit, C1 is the first reserved capacity, C is the characteristic value storage requirement of any processing layer allocated to the training circuit, To round down.
11. The control circuit of the neural network according to claim 9 or 10, characterized in that: The processor is further configured to determine the processing layer combination with the shortest first training duration in the second processing layer combinations as the first processing layer combination; Wherein, the first training duration includes the feature value recalculation time; or, The first training duration includes forward training time, reverse training time and feature value recalculation time.
12. The control circuit of the neural network according to claim 11, characterized in that: The training duration of each processing layer allocated to the training circuit is equal; The processor is further configured to calculate the first step length and the second step length according to the following formula: Wherein, step1 is the first step length, step2 is the second step length, M is the total number of processing layers allocated to the training circuit, N is the number of processing layers in the first processing layer combination, To round down; The processor is further configured to select, from the processing layer combinations corresponding to the first step length and the second step length, the processing layer combination with the shortest first training duration as the first processing layer combination; or The processor is further configured to select, from the processing layer combinations corresponding to the first step length and the second step length, a processing layer combination with the smallest number of eigenvalue recalculations as the first processing layer combination.
13. The control circuit of the neural network according to any one of claims 8 to 10, characterized in that: The storage capacity of the training circuit meets the storage requirements of the feature values corresponding to the first processing layer combination during the training process, including: The storage capacity of the training circuit meets the storage requirements and the first reserved capacity of the characteristic values corresponding to the first processing layer combination during the training process; wherein the first reserved capacity is used to store the intermediate data generated during the training process.
14. A training device, characterized in that: The training device includes: a processor coupled to a memory; The memory is used to store computer programs; The processor is used to execute the computer program stored in the memory so that the training device implements the method as described in any one of claims 1-6.
15. A training device, characterized in that: The training device comprises: a control circuit as described in any one of claims 8 to 13, and one or more training circuits.
16. The training device according to claim 15, characterized in that The control circuit is a central processing unit (CPU), and the training circuit is a graphics processing unit (GPU).
17. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a program or an instruction. When the program or the instruction is executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 6.
18. A computer program product, characterized in that The computer program product comprises: a computer program code, and when the computer program code is run on a computer, the computer is caused to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Deep learning heterogeneous computing method and system based on layer width memory allocation
CN109976903A
Artificial neural network adjustment method and device
CN110413255A