Neural network mixing precision training method, device and equipment based on sampling
By using two-stage sampling and optimization methods during the training process of neural network model, the mixed accuracy training strategy is adjusted, and the problem of high cost of neural network training in the existing technology is solved, efficient and low-cost mixed accuracy training is achieved, ensuring the convergence and performance of the model.
Patent Information
- Application Number
- CN202510015985.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology cannot effectively reduce the cost of neural network training while ensuring the convergence of neural network models. Moreover, the dimensions considered by the hybrid accuracy training strategy are not comprehensive enough, the accuracy decisions are not flexible enough, and the cost of use is high.
The sampling-based hybrid accuracy training method is adopted, and two-stage sampling and optimization are carried out: first, at the round level, the operator with a significant impact on convergence is adjusted according to expert knowledge, to ensure that the loss value of the adjustment scheme is within an acceptable range, and the adjustment scheme with the best performance is selected as the initial hybrid accuracy scheme; then the adjustment is further refined at the batch level, traversing and collecting the performance data of all hybrid accuracy schemes, and selecting the hybrid accuracy scheme with the best performance is selected as the post-training input.
It realizes the efficiency and accuracy of neural network model training while ensuring convergence, reduces the consumption of computing resources, and achieves dual optimization of model performance and training cost.
Smart Images

Figure CN120105090A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a sampling-based neural network mixed precision training method, device and equipment. Background Art
[0002] In the field of deep learning, model training is a computationally intensive and memory intensive process, especially when dealing with large-scale data sets and complex models. In order to improve training efficiency and reduce computing resources and memory consumption, mixed precision training technology has emerged. Mixed precision training refers to the use of floating point representations of different precisions, such as FP32 (32-bit floating point) and FP16 (16-bit floating point) during training to increase computing speed and save memory. By reasonably setting the precision of the operator, the training efficiency can be significantly improved while ensuring model convergence.
[0003] Currently, deep learning frameworks such as PyTorch and TensorFlow have built-in support for mixed precision training. These frameworks usually implement mixed precision training through two steps: automatic data precision conversion and gradient scaling. The automatic precision conversion mechanism selects the calculation precision for the operator according to the preset "black and white list" strategy. Operators in the white list (such as Conv2d, Linear, etc.) are usually set to FP16 by default to take advantage of hardware acceleration, while operators in the black list (such as Softmax, Layer_norm, etc.) are forced to be set to FP32 because they may cause numerical overflow. In addition, some research works have further considered the cost of converting from FP16 to FP32, and proposed methods for predicting tensor conversion costs based on offline training models, as well as distributed mixed precision strategies.
[0004] Although existing technologies have improved the efficiency of mixed-precision training to a certain extent, the current deep learning framework and mixed-precision training work cannot achieve efficient neural network training at a low cost while ensuring convergence. Summary of the invention
[0005] The present invention provides a sampling-based neural network mixed precision training method, device and equipment, which solves the problem of high cost of neural network model training in the prior art and realizes efficient and low-cost mixed precision training.
[0006] The present invention provides a neural network mixed precision training method based on sampling, comprising the following steps: Input the neural network model to be trained into the round-based sampling phase; In the round-based sampling phase, operators whose influence on convergence exceeds a preset influence threshold are adjusted based on expert knowledge in units of training rounds, and each adjustment scheme is run for a complete training round, and the loss value of each adjustment scheme is recorded. The adjustment schemes whose loss values are within the acceptable range are retained, and the adjustment scheme with the best performance is selected from the retained adjustment schemes as the initial mixed precision scheme; Inputting the initial mixed precision scheme into a batch-based sampling phase; In the batch-based sampling phase, taking the initial mixed precision scheme as a benchmark, adjusting the adjustable operator precision in the initial mixed precision scheme, traversing all mixed precision schemes and collecting performance data in units of training batches, and selecting the mixed precision scheme with the best performance as the input scheme for post-training according to the performance data; An input scenario for executing the post-training.
[0007] According to a sampling-based neural network mixed precision training method provided by the present invention, the adjustment of operators whose influence on convergence exceeds a preset influence threshold according to expert knowledge specifically includes: recording operators with conversion tendency to form a first-stage operation list; determining operators belonging to a whitelist and operators belonging to a blacklist in the first-stage operation list; checking whether the input tensors of the operators in the whitelist meet the tensor core usage conditions, and if so, adjusting the precision to 16-bit floating point numbers; testing the convergence of the operators in the blacklist converted to 16-bit floating point numbers, and retaining the original precision if the convergence does not meet the preset convergence conditions.
[0008] According to a sampling-based neural network mixed precision training method provided by the present invention, the adjustment scheme for retaining the loss value within an acceptable range specifically includes: when the loss value is less than 101% of the loss value obtained when all operators are set to 32-bit floating point precision, it is determined that the convergence of the adjustment scheme is within an acceptable range, and the adjustment scheme with convergence within the acceptable range is retained.
[0009] According to a sampling-based neural network mixed precision training method provided by the present invention, the adjusting of the adjustable operator precision in the initial mixed precision scheme, taking training batches as units, traversing all mixed precision schemes, specifically includes: constructing a search space for operators other than the operators in the first-stage operation list; setting the precision of operators other than the operators in the first-stage operation list in the mixed precision scheme to default to 32-bit floating point numbers; and traversing the mixed precision scheme with the positions of the operators in the first-stage operation list as split points.
[0010] According to a sampling-based neural network mixed precision training method provided by the present invention, in the process of traversing the mixed precision scheme, it includes: when the precision of two non-adjacent operators in the first-stage operation list is consistent, the operator between the two operators also maintains the precision of 32-bit floating point numbers; when the precision of two non-adjacent operators in the first-stage operation list is inconsistent, enumeration is performed to adjust the precision of the operator sandwiched between the two operators.
[0011] According to a sampling-based neural network mixed precision training method provided by the present invention, after adjusting the adjustable operator precision in the initial mixed precision scheme, the method also includes: after adjusting the adjustable operator precision in the initial mixed precision scheme, a plurality of mixed precision schemes are obtained, each mixed precision scheme defines a new model for testing; during the testing process, the data set is divided according to the batch size, and the same amount of calculation is divided for each batch, and then a single batch is extracted for testing, and a different mixed precision scheme is assigned to each batch.
[0012] According to a sampling-based neural network mixed precision training method provided by the present invention, when the neural network model is constructed, the method also includes: spreading out the model operators through sequential programming; passing in the mixed precision scheme in the form of a string; modifying the forward function of the neural network model and automatically inserting the precision conversion function.
[0013] The present invention also provides a neural network mixed precision training device based on sampling, comprising the following modules: A first input module, used to input the neural network model to be trained into the round-based sampling phase; A first sampling module is used to adjust the operators whose influence on convergence exceeds a preset influence threshold according to expert knowledge in a round-based sampling phase, taking training rounds as units, and run each adjustment scheme for a complete training round, record the loss value of each adjustment scheme, retain the adjustment schemes whose loss values are within an acceptable range, and select the adjustment scheme with the best performance from the retained adjustment schemes as the initial mixed precision scheme; A second input module, configured to input the initial mixed precision scheme into a batch-based sampling phase; A second sampling module is used to adjust the adjustable operator precision in the initial mixed precision scheme based on the initial mixed precision scheme in the batch-based sampling phase, traverse all mixed precision schemes and collect performance data in units of training batches, and select the mixed precision scheme with the best performance as the input scheme for post-training according to the performance data; An execution module is used to execute the post-training input scheme.
[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the sampling-based neural network mixed precision training method described above is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the sampling-based mixed precision training methods for a neural network as described above.
[0016] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned sampling-based mixed precision training methods for a neural network.
[0017] The present invention provides a sampling-based neural network mixed precision training method, device and equipment, including the following beneficial effects: through two-stage sampling and optimization, firstly, based on the round adjustment, the operator that has a significant impact on convergence is ensured to have an acceptable loss value, and the adjustment scheme with the best performance is selected as the initial mixed precision scheme; then, the adjustment is further refined at the batch level, the performance data of all mixed precision schemes are traversed and collected, and finally the mixed precision scheme with the best performance is selected as the post-training input. This method not only improves the efficiency and accuracy of neural network model training, but also significantly reduces the consumption of computing resources through the effective selection of mixed precision training schemes, and achieves dual optimization of model performance and training cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 This is one of the flow charts of the sampling-based neural network mixed precision training method provided by the present invention.
[0020] Figure 2 This is the second flow chart of the sampling-based neural network mixed precision training method provided by the present invention.
[0021] Figure 3 It is a schematic diagram of the overall sampling and post-training of the sampling-based neural network mixed precision training method provided by the present invention.
[0022] Figure 4It is a schematic diagram of the relationship between the sampling-based neural network mixed precision training method provided by the present invention and the deep learning framework.
[0023] Figure 5 It is a structural schematic diagram of a sampling-based neural network mixed precision training device provided by the present invention.
[0024] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] Decimals are represented by floating-point numbers in computers. There are different precision standards for the distribution range and accuracy of the values that need to be represented. For many scientific calculations, higher precision is pursued, but in some cases such precise numerical expression is not required, which gives room for the use of lower precision FP32 and FP16. In the field of deep learning, the weights saved in some calculation processes and the numerical distribution characteristics of some layers do not require too wide an expression range and accuracy in many cases. Therefore, mixed precision technology is widely used in deep learning. The FP32 / FP16 mixed precision training method is widely used in the training stage. By converting some operators into FP16 precision calculations in the training stage, but saving the weights as FP32 storage, combined with gradient scaling, the model training speed is effectively improved and memory is saved while ensuring convergence.
[0027] The mixed precision training strategy in the current deep learning framework does not take into account the problem dimension enough. It is mainly determined by whether the structural characteristics of the model itself are prone to non-convergence, or more specifically, whether the operators at each layer are prone to numerical overflow. Such a decision means that it is separated from the data set, the numerical distribution characteristics processed by the same operator in the model at different training stages, and the size of the input tensor. In addition, the computing power of different GPU hardware for processing FP32 and FP16 is also different, and there is a lack of consideration for the intermediate conversion cost. The current solutions for mixed precision training of neural networks can be summarized in the following two aspects: (1) Mixed Precision Training Scheme in Deep Learning Framework This work usually combines the two steps of automatic data precision conversion and gradient scaling to implement mixed precision training on GPU. Taking PyTorch as an example, using it for neural network training means using both automatic precision conversion (Autocast) and gradient scaler (Gradscaler). When automatic precision conversion is used, the calculation precision will be automatically selected for the operator in the part it modifies according to the setting of the "black and white list" mechanism. Some operators whose calculation results are not easy to overflow or underflow are in the white list and will be automatically converted to FP16 during calculation. On the contrary, some operators whose calculation results are easy to overflow or underflow are in the black list and will be forced to be converted to FP32 during calculation. Some operators outside the black and white list will need to select precision in a way that ensures that there will be no overflow in combination with the context of the operator, usually a more conservative FP32. TensorFlow also has a similar idea, and the mixed precision strategy needs to be explicitly set in the code. Therefore, the current way of conducting mixed precision training in deep learning frameworks mainly considers two dimensions: 1) time performance, and 2) whether the value will overflow. However, there is no fine-grained adjustment and adaptation for the precision conversion overhead and the large differences in the numerical distribution ranges of different models and data sets.
[0028] (2) Mixed Precision Training Scheme in Research Work This research further considers the cost of converting FP16 to FP32. Through offline training models, online prediction of tensor conversion costs, and certain padding of tensors to meet the use conditions of tensor cores, a training framework that can perceive the input tensor size and precision conversion costs is proposed. This type of work with offline operations is intrusive in modifying the framework, which is not flexible enough in the face of the rapidly developing deep learning framework, and the adaptation cost of retraining prediction models for different GPU architectures is also relatively high. There are also research works that propose a distributed mixed precision strategy, where forward and back propagation use FP16, and gradients and weights are completely FP32. However, this method does not consider the convergence brought by different operators in training in detail, and may lead to non-convergence for some models that contain special activation functions and reduction operations. Therefore, the defects of the mixed precision training scheme in the current research work are mainly in two points: 1) It is difficult to adapt to the new GPU architecture, the constantly updated deep learning framework, and the model training with different numerical distributions at low cost. 2) Under the premise of ensuring convergence, there is still room for performance improvement.
[0029] In summary, the current deep learning framework and mixed precision training cannot achieve efficient neural network training at a low cost while ensuring convergence. Efficient neural network training that ensures convergence mainly faces the following challenges: 1) The dimensions considered are not comprehensive enough, and the decisions on different data sets, different hardware, and input tensor sizes are not fine-grained enough. 2) The precision decision is not flexible enough, and the appropriate precision of a certain operator is limited in advance based only on expert knowledge. 3) The cost of use is high, and a large number of offline operations and screening models built for the architecture will make it difficult to achieve the same good results after updating the framework and hardware.
[0030] The present invention provides a neural network convergence-aware mixed precision training method and device based on sampling to solve the problems that the dimensions considered in the existing neural network mixed precision training are not comprehensive enough, the precision decision is not flexible enough, and the cost of existing research work is high. It provides an efficient method to ensure convergence for the mixed precision training of neural networks.
[0031] Combine the following Figure 1-Figure 6 Embodiments of the present invention are described in detail.
[0032] Figure 1 This is one of the flow charts of the sampling-based mixed precision training method for a neural network provided by the present invention, such as Figure 1 As shown, the method comprises the following steps: S110, inputting the neural network model to be trained into the round-based sampling phase.
[0033] S120. In the round-based sampling phase, operators whose impact on convergence exceeds a preset impact threshold are adjusted based on expert knowledge in units of training rounds, and each adjustment scheme is run for a complete training round. The loss value of each adjustment scheme is recorded, and the adjustment schemes with loss values within an acceptable range are retained. Among the retained adjustment schemes, the adjustment scheme with the best performance is selected as the initial mixed precision scheme.
[0034] According to a sampling-based neural network mixed precision training method provided by the present invention, operators whose influence on convergence exceeds a preset influence threshold are adjusted according to expert knowledge, specifically including: recording operators with conversion tendencies to form a first-stage operation list; determining operators belonging to a whitelist and operators belonging to a blacklist in the first-stage operation list; checking whether the input tensors of the operators in the whitelist meet the tensor core usage conditions, and if so, adjusting the precision to 16-bit floating point numbers; testing the convergence of the operators in the blacklist converted to 16-bit floating point numbers, and retaining the original precision if the convergence does not meet the preset convergence conditions.
[0035] According to a sampling-based neural network mixed precision training method provided by the present invention, an adjustment scheme that retains the loss value within an acceptable range specifically includes: when the loss value is less than 101% of the loss value obtained when all operators are set to 32-bit floating point precision, it is determined that the convergence of the adjustment scheme is within an acceptable range, and the adjustment scheme whose convergence is within the acceptable range is retained.
[0036] Specifically, Figure 2 The figure shows the second flow chart of the sampling-based neural network mixed precision training method provided by the present invention, step S11 is based on round sampling: by round-based sampling to ensure convergence, only operators that meet the conditions and have a greater impact on convergence are adjusted according to expert knowledge, so as to find a mixed precision solution with relatively good performance under the premise of ensuring convergence; It should be noted that the deep learning framework uses a list mechanism when using the built-in mixed precision solution for the operators that make up the model. More specifically, it is determined by whether the calculation structure of the operator is prone to overflow or underflow, that is, it is determined by numerical security. For example, for the operator Linear, it exists in the whitelist of the framework, which means that the operator is generally numerically safe, but sometimes due to different contextual precision conditions, it brings additional precision conversion overhead, or the FP16 (16-bit floating point precision) operator cannot use the tensor core to enjoy the hardware acceleration effect due to certain dimensions, which makes the tendency to convert to FP16 unreasonable. Or for the operator exp, it exists in the blacklist of the framework, but sometimes due to different specific parameters of the model and differences in the numerical distribution of the data set, it does not cause the overflow expected by the conservative strategy.
[0037] In an optional embodiment, before the optimal operator-level mixed precision scheme is put into the machine learning model, for the model defined by the deep learning framework, the precision that each operator will tend to use (16-bit floating point precision FP16 or 32-bit floating point precision FP32) is determined according to the list mechanism in the deep learning framework; for the neural network model based on round sampling, each mixed precision scheme must thoroughly traverse the data set once to obtain the loss value and record the processing time; for the neural network model based on batch sampling, on the basis of the existing mixed precision scheme, other operator combinations that have little effect on convergence are enumerated, and a search space is constructed to find the training mixed precision scheme with the best performance.
[0038] As shown in Table 1, the first-stage operation list Stage1_OP_List contains operators with obvious conversion trends. According to the list mechanism of the deep learning framework and the usage of tensor cores, operators with obvious conversion trends are extracted and put into the operator list Stage1_OP_List; the precision of the operators Conv2d, Linear, Matmul, and ReLU in the whitelist of the deep learning framework is set to FP16 by default; the operators in the whitelist are checked to see whether their input tensors meet the requirements for enabling tensor cores; the operators in the blacklist such as Softmax and Layer_norm are set to try to convert to FP16 and sense whether they converge.
[0039] Table 1
[0040] It should be noted that sampling is performed in rounds based on the operator list Stage1_OP_List, including: for operators with clear conversion tendencies in the deep learning framework, round-based sampling first needs to obtain the information of these operators and save them in Stage1_OP_list when performing the first benchmark sampling; in round-based sampling, the operators in Stage1_OP_list are mainly adjusted, but all combinations are not directly enumerated; the operator in the input model is checked. If it exists in the whitelist of the deep framework, its parameters will be further checked to see if they meet the requirements for using tensor cores. If so, its calculation accuracy is forced to be set to FP16; if it does not meet the requirements for using tensor cores, this operator may be adjusted to FP32, and needs to be tested as an optional mixed precision scheme in round-based sampling; if the operator is in the blacklist of the deep learning framework, it is also tested as an optional mixed precision scheme.
[0041] like Figure 3 The darker operator-3 and operator-5 in the model on the left; if the requirements for using Tensor Cores are not met, such as Figure 3 The light-colored operators -0 and -7 on the left in the middle may be adjusted to FP32, and need to be tested as an optional solution in round-based sampling. On the other hand, if the operator is in the blacklist of the deep learning framework, such as operator -5, then it also gives us the possibility of trying it, also as an optional solution. There are three operators that need to try different precision combinations, namely operator -0, operator -4 and operator -7. Although operator -3 and operator -5 are in Stage1_OP_list, they can be accelerated using the upper tensor core, so they are not tried. In this embodiment, the loss values of operator -0, operator -4 and operator -7 at different precisions are recorded. iand running time; compared with the loss value loss obtained by the first sampling based on the round, that is, all operators are run completely with FP32 precision 0 In comparison, when loss i <loss 0 *101%, it is considered to be the convergence safety range; the loss within this convergence safety range i The set is {loss selected}, and select the solution with the shortest running time, that is, the best speed performance, from this set.
[0042] S130. Input the initial mixed precision solution into the batch-based sampling phase.
[0043] S140. In the batch-based sampling stage, the initial mixed precision scheme is used as a benchmark to adjust the adjustable operator precision in the initial mixed precision scheme. All mixed precision schemes are traversed and performance data is collected in training batches. The mixed precision scheme with the best performance is selected as the input scheme for post-training based on the performance data.
[0044] According to a sampling-based neural network mixed precision training method provided by the present invention, the adjustable operator precision in the initial mixed precision scheme is adjusted, and all mixed precision schemes are traversed in units of training batches, specifically including: constructing a search space for operators other than the operators in the first-stage operation list; setting the precision of operators other than the operators in the first-stage operation list in the mixed precision scheme to 32-bit floating point numbers by default; and traversing the mixed precision scheme with the position of the operator in the first-stage operation list as the dividing point.
[0045] According to a sampling-based neural network mixed precision training method provided by the present invention, in the process of traversing the mixed precision scheme, it includes: when the precision of two non-adjacent operators in the first-stage operation list is consistent, the operator between the two operators also maintains the precision of 32-bit floating point numbers; when the precision of two non-adjacent operators in the first-stage operation list is inconsistent, enumeration is performed to adjust the precision of the operator sandwiched between the two operators.
[0046] According to a sampling-based neural network mixed precision training method provided by the present invention, after adjusting the adjustable operator precision in the initial mixed precision scheme, the method also includes: after adjusting the adjustable operator precision in the initial mixed precision scheme, a plurality of mixed precision schemes are obtained, each mixed precision scheme defines a new model for testing; during the testing process, the data set is divided according to the batch size, and the same amount of calculation is divided for each batch, and then a single batch is extracted for testing, and a different mixed precision scheme is assigned to each batch.
[0047] Specifically, Figure 2Step S12 is batch-based sampling: The optimal performance is searched through a batch-based sampling strategy, which combines the data set, model structure and current GPU hardware characteristics to obtain the optimal mixed precision solution for the model with a single batch at a low cost.
[0048] Specifically, in the batch sampling stage, each mixed precision scheme defines a new model for testing, including: the data set is divided according to the batch size, the same amount of computation is allocated to each batch, and a single batch is extracted for testing; each batch is used as a relatively lightweight and low-cost basic test unit, and a different mixed precision scheme is assigned to a batch, which means that multiple models are defined, but each model only runs the first batch in a round.
[0049] like Figure 3 As shown in the figure, operators with a certain precision conversion tendency in the deep learning framework are identified, that is, the operator precision in Stage1_OP_list. Some of these operators are numerically safe and have the potential to use tensor cores, while others are not numerically safe, so they are set to FP32 in the preliminary plan. The remaining operators do not have such a conversion tendency, and the specific plan needs to be determined based on the precision of the context. These operators are interspersed in the model, such as Figure 3 There are no colored operators -1, -2, and -6 in the left model. Considering the acceleration effect of specific GPU hardware, the size of the input tensor, and the precision conversion cost of its context, other solutions need to be tried.
[0050] Since the precision of operators in Stage1_OP_list has been determined, the precision of only some operators in the model has not been determined. It is necessary to consider the context conversion to enumerate the remaining operator precision schemes. Determine the precision scheme with the best performance within the acceptable loss value floating range. In fact, this scheme is really determined to be the operator in Stage1_OP_list. The operator to be adjusted based on batch sampling is still the default FP32 precision for the time being. Use the operator position in Stage1_OP_list as the split point for traversal. When the precision of two non-adjacent operators in Stage1_OP_list is consistent, for example, both are FP32, then the operator between the two operators also maintains FP32 precision, although in fact executing the middle operator alone may obtain better performance in FP16. When the precision of two non-adjacent operators in Stage1_OP_list is inconsistent, then enumeration can be performed to adjust the precision of the middle sandwiched operator. Such enumeration has been tested and combines the processing power of the current GPU hardware, the cost difference of FP16 / FP32 conversion, and the fixed cost of the conversion itself.
[0051] S150, executing the input scheme of post-training.
[0052] According to a sampling-based neural network mixed precision training method provided by the present invention, when a neural network model is constructed, the model operators are spread out through sequential programming; the mixed precision scheme is passed in the form of a string; the forward function of the neural network model is modified, and the precision conversion function is automatically inserted.
[0053] Specifically, if Figure 2 In step S13, the optimization scheme is applied: the optimal operator-level mixed precision scheme after two-stage sampling and screening is used as the mixed precision scheme for post-training.
[0054] Specifically, the coupling work with the deep learning framework includes: using sequential programming when building the model to spread out the model operators to facilitate forced modification of the precision of each operator; in the decision-making process, the specific precision setting of each operator, that is, the mixed precision scheme, is passed in the form of a "01..." string; the forward function of the model is modified to automatically insert a precision conversion function according to the scheme to prevent data type mismatch problems during calculations; and the framework works in a non-invasive way with the deep learning framework.
[0055] like Figure 4 The figure shows the relationship between the sampling-based neural network convergence-aware mixed precision training method and the deep learning framework. It shows the whole process of model training and application from model and dataset input, through sampling-based training methods and optimization schemes, with the support of deep learning framework and GPU hardware, and continuously improves the model through performance feedback.
[0056] GPU hardware is a commonly used high-performance computing hardware used to accelerate the deep learning training process.
[0057] In summary, the embodiment of the present invention uses two-stage pre-run based on rounds and batches to sample the convergence and performance of model training, and optimizes the operator-level optimal mixed precision training strategy that ensures convergence on the current GPU platform. The present invention effectively improves the efficiency of mixed precision training of neural networks while ensuring convergence, and is applicable to mixed precision training on GPU platforms.
[0058] The sampling-based neural network mixed precision training device provided by the present invention is described below. The sampling-based neural network mixed precision training device described below and the sampling-based neural network mixed precision training method described above can refer to each other.
[0059] like Figure 5 The present invention provides a neural network mixed precision training device based on sampling, comprising: A first input module 510, for inputting the neural network model to be trained into the round-based sampling phase; A first sampling module 520 is used to adjust the operators whose influence on convergence exceeds a preset influence threshold according to expert knowledge in the round-based sampling phase, with training rounds as the unit, and run each adjustment scheme for a complete training round, record the loss value of each adjustment scheme, retain the adjustment schemes whose loss values are within the acceptable range, and select the adjustment scheme with the best performance from the retained adjustment schemes as the initial mixed precision scheme; A second input module 530, for inputting an initial mixed precision scheme into a batch-based sampling phase; A second sampling module 540 is used to adjust the adjustable operator precision in the initial mixed precision scheme based on the initial mixed precision scheme in the batch-based sampling phase, traverse all mixed precision schemes and collect performance data in units of training batches, and select the mixed precision scheme with the best performance as the input scheme for post-training according to the performance data; The execution module 550 is used to execute the input scheme of the post-training.
[0060] Figure 6 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 6 As shown, the electronic device may include: a processor (processor) 610 , a communication interface (Communications Interface) 620 , a memory (memory) 630 and a communication bus 640 , wherein the processor 610 , the communication interface 620 , and the memory 630 communicate with each other through the communication bus 640 . The processor 610 can call the logic instructions in the memory 630 to execute the sampling-based neural network mixed precision training method, which includes: inputting the neural network model to be trained into the round-based sampling stage; in the round-based sampling stage, adjusting the operators whose influence on convergence exceeds the preset influence threshold according to expert knowledge in units of training rounds, and running each adjustment scheme for a complete training round, recording the loss value of each adjustment scheme, retaining the adjustment schemes with loss values within an acceptable range, and selecting the adjustment scheme with the best performance from the retained adjustment schemes as the initial mixed precision scheme; inputting the initial mixed precision scheme into the batch-based sampling stage; in the batch-based sampling stage, adjusting the adjustable operator precision in the initial mixed precision scheme based on the initial mixed precision scheme, traversing all mixed precision schemes and collecting performance data in units of training batches, and selecting the mixed precision scheme with the best performance as the input scheme for post-training based on the performance data; and executing the input scheme for post-training.
[0061] In addition, the logic instructions in the above-mentioned memory 630 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0062] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the sampling-based neural network mixed precision training method provided by the above methods, the method comprising: inputting the neural network model to be trained into a round-based sampling stage; in the round-based sampling stage, adjusting the operators whose influence on convergence exceeds a preset influence threshold according to expert knowledge in units of training rounds, and running each adjustment scheme for a complete training round, recording the loss value of each adjustment scheme, retaining the adjustment schemes with loss values within an acceptable range, and selecting the adjustment scheme with the best performance from the retained adjustment schemes as the initial mixed precision scheme; inputting the initial mixed precision scheme into a batch-based sampling stage; in the batch-based sampling stage, adjusting the adjustable operator precision in the initial mixed precision scheme based on the initial mixed precision scheme, traversing all mixed precision schemes and collecting performance data in units of training batches, and selecting the mixed precision scheme with the best performance as the input scheme for post-training based on the performance data; and executing the input scheme for post-training.
[0063] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented by executing the sampling-based neural network mixed precision training method provided by the above methods when the computer program is executed by the processor, the method comprising: inputting the neural network model to be trained into a round-based sampling stage; in the round-based sampling stage, adjusting the operators whose influence on convergence exceeds a preset influence threshold according to expert knowledge in units of training rounds, and running each adjustment scheme for a complete training round, recording the loss value of each adjustment scheme, retaining the adjustment schemes with loss values within an acceptable range, and selecting the adjustment scheme with the best performance from the retained adjustment schemes as the initial mixed precision scheme; inputting the initial mixed precision scheme into a batch-based sampling stage; in the batch-based sampling stage, adjusting the adjustable operator precision in the initial mixed precision scheme based on the initial mixed precision scheme, traversing all mixed precision schemes and collecting performance data in units of training batches, and selecting the mixed precision scheme with the best performance as the input scheme for post-training based on the performance data; and executing the input scheme for post-training.
[0064] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Those of ordinary skill in the art may understand and implement it without creative effort.
[0065] Through the description of the above implementation modes, those skilled in the art can clearly understand that each implementation mode can be implemented by means of software plus a necessary general hardware platform, or of course by hardware. Based on such an understanding, the above technical solution can essentially or in other words be embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiment.
[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A neural network mixed precision training method based on sampling, characterized in that: include: Input the neural network model to be trained into the round-based sampling phase; In the round-based sampling phase, operators whose influence on convergence exceeds a preset influence threshold are adjusted based on expert knowledge in units of training rounds, and each adjustment scheme is run for a complete training round, and the loss value of each adjustment scheme is recorded. The adjustment schemes whose loss values are within the acceptable range are retained, and the adjustment scheme with the best performance is selected from the retained adjustment schemes as the initial mixed precision scheme; Inputting the initial mixed precision scheme into a batch-based sampling phase; In the batch-based sampling phase, taking the initial mixed precision scheme as a benchmark, adjusting the adjustable operator precision in the initial mixed precision scheme, traversing all mixed precision schemes and collecting performance data in units of training batches, and selecting the mixed precision scheme with the best performance as the input scheme for post-training according to the performance data; An input scenario for executing the post-training.
2. The sampling-based mixed precision training method for a neural network according to claim 1, characterized in that: The operator whose influence on convergence exceeds a preset influence threshold is adjusted according to expert knowledge, specifically including: Record operators with conversion tendencies to form a first-stage operation list; Determine the operators in the first-stage operation list that belong to the whitelist and the operators that belong to the blacklist; Check whether the input tensor of the operator in the whitelist meets the tensor core usage conditions, and if so, adjust the precision to 16-bit floating point numbers; The convergence of the operators in the blacklist converted to 16-bit floating point numbers is tested, and if the convergence does not meet the preset convergence condition, the original precision is retained.
3. The sampling-based mixed precision training method for a neural network according to claim 1, characterized in that: The adjustment plan for keeping the loss value within an acceptable range specifically includes: When the loss value is less than 101% of the loss value obtained when all operators are set to 32-bit floating point precision, the adjustment scheme is deemed to have convergence within an acceptable range, and the adjustment scheme with convergence within the acceptable range is retained.
4. The sampling-based mixed precision training method for a neural network according to claim 2, characterized in that: The adjusting of the adjustable operator precision in the initial mixed precision scheme, taking training batches as units, traverses all mixed precision schemes, specifically including: Constructing a search space for other operators except the operators in the operation list of the first stage; The precision of all operators in the mixed precision scheme except the operators in the first stage operation list is set to 32-bit floating point numbers by default; The mixed precision scheme is traversed using the positions of the operators in the first-stage operation list as the split points.
5. The sampling-based mixed precision training method for a neural network according to claim 4, characterized in that: In the process of traversing the mixed precision solution, including: When the precision of two non-adjacent operators in the first stage operation list is the same, the operator between the two operators also maintains the precision of 32-bit floating point numbers; When the precisions of two non-adjacent operators in the first-stage operation list are inconsistent, enumeration is performed to adjust the precision of the operator sandwiched between the two operators.
6. The sampling-based mixed precision training method for a neural network according to claim 1, characterized in that: After adjusting the adjustable operator precision in the initial mixed precision scheme, the method further includes: After adjusting the adjustable operator precision in the initial mixed precision scheme, multiple mixed precision schemes are obtained, each of which defines a new model for testing; During the test, the dataset is divided into batch sizes, and the same amount of computation is allocated to each batch. Then a single batch is extracted for testing, and a different mixed precision scheme is assigned to each batch.
7. The sampling-based mixed precision training method for a neural network according to claim 1, characterized in that: When the neural network model is constructed, the method further includes: Spread out the model operators through sequential programming; Pass the mixed precision scheme in string form; Modify the forward function of the neural network model and automatically insert the precision conversion function.
8. A neural network mixed precision training device based on sampling, characterized in that: include: A first input module, used for inputting the neural network model to be trained into the round-based sampling phase; A first sampling module is used to adjust the operators whose influence on convergence exceeds a preset influence threshold according to expert knowledge in a round-based sampling phase, taking training rounds as units, and run each adjustment scheme for a complete training round, record the loss value of each adjustment scheme, retain the adjustment schemes whose loss values are within an acceptable range, and select the adjustment scheme with the best performance from the retained adjustment schemes as the initial mixed precision scheme; A second input module, configured to input the initial mixed precision scheme into a batch-based sampling phase; A second sampling module is used to adjust the adjustable operator precision in the initial mixed precision scheme based on the initial mixed precision scheme in the batch-based sampling phase, traverse all mixed precision schemes and collect performance data in units of training batches, and select the mixed precision scheme with the best performance as the input scheme for post-training according to the performance data; An execution module is used to execute the post-training input scheme.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the sampling-based neural network mixed precision training method as described in any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the sampling-based neural network mixed precision training method as described in any one of claims 1 to 7 is implemented.