Model Training Method and Related Devices

By dynamically adjusting precision ranges in neural network training based on overflow information, the method addresses training stagnation and efficiency issues, enhancing model training performance.

JP2025522114AActive Publication Date: 2025-07-10HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025501773
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-15
Filing Date
2023-07-12
Publication Date
2025-07-10
Estimated Expiration
2043-07-12

AI Technical Summary

Technical Problem

The manual setting of data format accuracy in neural network training based on experience leads to overdependence on professional ability and results in training stagnation or failure due to overflow issues in low-precision training.

Method used

A method for adjusting the precision range used in model training in real time by recalculating parameters when overflow occurs, allowing for dynamic adjustment of accuracy ranges based on overflow information.

Benefits of technology

Reduces training stagnation and failure by automatically adjusting precision ranges, improving training efficiency and reducing memory usage while minimizing overflow issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522114000001_ABST
    Figure 2025522114000001_ABST
Patent Text Reader

Abstract

This application discloses a model training method. The method is applicable to a dynamic computational graph scenario and also applicable to a static computational graph scenario. The method includes: obtaining training data; using the training data as the input of the model, and in the training process of the model, calculating parameters by using a first accuracy range to obtain a calculated value; and when the calculated value of the parameters overflows the first accuracy range, recalculating the parameters by using a second accuracy range, and performing iterative training on the model one or more times by using the recalculated parameters, where the second accuracy range includes the first accuracy range, or the second accuracy range partially overlaps with the first accuracy range. When the calculated value of the parameters overflows the accuracy range, the accuracy range used in the training process of the model is adjusted in real time, so that the problem of training stagnation caused by overflow in low-precision training can be effectively solved. In addition, the initialization solution means is customized without relying on manual experience, and the training accuracy can be automatically adjusted layer by layer in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a model training method and related devices.

Background Art

[0002] Artificial intelligence (AI) is a theory, method, technology, and application system in which human intelligence is simulated and extended by using digital computers, or a machine that perceives the environment, acquires knowledge, and is controlled by a digital computer to achieve optimal results by using knowledge. A neural network is a dynamic system established manually and using a directed graph as a topological structure. A neural network is an information processing system that processes information by using continuous or discontinuous inputs as status responses and aims to simulate the structure and function of the human brain. After decades of development, artificial neural networks have been widely used in many fields such as pattern recognition, automatic control, signal processing, decision-making support, artificial intelligence, and scientific computing, and have achieved extensive success. In particular, in many fields such as image processing, audio and video processing, and natural language processing, artificial neural networks are in a stage of rapid development and play an indispensable role.

[0003] Currently, in the parameter memory or parameter calculation in the training process of most neural networks, the accuracy of the data format used is mainly set through manual experience. For example, based on experience, the setter determines whether to use 16-bit half-precision floating-point numbers (FP16) or 32-bit single-precision floating-point numbers (FP32) for each network layer.

[0004] However, the setting of the applicable accuracy for the structure of the network layer based on manual experience overly depends on the professional ability of the setter.

Summary of the Invention

[0005] This application provides a model training method and related devices for effectively solving the problem of training stagnation caused by overflow in low-precision training by adjusting the precision range used in the model training process in real time when the calculated value of the parameter overflows the precision range.

[0006] The first aspect of the embodiments of this application provides a model training method. The method is applicable to a dynamic computational graph scenario and is also applicable to a static computational graph scenario. The dynamic computational graph scenario can be understood as updating the computational graph after the network structure of each layer of the model is calculated. The static computational graph scenario can be understood as updating the computational graph after the network structures of all layers of the model are calculated. The main difference is the opportunity to update the computational graph. The calculation of the computational graph is applicable to the method provided in the embodiments of this application. The method can be executed by a training device or can be executed by a component (such as a processor, a chip, or a chip system) of the training device. The method includes the steps of: obtaining training data; using the training data as the input of the model and obtaining a calculated value by calculating the parameters by using a first precision range in the model training process; and when the calculated value overflows the first precision range, recalculating the parameters by using a second precision range and performing iterative training on the model one or more times by using the recalculated parameters, where the second precision range includes the first precision range or the second precision range partially overlaps with the first precision range.

[0007] In the present embodiment of the present application, when the calculated value of a parameter overflows within the first accuracy range during the training process of the model, the parameter is recalculated by using the second accuracy range. That is, the accuracy range is automatically adjusted in real time by using the overflow information of the calculated value of the parameter. As a result, the memory occupied by model training can be reduced, and the model training efficiency can be improved. In this way, problems such as training stagnation or training failure caused by the parameter overflowing within the first accuracy range are reduced. In addition, in the present embodiment of the present application, compared with the prior art method in which the type of the network layer of the model needs to be used to determine whether to use high-precision floating-point numbers or low-precision floating-point numbers, the accuracy range applicable to the parameter can be adjusted in real time by using the overflow information of the parameter, and the overflow problem caused by the calculation of low-precision floating-point numbers is reduced.

[0008] Optionally, in a possible implementation of the first aspect, the model includes a plurality of network structures, and the step of recalculating the parameter by using the second accuracy range includes recalculating the parameter starting from the network structure of the first layer of the model by using the second accuracy range.

[0009] In this possible implementation, when the calculated value of the current layer's network structure overflows, a new accuracy range can be selected to perform recalculation starting from the network structure of the first layer of the model, thereby reducing the problem of calculation error caused by the overflow of the calculated value.

[0010] Optionally, in a possible implementation of the first aspect, the model includes a plurality of network structures, and the step of recalculating the parameter by using the second accuracy range includes recalculating the parameter starting from the current network structure whose calculated value overflows by using the second accuracy range.

[0011] In this possible implementation, if the calculated value of the network structure of the current layer overflows, a new precision range may be selected to perform recalculation on the network structure of the current layer, thereby reducing the problem of calculation errors caused by the overflow of the calculated value.

[0012] Optionally, in a possible implementation of the first aspect, the parameter relates to the loss function of the model, or the parameter relates to the calculation of the model in the forward propagation process, or the parameter relates to the calculation of the model in the backpropagation process.

[0013] Optionally, in a possible implementation of the first aspect, the model includes a plurality of network structures, and the parameter is one or more of the following: intermediate features calculated by the plurality of network structures in the forward propagation process or the value of the loss function of the model, where the intermediate feature is an output feature of any one of the plurality of network structures; and gradients calculated by the plurality of network structures in the backpropagation process, where the gradient includes the gradient of the intermediate feature and / or the weight gradient of the model.

[0014] In this possible implementation, the parameter can be a parameter that needs to be calculated in the forward propagation process or the backpropagation process of the model in the training process, or can be a parameter output by an individual layer of the model, or can be a parameter obtained after the calculations of all layers of the entire model are completed, or the like. This improves the applicable scenarios of the method in the training process. In other words, the method provided in the present embodiment of the present application can be used to perform precision adjustment for all calculations in the training process of the model.

[0015] Optionally, in a possible implementation of the first aspect, when the parameter includes a gradient, the calculated value is a value obtained by dividing the gradient by a scaling factor, and the scaling factor is used to reduce the probability of the gradient overflowing. The method further comprises: updating the scaling factor by using a first factor, where the updated scaling factor replaces the non-updated scaling factor and is used to perform the next iterative training of the model, and the first factor is a positive number less than 1.

[0016] In this possible implementation, if the parameter calculated in the backpropagation process (i.e., the value obtained by dividing the gradient by the scaling factor) overflows, another precision range is used to recalculate the weight gradient, the scaling factor is updated, and the overflow of the calculated value in subsequent iterative training is reduced.

[0017] Optionally, in a possible implementation of the first aspect, the minimum value of the scaling factor is a preset threshold greater than or equal to 1, and the preset threshold is used to reduce the probability of the gradient overflowing.

[0018] In this possible implementation, in the reverse calculation process, the lower limit of the value of the scaling factor is set to reduce the risk of subsequent parameter precision underflow.

[0019] Optionally, in a possible implementation of the first aspect, when the parameter includes intermediate features calculated in the forward propagation process, the above step of recalculating the parameter by using a second precision range comprises: calculating the intermediate features of the overflow layer by using the second precision range, or calculating the intermediate features layer by layer starting from the network structure of the first layer, where the overflow layer is a network structure in a plurality of network structures and the calculated value of the intermediate features overflows the first precision range.

[0020] In this possible implementation, when the intermediate feature overflows the first accuracy range, the second accuracy range can be used to perform recalculation for the overflow layer or to perform recalculation starting from the network structure of the first layer. That is, the accuracy ranges of some layers can be modified, or the accuracy ranges of all layers can be modified. This makes the solution flexible.

[0021] Optionally, in a possible implementation of the first aspect, when the parameter includes the value of the loss function of the model, the above step of recalculating the parameter by using the second accuracy range includes calculating the value of the loss function by using the second accuracy range, or performing calculations layer by layer starting from the network structure of the first layer until the value of the loss function is obtained.

[0022] In this possible implementation, when the value of the loss function overflows the first accuracy range, the second accuracy range can be used to recalculate the value of the loss function or to perform calculations layer by layer starting from the network structure of the first layer until the value of the loss function is obtained. That is, the accuracy ranges of some layers can be modified, or the accuracy ranges of all layers can be modified. This makes the solution flexible.

[0023] Optionally, in a possible implementation of the first aspect, the above step of training the model by using the recalculated parameter includes: in the Nth iteration in the training process of the model, obtaining the number of overflow times of a plurality of network structures in the model based on the first accuracy range, where N is a positive integer greater than or equal to 1; and when the number of overflow times is greater than or equal to the second threshold, determining that the initial accuracy range in the next iterative training process is changed from the first accuracy range to the second accuracy range, and clearing the number of overflow times to zero.

[0024] In this possible implementation, the initial accuracy range is adjusted by recording the number of overflows. When the first accuracy range affects the training of the model, the initial accuracy range is adjusted from the first accuracy range to the second accuracy range to ensure the accuracy of subsequent model training.

[0025] Optionally, in a possible implementation of the first aspect, the number of overflows includes: the number of overflows of a plurality of network structures in the forward propagation process and / or the number of overflows of a plurality of network structures in the backpropagation process.

[0026] In this possible implementation, the determination condition for adjusting the accuracy range (i.e., the determination of the number of overflows) can be the number of overflows in the entire training process or the number of overflows in forward or backward propagation. This improves the applicable range of the method.

[0027] Optionally, in a possible implementation of the first aspect, the parameter relates to the loss function. The loss function can vary according to different training methods of the model. In supervised learning, the loss function is used to represent the difference between the output of the model and the label to which the training data belongs. In unsupervised training, the loss function can be a user-defined function. For example, when the task of the model is a classification task, the loss function is used to represent the difference between the output and the input of the model (or the clustering result or the like). Alternatively, in unsupervised training, it is understood that the output of the model can be fed back to the input of the model. For example, the label is the training data (i.e., the output obtained by the model is sent to another network and the training data is fed back). It can be understood that the loss function is not limited in this embodiment of the present application. The loss function can also be understood as the optimization target function of the model and specifically can be set based on actual requirements.

[0028] The second aspect of the embodiments of the present application provides a training device. The training device is applicable to a dynamic computational graph scenario and also applicable to a static computational graph scenario. The training device includes: an acquisition unit configured to acquire training data; and a calculation unit configured to use the training data as an input to the model and calculate parameters by using a first accuracy range to obtain a calculated value in the training process of the model. The calculation unit is further configured to: when the calculated value overflows the first accuracy range, recalculate the parameters by using a second accuracy range, and perform iterative training on the model one or more times by using the recalculated parameters, where the second accuracy range includes the first accuracy range, or the second accuracy range partially overlaps with the first accuracy range.

[0029] Optionally, in a possible implementation of the second aspect, the model includes a plurality of network structures, and specifically, the calculation unit is configured to recalculate the parameters starting from the first-layer network structure of the model by using the second accuracy range.

[0030] Optionally, in a possible implementation of the second aspect, the calculation unit is specifically configured to recalculate the parameters starting from the current network structure where the calculated value overflows by using the second accuracy range.

[0031] Optionally, in a possible implementation of the second aspect, the model includes a plurality of network structures, and the parameters include one or more of the following: intermediate features calculated by the plurality of network structures in the forward propagation process or the value of the loss function of the model, where the intermediate features are the output features of any one of the plurality of network structures; and gradients calculated by the plurality of network structures in the backpropagation process, where the gradients include the gradients of the intermediate features and / or the weight gradients of the model.

[0032] Optionally, in a possible implementation of the second aspect, when the parameter includes a gradient calculated in the backpropagation process, the calculated value is a value obtained by dividing the gradient by a scaling factor, and the scaling factor is used to reduce the probability that the gradient overflows. The calculation unit further updates the scaling factor by using a first coefficient, where the updated scaling factor replaces the non-updated scaling factor and is used to perform the next iterative training of the model, the first coefficient is a positive number less than 1, and the minimum value of the scaling factor is a preset threshold greater than or equal to 1.

[0033] Optionally, in a possible implementation of the second aspect, when the parameter includes intermediate features calculated in the forward propagation process, the calculation unit specifically calculates the intermediate features of the overflow layer by using a second precision range, or calculates the intermediate features layer by layer starting from the network structure of the first layer, where the overflow layer is a network structure in a plurality of network structures where the calculated value of the intermediate features exceeds the first precision range.

[0034] Optionally, in a possible implementation of the second aspect, when the parameter includes the value of the loss function of the model, the calculation unit specifically calculates the value of the loss function by using a second precision range, or is configured to perform calculations layer by layer starting from the network structure of the first layer until the value of the loss function is obtained.

[0035] Optionally, in a possible implementation of the second aspect, the computing unit is specifically configured to obtain the number of overflows of a plurality of network structures in the model based on the first accuracy range in the Nth iteration in the training process of the model, where N is a positive integer greater than or equal to 1. Specifically, the computing unit is configured to: when the number of overflows is greater than or equal to a second threshold, determine that the initial accuracy range in the next iterative training process is changed from the first accuracy range to the second accuracy range, and erase the number of overflows to zero.

[0036] Optionally, in a possible implementation of the second aspect, the number of overflows includes: the number of overflows of a plurality of network structures in the forward propagation process and / or the number of overflows of a plurality of network structures in the backward propagation process.

[0037] The third aspect of the present application provides a training device including a processor. The processor is coupled to a memory, and the memory is configured to store a program or instructions. When the program or instructions are executed by the processor, the training device can implement the method in any one of the first aspect or the possible implementations of the first aspect.

[0038] The fourth aspect of the present application provides a computer-readable medium. The computer-readable medium stores a computer program or instructions. When the computer program or instructions are executed on a computer, the computer can execute the method in any one of the first aspect or the possible implementations of the first aspect.

[0039] The fifth aspect of the present application provides a computer program product. When the computer program product is executed on a computer, the computer can execute the method in any one of the first aspect or the possible implementations of the first aspect.

[0040] For the technical effects brought about by the second aspect, the third aspect, the fourth aspect, the fifth aspect, or any conceivable implementation thereof, reference may be made to the technical effects brought about by the first aspect, or different conceivable implementations of the first aspect. Details are not described herein.

[0041] From the above technical solutions, it can be seen that this application has the following features: When the calculated value of a parameter overflows within the first accuracy range during the training process of the model, the parameter is recalculated by using the second accuracy range. That is, the accuracy range is automatically adjusted in real time by using the overflow information of the calculated value of the parameter. As a result, the memory occupied by model training can be reduced, and the model training efficiency can be improved. In this way, problems such as training stagnation or training failure caused by the parameter overflowing within the first accuracy range are reduced. In addition, in the embodiments of this application, compared with the conventional art method that the type of the network layer of the model needs to be used to determine whether to use high-precision floating-point numbers or low-precision floating-point numbers, the accuracy range applicable to the parameter can be adjusted in real time by using the overflow information of the parameter, and the overflow problem caused by the calculation of low-precision floating-point numbers is reduced.

Brief Description of Drawings

[0042]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6A

Figure 6B

Figure 7A

Figure 7B

Figure 7C

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0043] The present application provides a model training method and related devices for effectively solving the problems of training stagnation or training failure caused by overflow in low-precision training by adjusting the precision range used in the model training process in real time when the calculated value overflows the precision range. In addition, the requirements for the network mixed-precision initialization solution means are low, the initialization solution means is customized without relying on manual experience, and the training precision can be automatically adjusted layer by layer in real time.

[0044] To facilitate understanding, the main related terms and concepts in the embodiments of the present application will be described below first. 1. Neural Network A neural network may include neural units. A neural unit may be an arithmetic unit that uses X s and the bias b as inputs, and the output of the arithmetic unit may be as follows:

Number

[0045] Here, s = 1, 2, …, or n, where n is a natural number greater than 1, and W s is the weight of X s and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce non-linear characteristics into the neural network and convert the input signal in the neural unit into an output signal. The output signal of the activation function can be used as the input of the next layer. The activation function can be the ReLU function. A neural network is a network formed by connecting many single neural units together. Specifically, the output of a neural unit can be the input of another neural unit. The input of each neural unit is connected to the local receptive field of the previous layer and can extract the features of the local receptive field. The local receptive field can be a region containing multiple neural units.

[0046] The operation in each layer of the neural network can be explained by using the mathematical formula y = a(Wx + b). From a physical perspective, the operation in each layer of the neural network can be understood as completing the transformation from the input space (a set of input vectors) to the output space (i.e., from the row space of the matrix to the column space) by performing five operations on the input space. The five operations are as follows: 1. Dimension increasing / dimension reduction; 2. Scaling up / scaling down; 3. Rotation; 4. Translation; and 5. “Bending”. Calculations 1, 2, and 3 are executed by Wx, calculation 4 is executed by +b, and calculation 5 is executed by a(). Since the objects to be classified are not single things but types of things, the word "space" is used here for the sake of expression. Space is the set of all individuals of this type of thing. W is a weight vector, and each value in the vector represents the weight value of one neuron in this layer of the neural network. The vector W determines the spatial transformation from the input space to the output space described above. In other words, the weight W in each layer controls how the space is transformed. The purpose of training a neural network is to ultimately obtain the weight matrix (the weight matrix formed by the vectors W in multiple layers) in all layers of the trained neural network. Therefore, the training process of a neural network is basically a way of learning to control spatial transformation, more specifically, learning the weight matrix.

[0047] A neural network is also called an Artificial Neural Network (ANN) and is a dynamic system that uses a manually established directed graph as its topological structure. A neural network is an information processing system that processes information by using continuous or discontinuous inputs as status responses and aims to simulate the structure and functions of the human brain. After decades of development, artificial neural networks have been widely used in many fields such as pattern recognition, automatic control, signal processing, decision-making support, artificial intelligence, and scientific computing, achieving extensive success. Generally, a network includes an input layer, a hidden layer, and an output layer.

[0048] 2. Loss Function In the process of training a neural network, the output of the neural network is expected to be as close as possible to the truly expected predicted value. Therefore, the predicted value of the current network can be compared with the truly expected target value. Next, the weight matrix of each layer of the neural network is updated based on the difference between the two values. (Of course, before the update is executed for the first time, there is usually an initialization process. Specifically, parameters are pre-configured for each layer of the neural network.) For example, if the predicted value of the network is large, the weight matrix is adjusted to make the predicted value smaller, and the adjustment is continuously executed until the neural network can obtain the truly expected target value through prediction. Therefore, it is necessary to define in advance "how to obtain the difference between the predicted value and the target value through comparison". This is the loss function or objective function. The loss function and objective function are important equations for measuring the difference between the predicted value and the target value. The loss function is used as an example. A high output value (loss) of the loss function indicates a large difference. Therefore, the training of a neural network is a process of minimizing the loss as much as possible.

[0049] 3. Forward Propagation The forward propagation of a neural network is a calculation process from the input layer to the hidden layer and then to the output layer. Starting from the input layer, the output (activation value) of the previous layer is used as the input of the next layer based on the topological structure of the network. The output of each layer is calculated layer by layer until the last output layer. This process is called the forward propagation of the network.

[0050] 4. Backward Propagation Backpropagation of a neural network is an abbreviation of "error backpropagation" and is a general method for training an artificial neural network in combination with an optimization method (such as gradient descent). This method is used to calculate the gradient of the loss function for all weights in the network. The gradient is fed back to the optimization method to update the weights and minimize the loss function.

[0051] 5. Computational Graph A computational graph usually uses arrows to indicate the order of computation. For example, taking the function y = 5(a + bc) as an example, the value of bc is first calculated and stored in the variable i. Next, a + i is calculated and a + i is stored in the variable j. Next, 5 * j is calculated and the calculation result of y is obtained.

[0052] 6. Precision Range The precision range is the precision range of the data type used by the computer. The precision range can be a specific precision or a dynamic range of precision. Below, an example where the common data type in a neural network is a floating-point (FP) number is used for illustration. In actual applications, the data type can alternatively be an integer (int), for example, int8 or int16.

[0053] FP is mainly used to represent decimal numbers and usually includes three parts: a sign, an exponent, and a mantissa. The sign can be a 1-bit (bit) representing positive or negative, and the exponent and mantissa can be multiple bits (bits). Generally, the mantissa represents the precision, and the exponent is used to represent the dynamic range (referred to as the precision range in the embodiments of the present application) that the precision can reach. When a floating-point number represents a decimal number, the decimal number in decimal notation cannot be accurately converted to a binary number and is truncated when stored in a computer with fixed bits. Therefore, when a floating-point number represents a decimal number, precision loss can occur.

[0054] Floating-point numbers generally can include three formats, specifically described below, namely, half-precision floating-point numbers, single-precision floating-point numbers, and double-precision floating-point numbers.

[0055] Half-precision floating-point numbers are a binary data type used by computers, occupy 16 bits (i.e., 2 bytes) in computer memory, and can also be abbreviated as FP16. The absolute value range of the values that can be represented by half-precision floating-point numbers is approximately [6×10 -8 ,65504]. The precision of FP16 is 2 -10 .

[0056] Single-precision floating-point numbers are a binary data type used by computers, occupy 32 bits (i.e., 4 bytes) in computer memory, and can also be abbreviated as FP32. The absolute value range of the values that can be represented by single-precision floating-point numbers is approximately [1.4×10 -45 ,1.7×10 38 . The precision of FP32 is 2 -23 .

[0057] Double-precision floating-point numbers are a binary data type used by computers, occupy 64 bits (i.e., 8 bytes) in computer memory, and can also be abbreviated as FP64. Double-precision floating-point numbers can represent 15 or 16 significant digits in decimal. The absolute value range of the values that can be represented by double-precision floating-point numbers is approximately [2.23×10 -308 ,1.80×10 38 . The precision is 2 of P64 -52 .

[0058] To present the above three types of floating-point numbers with different precisions more intuitively, the structures of the three types of floating-point numbers are shown in Table 1. Table 1​

Table 1

[0059] In the 16 bits occupied by FP16, the sign occupies 1 bit, the exponent occupies 5 bits, and the mantissa occupies 10 bits; in the 32 bits occupied by FP32, the sign occupies 1 bit, the exponent occupies 8 bits, and the mantissa occupies 23 bits; in the 64 bits occupied by FP64, the sign occupies 1 bit, the exponent occupies 11 bits, and the mantissa occupies 52 bits.

[0060] In actual applications, it can be understood that in order to represent floating-point numbers with higher precision, the format of floating-point numbers, storage formats that occupy more bits, and the like can be further extended. For example, there is a floating-point number that occupies 128 bits (which can be abbreviated as FP128). This is not particularly limited herein.

[0061] 7. Overflow The overflow in the embodiments of the present application includes overflow and underflow. Overflow means that the absolute value of the calculated value is excessively large and exceeds the maximum value that can be represented by the precision range. Underflow means that the absolute value of the calculated value is excessively small and is smaller than the positive value closest to 0 that can be represented by the precision range.

[0062] Overflow can include storage overflow and calculation overflow. The embodiments of the present application are mainly applicable to the scenario of calculation overflow.

[0063] For example, when the first precision range is 1 to 5, the values of the two parameters are 3, and the two parameters are assumed to be within the first precision range when stored. However, the parameter calculation can overflow the first precision range. For example, the above two parameters are added (3 + 3 = 6), and 6 is greater than the maximum value 5 that can be represented by the first precision range, that is, the calculated value of the parameter overflows the first precision range. It can be understood that the example is not intended to limit the boundary values of the precision range.

[0064] Specifically, FP16 is used as an example to explain the overflow situation. Overflow includes that when the calculated value is a positive number, the calculated value is greater than the largest positive number that can be represented by FP16; or when the calculated value is a negative number, the calculated value is less than the smallest negative number that can be represented by FP16. Underflow includes that when the calculated value is a positive number, the calculated value is less than the smallest positive number that can be represented by FP16; or when the calculated value is a negative number, the calculated value is greater than the largest negative number that can be represented by FP16.

[0065] 8. Mixed Precision (MP) Currently, most models are trained by using 32-bit single-precision floating-point numbers (FP32). In the mixed-precision training method, model training is performed by using half-precision or even lower precision and single precision, thereby reducing the memory required for model training. In addition, since low-precision operations are faster than single-precision operations, the hardware efficiency is further improved.

[0066] From the above, it can be seen that the key point of mixed precision is a policy for ensuring accuracy and improving training efficiency by setting which parts of the network to train using high precision and which parts to train using low precision. In other words, the key point of mixed precision is how to specifically combine single precision and high precision for training.

[0067] Currently, the mixed-precision training method mainly includes the following two methods.

[0068] In the first method, the calculation precision is specified based on the type of each layer in the neural network. Some types of layers use high precision for calculation, and some types of layers use low precision for calculation.

[0069] In the second method, the precision is dynamically selected by determining whether the quantization error exceeds a threshold. The quantization error can be measured at different points in the network or over time as the training is executed. For example, the quantization error is calculated by comparing the training result with a baseline value. The baseline value can be determined by using various methods, such as training the same network using high precision by using full-precision floating-point values, repeatedly calculating subsets by using high precision, and analyzing or sampling data statistics for related calculations.

[0070] However, in the first method, the same type of layer has different requirements for calculation precision in different networks or in different training phases of the same network. Therefore, it is clear that specifying the applicable precision for calculation based on the type is not sufficiently flexible or intelligent.

[0071] In the second method, the adjustment of the training precision depends on the precision baseline value. The baseline value needs to be repeatedly calculated by using high precision or obtained by training the same network by using full-precision floating-point values. As a result, the solution is still not fully automated and can significantly increase the amount of calculation for network training due to the construction of the baseline value. This is contrary to the purpose of using low-precision training to reduce calculations and speed up training.

[0072] In view of this, an embodiment of the present application provides a model training method and related device for adjusting the precision range used in the model training process in real time when the parameters of the calculated values in the neural network overflow the precision range. As a result, the training stagnation caused by overflow in low-precision training in the prior art can be effectively solved. In addition, the requirements for the network mixed-precision initialization solving means are low, and the initialization solving means is customized without relying on manual experience (for example, manual experience is not required to perform precision adjustment in different layers of each network), and the training precision can be automatically adjusted layer by layer in real time in the training process based on whether overflow occurs.

[0073] Referring to the accompanying drawings, the model training method and related device provided in the embodiments of the present application will be described in detail below.

[0074] First, the system architecture provided in the embodiments of the present application will be described.

[0075] Please refer to FIG. 1. Embodiments of the present invention provide a system architecture 100. As shown in the system architecture 100, the data collection device 160 is configured to collect training data. The training data in the present embodiment of the present application may include one or more of images, voices, texts, and the like. The training data is stored in the database 130, and the training device 120 obtains the target model / rule 101 through training based on the training data maintained in the database 130. The target model / rule 101 can be used to implement computer vision tasks (e.g., classification, segmentation, detection, and image generation). The target model / rule 101 in the present embodiment of the present application may specifically be a neural network or the like. It should be noted that in actual applications, the training data maintained in the database 130 is not necessarily collected by the data collection device 160 and may be received from another device. In addition, it should be noted that the training device 120 does not necessarily train the target model / rule 101 completely based on the training data maintained in the database 130, but may obtain training data from the cloud or another location to execute model training. The above description should not be construed as a limitation to the present embodiment of the present application.

[0076] The target model / rule 101 obtained through training by the training device 120 can be applied to different systems or devices. For example, it can be applied to the execution device 110 shown in FIG. 1. The execution device 110 can be a terminal, such as a mobile phone terminal, a tablet computer, a notebook computer, an augmented reality (AR) device / virtual reality (VR) device, or an in-vehicle terminal. Of course, the execution device 110 can alternatively be a server, a cloud, or the like. In FIG. 1, the I / O interface 112 is configured for the execution device 110 and is configured to exchange data with external devices. The user can input data into the I / O interface 112 by using the client device 140. The input data corresponds to the training data. The input data in the present embodiment of the present application can also include one or more of images, sounds, texts, and the like. In addition, the input data can be input by the user, or uploaded by the user by using a photographing device, or of course, can be derived from a database or the like. This is not particularly limited herein.

[0077] The preprocessing module 113 is configured to perform preprocessing (e.g., splitting, selection, and conversion) based on the input data received by the I / O interface 112. For example, the input data is split and a plurality of data blocks (patches) are obtained.

[0078] In the related processing procedures where the execution device 110 preprocesses the input data or the calculation module 111 of the execution device 110 executes calculations, the execution device 110 can call data, code, and the like in the data storage system 150 and implement corresponding processing, or store data, instructions, and the like obtained through the corresponding processing in the data storage system 150.

[0079] Finally, the I / O interface 112 returns the processing result (e.g., classification result, segmentation result, or detection result) to the client device 140 and provides the processing result to the user.

[0080] Note that the training device 120 may generate corresponding target models / rules 101 based on different training data for different targets or different tasks. The corresponding target models / rules 101 can be used to achieve the above-mentioned target or complete the above-mentioned task and provide the necessary results to the user.

[0081] In the case shown in FIG. 1, the user may manually provide the input data. The input data may be manually provided by using the screen provided by using the I / O interface 112. In another case, the client device 140 may automatically send the input data to the I / O interface 112. If the client device 140 needs to obtain approval from the user to automatically send the input data, the user may set the corresponding permission on the client device 140. The user may view the results output by the execution device 110 on the client device 140. The results may be presented in a specific manner of display, sound, action, or the like. Alternatively, the client device 140 is used as a data collection end, collects the input data input to the I / O interface 112 and the output values output from the I / O interface 112 shown in the figure as new sample data, and may store the new sample data in the database 130. Of course, the client device 140 may alternatively not perform the collection. Instead, the I / O interface 112 directly stores the input data input to the I / O interface 112 and the output values output from the I / O interface 112 shown in the figure as new sample data in the database 130.

[0082] Note that FIG. 1 is merely a diagram of a system architecture according to an embodiment of the present invention. The positional relationships among the devices, components, modules, and the like shown in the figure do not constitute a limitation. For example, in FIG. 1, the data storage system 150 is an external memory with respect to the execution device 110. In another case, the data storage system 150 may alternatively be arranged in the execution device 110.

[0083] Hereinafter, the hardware structure of the chip provided in the embodiment of the present application will be described.

[0084] FIG. 2 shows the hardware structure of a chip according to an embodiment of the present invention. The chip includes a neural network processing unit 20. The chip may be arranged in the execution device 110 shown in FIG. 1 and is configured to complete the calculation operation of the calculation module 111. Alternatively, the chip may be arranged in the training device 120 shown in FIG. 1 and is configured to complete the training operation of the training device 120 and output the target model / rule 101.

[0085] The neural network processing unit 20 can be any processor suitable for large-scale exclusive OR operation processing, for example, a neural-network processing unit (NPU), a tensor processing unit (TPU), or a graphics processing unit (GPU). The NPU is used as an example. The neural network processing unit 20 functions as a coprocessor and is mounted on a host central processing unit (CPU) (host CPU). The host CPU assigns tasks. The core part of the NPU is the arithmetic circuit 203, and the controller 204 controls the arithmetic circuit 203 to extract data in the memory (weight memory or input memory) and execute the operation.

[0086] In some implementations, the arithmetic circuit 203 includes a plurality of processing engines (PEs). In some implementations, the arithmetic circuit 203 is a two-dimensional systolic array. The arithmetic circuit 203 can alternatively be a one-dimensional systolic array or another electronic circuit capable of performing arithmetic operations such as multiplication and addition. In some implementations, the arithmetic circuit 203 is a general matrix processor.

[0087] For example, assume there are an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit 203 fetches data corresponding to the matrix B from the weight memory 202 and buffers the data for each PE in the arithmetic circuit. The arithmetic circuit fetches the data of the matrix A from the input memory 201, performs matrix operations using the matrix B, and stores the obtained partial result or the obtained final result of the matrix in the accumulator 208.

[0088] The vector calculation unit 207 can perform further processing on the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, or value comparison. For example, the vector calculation unit 207 can be configured to perform network calculations such as pooling, batch normalization, or local response normalization in non-convolutional / non-FC layers in a neural network.

[0089] In some implementations, the vector calculation unit 207 can store the processed output vector in the integrated memory 206. For example, the vector calculation unit 207 can apply a non-linear function to the output of the arithmetic circuit 203, such as a vector of cumulative values, to generate activation values. In some implementations, the vector calculation unit 207 generates normalization values, combination values, or both. In some implementations, the processed output vector can be used as an activation input to the arithmetic circuit 203, for example, in subsequent layers in a neural network.

[0090] The integrated memory 206 is configured to store input data and output data.

[0091] Regarding the weight data, the direct memory access controller (DMAC) 205 transfers the input data in the external memory to the input memory 201 and / or the integrated memory 206, stores the weight data in the external memory in the weight memory 202, and stores the data in the integrated memory 206 in the external memory.

[0092] The bus interface unit (BIU) 210 is configured to implement the interaction between the host CPU, the DMAC, and the instruction fetch buffer 209 through the bus.

[0093] The instruction fetch buffer 209 connected to the controller 204 is configured to store the instructions used by the controller 204.

[0094] The controller 204 is configured to call the instructions buffered in the instruction fetch buffer 209 to control the operation process of the arithmetic accelerator.

[0095] Generally, the integrated memory 206, the input memory 201, the weight memory 202, and the instruction fetch buffer 209 are each on-chip memories. The external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (abbreviated as DDR SDRAM), a high bandwidth memory (HBM), or another readable and writable memory.

[0096] Hereinafter, with reference to the accompanying drawings, the model training method and data processing method in the embodiments of the present application will be described in detail.

[0097] First, the application scenarios of the model training method provided in the embodiments of the present application will be described. The method is applicable to the dynamic computational graph scenario and also applicable to the static computational graph scenario. The dynamic computational graph scenario can be understood as updating the computational graph after the network structure of each layer of the model is calculated. The static computational graph scenario can be understood as updating the computational graph after the network structures of all layers of the model are calculated. The main difference is the opportunity to update the computational graph. The calculation of the computational graph is applicable to the model training method (or referred to as the calculation accuracy adjustment method) provided in the embodiments of the present application.

[0098] Hereinafter, with reference to FIG. 3, the model training method in the embodiments of the present application will be described in detail. The method can be executed by a training device or by a component of the training device (for example, a processor, a chip, or a chip system). The training device can be a cloud device or a terminal device. For example, the training device can be a device such as a computer or a server having a robust computing ability to execute the model training method, or can be a system including a cloud device and a terminal device. For example, the training method can be executed by the training device 120 in FIG. 1 and the neural network processing unit 20 in FIG. 2.

[0099] Optionally, the model training method can be processed by the CPU or by both the CPU and the GPU; or the GPU may not be used, and another processor suitable for neural network calculation may be used. This is not limited in the present application.

[0100] Figure 3 is a schematic flowchart of the model training method according to the embodiment of the present application. The method may include steps 301 to 303. Hereinafter, steps 301 to 303 will be described in detail.

[0101] Step 301: Obtain training data.

[0102] In the present embodiment of the present application, the training device may obtain training data in a plurality of ways. The training data may be transmitted by another device (such as a server or a service device) and then received, selected from a database, taken by a user, or obtained in another way. This is not particularly limited herein.

[0103] The training data in the present embodiment of the present application may include one or more of images, voices, texts, and the like. Specifically, the training data relates to the scenario to which the model is applied. For example, when the function of the model is speech recognition, the specific form of the training data may be speech data or the like. In another example, when the function of the model is image classification, the specific form of the training data may be image data or the like. In another example, when the function of the model is speech prediction, the specific form of the training data may be text data or the like. It can be understood that the above-mentioned multiple cases are merely examples and not necessarily in a one-to-one correspondence. For example, for speech recognition, the specific form of the training data may be image data, text data, or the like (for example, when the model is applied to a scenario in the education field of displaying images and playing voices, the function of the model is to recognize the voice corresponding to the image, and the specific form of the training data may be image data). In actual applications, there are other scenarios. For example, when the model is applied to a video recommendation scenario, the training data may be word vectors corresponding to the video or the like. In some application scenarios, the training data may alternatively include data of different modalities at the same time. For example, in an autonomous driving scenario, the training data may include image / video data collected by a camera and may further include voice / text data or the like for sending commands by the user. The specific form or type of the training data is not limited in the present embodiment of the present application.

[0104] When the training of the model is supervised learning, it can be understood that the training data obtained in this step is training data that holds labels. When the training of the model is unsupervised training, the training data obtained in this step is training data that does not hold labels.

[0105] Step 302: Use the training data as the input of the model, and in the training process of the model, calculate the parameters by using the first accuracy range and obtain the calculated values.

[0106] After obtaining the training data, the training device uses the training data as the input of the model, calculates the parameters by using the first accuracy range, and obtains the calculated values. The parameters relate to the loss function of the model. In supervised learning, the loss function is used to represent the difference between the output of the model and the label to which the training data belongs. In unsupervised training, the loss function can be a user-defined function. For example, when the task of the model is a classification task, the loss function is used to represent the difference between the output of the model and the input (or clustering result or the like). Alternatively, in unsupervised training, it is understood that the output of the model can be fed back to the input of the model. For example, the label is the training data (that is, the output obtained by the model is sent to another network and the training data is fed back). It can be understood that the loss function is not limited in the present embodiment of the present application. The loss function can also be understood as the optimization objective function of the model, and specifically, it can be set based on actual requirements.

[0107] The parameters in the present embodiment of the present application include the following possibilities: the parameters relate to the loss function of the model, the parameters relate to the calculation of the model in the forward propagation process, and the parameters relate to the calculation of the model in the backpropagation process, and may include one or more of them.

[0108] Optionally, the model includes a plurality of network structures. In addition, the model in the present embodiment of the present application is specifically an artificial neural network. The specific number or specific structure of the layers included in the artificial neural network can be set based on actual requirements. It is not limited here.

[0109] The accuracy ranges (e.g., the first accuracy range and the second accuracy range) in the present embodiment of the present application can be the accuracy ranges of data types (e.g., int or float). For the sake of facilitating the subsequent description, in the present embodiment of the present application, only an example where the accuracy range is the FP accuracy range is used in the description. Of course, in actual applications, the accuracy range can alternatively be the accuracy range of another data type (e.g., int).

[0110] Optionally, the first accuracy range is the dynamic reachable range of FP16 precision. For the description of FP16 and the accuracy range, please refer to the description in the above related terms. Details are not described in this specification.

[0111] In the present embodiment of the present application, the accuracy ranges used by the network structures of all layers of the model can be the same or different. In other words, the first accuracy range can be the accuracy range refined for each layer of the model, or the same accuracy range used by all layers of the model. That is, later replacing the first accuracy range with the second accuracy range for parameter recalculation can be layer-granularity accuracy range adjustment (i.e., only the accuracy range of a specific layer is adjusted, the accuracy ranges of non-specific layers are not adjusted, and the specific layer can also become an overflow layer), or accuracy range adjustment for the entire model (i.e., all layers).

[0112] In a conceivable implementation, the network structures of all layers of the model use the same accuracy range, and in this case, the first accuracy range is the same accuracy range. Alternatively, the present embodiment of the present application is understood to be applicable to a training scenario using single precision.

[0113] In another possible implementation, at least two layers of the network structure in the model use different accuracy ranges. For example, the model is trained by using mixed precision. In this case, the first accuracy range is low precision or high precision in mixed precision. Generally, in mixed precision training, the probability of a high-precision overflow problem is relatively low by default, and the first accuracy range may specifically be a low-precision range in mixed precision. Alternatively, it should be understood that the present embodiment of the present application is applicable to a training scenario using mixed precision. That is, the accuracy adjustment is then performed only for specific layers.

[0114] In the present embodiment of the present application, the parameters have multiple cases. These cases will be described separately below.

[0115] In the first case, the parameter is a parameter in the forward propagation process.

[0116] Thus, the parameter can be an intermediate feature or a value of a loss function calculated by a plurality of network structures in the forward propagation process. The intermediate feature can also be understood as an activation value obtained by a plurality of network structures using an activation function. The value of the loss function can also be understood as a difference value between the model output and the label calculated by the network structures of all layers after the forward propagation is completed.

[0117] For example, the model shown in FIG. 4 is used as an example. The model includes a three-layer network structure, which are respectively an input layer (including three neurons), a hidden layer (including four neurons), and an output layer (including two neurons).

Number

Number

Number

Number

[0118] Equation 1:

Number

Number

[0119] The activation value can be calculated layer by layer by using Equation 2, and finally, the output f(x) of the model can be obtained based on the input X. The value of the loss function is calculated based on the difference between the output f(x) of the model and the label value Y of X. The loss function can be Mean Square Error (MSE) loss, Mean Absolute Error loss, Cross Entropy loss, Hinge loss, or the like, and specifically, it can be set based on actual requirements. The structure of the loss function is not limited.

[0120] It can be understood that the model structure shown in Figure 4 is merely an example and does not constitute a limitation on the model mentioned in the present embodiment of the present application. In addition, Equations 1 and 2 are merely examples of the expression forms in the forward propagation process. In actual applications, the forward propagation process can alternatively have another expression form. This is not particularly limited herein.

[0121] In the above example, the parameter may include intermediate features and / or the value of the loss function. The a l calculated at each layer can be understood as an intermediate feature, and the difference between f(x) and Y is the value of the loss function.

[0122] In the second case, the parameter is the parameter in the backpropagation process.

[0123] Thus, the parameter can be the gradient calculated by multiple network structures in the backpropagation process. The gradient includes one or more of the following: the gradient of the intermediate feature, the gradient of the weights in the model, the gradient of the loss, and the like.

[0124] The backpropagation process can be understood as a process of continuously adjusting the weights by using the value of the loss function.

[0125] For example, continuing with the above example, the backpropagation process uses an optimizer and a preset learning rate to minimize the loss function to obtain optimal parameters (e.g., weights

Number

[0126] It can be understood that the above two cases of the parameter are just examples. In actual applications, the parameter may alternatively have other cases. This specification does not particularly limit this.

[0127] Step 303: If the calculated value overflows the first accuracy range, recalculate the parameter by using the second accuracy range, and perform iterative training on the model multiple times by using the recalculated parameter.

[0128] If the calculated value obtained by calculating the parameters by the training device using the first accuracy range in step 302 overflows the first accuracy range, the parameters are recalculated using the second accuracy range, and iterative training is performed on the model one or more times by using the recalculated parameters. The second accuracy range includes the first accuracy range, or the second accuracy range partially overlaps with the first accuracy range.

[0129] When precision overflow occurs in an iteration, compared with the prior art where the training data used in the overflow is discarded and the problem of training stagnation occurs, the method provided in the present embodiment of the present application can continue training by adjusting the accuracy range in a timely manner, and as a result, the model training efficiency can be improved.

[0130] In the present embodiment of the present application, the second accuracy range includes the first accuracy range, or the second accuracy range partially overlaps with the first accuracy range.

[0131] For example, the second accuracy range includes the first accuracy range. An example where the first accuracy range is FP16 and the second accuracy range is FP32 is used. The first accuracy range can be (-65504, 65504), and the second accuracy range can be [-1.7×10 38 , 1.7×10 38 .

[0132] For example, the second accuracy range partially overlaps with the first accuracy range. An example where the first accuracy range is FP16 and the second accuracy range is int16 is used. The first accuracy range is (-65504, 65504), and the second accuracy range can be integers from -32768 to 32767.

[0133] It can be understood that the examples do not limit the boundary values of the accuracy range.

[0134] Optionally, the first accuracy range and the second accuracy range may alternatively be concepts of relative high and low ranges. When the first accuracy range is a high-accuracy range (hereinafter referred to as high accuracy), the second accuracy range is a low-accuracy range (hereinafter referred to as low accuracy). When the first accuracy range is low accuracy, the second accuracy range is high accuracy. Generally, overflow problems occur in low accuracy.

[0135] In the present embodiment of the present application, only the example where the first accuracy range is low accuracy and the second accuracy range is high accuracy is used for the purpose of explanation. Of course, in actual applications, adjustments between high accuracy and low accuracy, or adjustments from low accuracy to high accuracy can also be implemented by using the method provided in the present embodiment of the present application.

[0136] Furthermore, recalculating the parameters by using the second accuracy range may include multiple cases, and the parameters can be recalculated starting from the network structure of the first layer of the model by using the second accuracy range. Alternatively, the parameters can be recalculated starting from the current network structure where the calculated value overflows by using the second accuracy range. This is not particularly limited herein.

[0137] In a possible implementation, when the parameters include intermediate features calculated in the forward propagation process, the above steps of recalculating the parameters by using the second accuracy range specifically include: calculating the intermediate features of the overflow layer by using the second accuracy range, or calculating the intermediate features layer by layer starting from the network structure of the first layer, where the overflow layer is the network structure in a plurality of network structures where the calculated value of the intermediate features overflows the first accuracy range.

[0138] In another possible implementation, when the parameter includes the value of the loss function of the model, the above step of recalculating the parameter by using the second precision range specifically includes calculating the value of the loss function by using the second precision range, or starting from the network structure of the first layer and performing calculations layer by layer until the value of the loss function is obtained.

[0139] For example, the model includes five layers of network structure. If the calculated value obtained by performing the parameter calculation of the fourth layer by using the first precision range overflows the first precision range, the second precision range can be used to perform the recalculation from the first layer to the fourth layer. Alternatively, the second precision range can be used to recalculate the parameters of the fourth layer.

[0140] Optionally, when the parameter includes a gradient, the calculated value is the value obtained by dividing the gradient by a scaling factor, and the scaling factor is used to reduce the probability of the gradient overflowing. The method further comprises: updating the scaling factor by using a first factor, where the updated scaling factor replaces the non-updated scaling factor and is used to perform the next iterative training of the model, and the first factor is a positive number less than 1.

[0141] Regarding overflow, refer to the explanations in the above related terms. Details are not described in this specification. The calculated value overflowing the first precision range can also be understood as the calculated value of the parameter not being able to be accurately represented by using the first precision range. This can cause subsequent rounding errors or overflow errors caused by a narrow precision range.

[0142] If the calculated value of a parameter overflows the first precision range, the parameter is recalculated using a second precision range different from the first precision range, and iterative training is performed on the model one or more times by using the updated parameter. For specific training processes, refer to the descriptions in the forward propagation process and the backpropagation process. For example, the loss function in forward propagation is calculated by using the parameter calculated by using the second precision range. In order to achieve the goal that the value of the loss function is smaller than the threshold, iterative training is performed on the model one or more times, and a trained model is obtained.

[0143] As described in step 302, note that recalculating using the second precision range can be a layer - granularity precision range adjustment or a precision range adjustment for the entire model (i.e., all layers). That is, if the calculation precision of all layers is initialized to the same precision range, the recalculation in this step is for all layers (or is understood as a calculation precision adjustment for all layers). If the calculation precision of all layers is initialized to different precision ranges (i.e., mixed - precision calculation), the recalculation in this step is for a specific layer (or is understood as a calculation precision adjustment for a specific layer). The calculation precision of a specific layer overflows the first precision range.

[0144] In addition, steps 301 - 303 in this embodiment can be executed one or more times. After updates are executed one or more times in the training process of the model, steps 301 - 303 can be executed once. Alternatively, steps 301 - 303 can be executed again when a preset cycle or a preset number of times is met.

[0145] In the present embodiment of the present application, when the calculated value of a parameter overflows the first accuracy range in the training process of a model, the parameter is recalculated by using the second accuracy range. That is, the accuracy range is automatically adjusted in real time by using the overflow information of the calculated value of the parameter. As a result, the memory occupied by model training can be reduced, and the model training efficiency can be improved. In this way, problems such as training stagnation caused by the parameter calculation overflowing the first accuracy range are reduced. In addition, in the present embodiment of the present application, compared with the method in the prior art in which the type of the network layer of the model needs to be used to determine whether to use high-precision floating-point numbers or low-precision floating-point numbers, the accuracy range applicable to the parameter can be adjusted in real time by using the overflow information of the parameter, and the overflow problem caused by the calculation of low-precision floating-point numbers is reduced.

[0146] Optionally, in the embodiment shown in FIG. 3, when the calculated value of the parameter does not overflow the first accuracy range, the value of the loss function obtained in the forward propagation process is multiplied by a scaling factor, and the gradient calculation in the backward propagation process is performed by using the first accuracy range. When the value obtained by dividing the gradient by the scaling factor overflows the first accuracy range, the scaling factor is updated by using the first factor, and the recalculation is performed in the forward propagation process and the backward propagation process by using the second accuracy range. The first factor is a positive number smaller than 1, and the updated scaling factor is used to replace the unupdated scaling factor to perform the next iterative training of the model.

[0147] To reduce the risk of subsequent parameter precision underflow, the lower limit of the value of the scaling factor can be limited. For example, the minimum value of the scaling factor is a preset threshold greater than or equal to 1.

[0148] Optionally, since the initial accuracy range (e.g., the first accuracy range) can be set for the model in the first iteration process, the initial accuracy range for each iteration can be adjusted based on the number of iterations and the number of overflows to improve the efficiency of the model in subsequent iterations.

[0149] Specifically, in the Nth iteration in the training process of the model, the number of overflows of a plurality of network structures in the model based on the first accuracy range is obtained, where N is a positive integer greater than or equal to 1. When the number of overflows is greater than or equal to the second threshold, it is determined that the initial accuracy range in the next iteration training process is changed from the first accuracy range to the second accuracy range, and the number of overflows is erased to become zero. The number of overflows includes the number of overflows of a plurality of network structures in the forward propagation process and / or the number of overflows of a plurality of network structures in the backpropagation process.

[0150] In addition, the parameters in the embodiment shown in FIG. 3 have multiple cases. In a conceivable implementation, when the parameter is an intermediate feature and the calculation of the intermediate feature overflows the first accuracy range, the second accuracy range can be used to recalculate the intermediate feature, and the second accuracy range can be further used to perform subsequent loss calculation and / or weight gradient calculation. In other words, the calculation accuracy adjustment can be an accuracy range adjustment for the overflow parameter, and further, can be an accuracy range adjustment for another parameter in the subsequent training process. This is not particularly limited herein. In another conceivable implementation, when the parameter is a loss and the calculation of the loss overflows the first accuracy range, the second accuracy range can be used to recalculate the loss (including recalculating the loss by using the second accuracy range, or recalculating from the first layer until the loss is obtained, that is, the recalculation range can include the current calculation, or can include multiple calculations before the current overflow calculation), and the second accuracy range can be further used to perform subsequent weight gradient calculation. In other words, in the training process, when the parameter calculated in the intermediate calculation overflows the accuracy range, the calculation accuracy adjustment can be an accuracy range adjustment for the overflow parameter, and further, can be an accuracy range adjustment for the parameter before the overflow parameter that does not overflow in the training process, and further, can be an accuracy range adjustment for another parameter in the subsequent training process. This is not particularly limited herein.

[0151] FIGS. 5A and 5B show another model training method according to an embodiment of the present application. The execution subject of the method is the same as that of the embodiment shown in FIG. 3, and details are not described here. The method may include steps 501 to 511. Hereinafter, steps 501 to 511 will be described in detail.

[0152] Step 501: Obtain training data.

[0153] For step 501, refer to the description of step 301 in the embodiment shown in FIG. 3. Details are not described in this specification.

[0154] Step 502: Determine whether the number of overflows (num) is greater than or equal to a second threshold value (N). If the number of overflows is greater than or equal to the second threshold value, step Pl 508 is executed. If the number of overflows is less than the second threshold value, step Pl 503 is executed.

[0155] In the present embodiment of the present application, only an example in which the number of overflows includes the number of times that the calculated value of each layer overflows the first accuracy range in the forward calculation process and the backward calculation process of the model is used for the purpose of explanation. In actual applications, it can be understood that the number of overflows may be the number of times that the intermediate features of each layer overflow the first accuracy range in the forward calculation of the model, or may be the number of times that the gradient overflows the first accuracy range in the backward calculation of the model. That is, the number of overflows may be a count in the overall forward and backward processes, or may be a separate count in the forward or backward process. This is not particularly limited herein.

[0156] In the present embodiment of the present application, forward calculation is the calculation of parameters (for example, intermediate features or losses) in the forward propagation process of the model. Backward calculation is the calculation of parameters (for example, gradients) in the backward propagation process of the model.

[0157] The training device determines whether the number of overflows is greater than or equal to a second threshold value (N), where N is an integer greater than or equal to 0. If the number of overflows is greater than or equal to the second threshold value, step Pl 508 is executed. If the number of overflows is less than the second threshold value, step Pl 503 is executed.

[0158] First, it can be understood that the number of overflows num is set to 0.

[0159] Step 503: Perform forward calculation by using the first accuracy range to obtain a loss.

[0160] If the number of overflows in step 502 is greater than the second threshold a smaller case then the execution of this step is triggered.

[0161] For the description of the loss calculation in step 503, refer to the description of step 302 in the embodiment shown in FIG. 3. Details are not described herein.

[0162] Step 504: Determine whether the loss overflows. If the loss overflows, step 509 is executed. If the loss does not overflow, the loss is multiplied by a scaling factor and step 505 is executed.

[0163] After obtaining the loss, the training device determines whether the loss overflows. If the loss overflows, step 509 is executed. If the loss does not overflow, the loss is multiplied by a scaling factor and step 505 is executed.

[0164] Step 505: Perform backward calculation by using the first accuracy range to obtain a weight gradient.

[0165] If the loss in step 504 does not overflow the first accuracy range, the loss is multiplied by a scaling factor and the execution of this step is triggered.

[0166] Step 505 is executed by multiplying by a scaling factor to prevent underflow of the calculated value.

[0167] For the description of the weight gradient calculation in step 505, refer to the description of step 302 in the embodiment shown in FIG. 3. Details are not described herein.

[0168] Step 506: Determine whether the value obtained by dividing the weight gradient by the scaling factor overflows. If the value obtained by dividing the weight gradient by the scaling factor overflows, step 510 is executed. If the value obtained by dividing the weight gradient by the scaling factor does not overflow, step 507 is executed.

[0169] After obtaining the weight gradient, the training device determines whether the value obtained by dividing the weight gradient by the scaling factor overflows the first precision range. If the value obtained by dividing the weight gradient by the scaling factor overflows, step 510 is executed. If the value obtained by dividing the weight gradient by the scaling factor does not overflow, step 507 is executed.

[0170] Step 507: Update the weights.

[0171] If the value obtained by dividing the weight gradient by the scaling factor in step 506 does not overflow the first precision range, and / or after step 508, the execution of this step is triggered.

[0172] Alternatively, if the value obtained by dividing the weight gradient by the scaling factor does not overflow the first precision range, the value is understood to be used to perform iterative updates on the weights.

[0173] Step 508: Perform forward and backward recomputation by using the second precision range.

[0174] The number of overflows in step 502 is greater than the second threshold a larger or equal case In combination, and / or after step 509, the execution of this step is triggered.

[0175] For step 508, refer to the description of step 303 in the embodiment shown in FIG. 3. Details are not described herein.

[0176] Step 509: Increment the overflow count by 1 (i.e., num + 1).

[0177] If the loss in step 504 overflows the first accuracy range and / or the value obtained by dividing the weight gradient by the scaling factor overflows the first accuracy range, the execution of this step is triggered.

[0178] Alternatively, it is understood that the overflow count is recorded, and in subsequent iterations, based on the comparison between the overflow count and the second threshold, it is determined whether to modify the initial accuracy range in each iteration process. For example, if the overflow count is greater than the second threshold, the initial accuracy range is adjusted from the first accuracy range to the second accuracy range. Of course, in addition to the overflow count, the number of iterations can also be considered for adjusting the initial accuracy range. For example, the number of iterations reaches 1000, and the overflow count is greater than 800 (i.e., the second threshold is 800). In this case, by setting the initial accuracy range to the first accuracy range, it indicates that the model training has been affected. To ensure the accuracy of subsequent model training, the initial accuracy range is adjusted to the second accuracy range, thereby reducing problems such as training stagnation caused by the parameters overflowing the first accuracy range.

[0179] Step 510: Update the scaling factor by using the first coefficient.

[0180] If the value obtained by dividing the weight gradient by the scaling factor in step 506 overflows the first accuracy range, the execution of this step is triggered. When the value obtained by dividing the weight gradient by the scaling factor overflows the first accuracy range, the scaling factor is updated by using the first coefficient. The first coefficient is a positive number less than 1.

[0181] Alternatively, when the value obtained by dividing the weight gradient by the scaling coefficient overflows the first precision range, the scaling coefficient is understood to be multiplied by a positive number less than 1 for adjustment.

[0182] Step 511: Determine that the scaling coefficient is greater than or equal to a preset threshold.

[0183] After determining that the value obtained by dividing the weight gradient by the scaling coefficient overflows the first precision range, when the training device updates the scaling coefficient by using the first coefficient, it can determine whether the scaling coefficient is less than a preset threshold. If the scaling coefficient is less than the preset threshold, the scaling coefficient is adjusted to the preset threshold. If the scaling coefficient is greater than or equal to the preset threshold, the scaling coefficient is not modified and step 509 is executed.

[0184] In this step, the lower limit of the value of the scaling coefficient is limited, and as a result, the risk of subsequent parameter underflow can be reduced, and the stability of model training can be improved.

[0185] It can be understood that setting the lower limit of the value of the adjusted scaling coefficient is for the next iteration. After the scaling coefficient is adjusted, step 509 is executed.

[0186] In addition, steps 501 to 511 in this embodiment can be executed multiple times. After being updated one or more times in the model training process, steps 501 to 511 can be executed once. Alternatively, steps 501 to 511 can be executed again when a preset cycle or a preset number of times is satisfied.

[0187] In the present embodiment of the present application, on the one hand, the accuracy range is automatically adjusted in real time by using the overflow information generated in the forward calculation process and / or the reverse calculation process. As a result, the memory occupied by model training can be reduced, and the model training efficiency can be improved. In this way, problems such as training stagnation caused by the parameter calculation overflowing the first accuracy range are reduced. In addition, in the present embodiment of the present application, compared with the method in the prior art in which the type of the network layer of the model needs to be used to determine whether to use high-precision floating-point numbers or low-precision floating-point numbers, the accuracy range to which the parameter is applicable can be adjusted in real time by using the overflow information of the parameter, and the overflow problem caused by the calculation of low-precision floating-point numbers is reduced. On the other hand, the initial accuracy range can be adjusted. When the first accuracy range affects the training of the model, the initial accuracy range is adjusted from the first accuracy range to the second accuracy range to ensure the accuracy of subsequent model training. On the other hand, in the reverse calculation process, the lower limit of the value of the scaling coefficient is set to reduce the risk of subsequent parameter accuracy underflow.

[0188] In the embodiments shown in FIGS. 5A and 5B, it can be understood that there are multiple processing means after the forward calculation overflows. For example, when the forward calculation overflows, the second accuracy range can be used to execute the forward calculation again, and the first accuracy range is used to execute the reverse calculation. For example, when the forward calculation overflows, the second accuracy range can be used to execute the forward calculation again and then execute the subsequent reverse calculation. That is, the accuracy adjustment can be an adjustment only for the overflow calculation or an accuracy adjustment for the entire calculation. This is not particularly limited herein.

[0189] FIGS. 6A and 6B show another model training method according to the embodiment of the present application. The execution subject of the method is the same as the execution subject of the embodiment shown in FIG. 3, and the details are not described here. The method may include steps 601 to 615. Hereinafter, steps 601 to 615 will be described in detail.

[0190] Step 601: Obtain training data.

[0191] Step 602: Determine whether the number of overflows (num) is greater than or equal to the second threshold value (N). If the number of overflows is greater than or equal to the second threshold value, step Pl 612 is executed. If the number of overflows is less than the second threshold value, step Pl 603 is executed.

[0192] Regarding Step 601 and Step 602, refer to the descriptions of Step 501 and Step 502 in the embodiments shown in FIGS. 5A and 5B. Details are not described herein.

[0193] Step 603: Perform forward calculation by using the first accuracy range to obtain intermediate features.

[0194] When the number of overflows in Step 602 is greater than a smaller case this step is triggered.

[0195] Step 604: Determine whether the intermediate features overflow. If the intermediate features overflow, Step 605 is executed. If the intermediate features do not overflow, Step 607 is executed.

[0196] After obtaining the intermediate features, the training device determines whether the calculated value of the intermediate features overflows. If the calculated value of the intermediate features overflows, Step 605 is executed. If the calculated value of the intermediate features does not overflow, Step 607 is executed.

[0197] Step 605: Increment the number of overflows by 1 (i.e., num + 1).

[0198] If the value obtained by dividing the intermediate feature in step 604, the loss in step 608, and / or the weight gradient by the scaling factor overflows the first precision range, the execution of this step is triggered.

[0199] Step 606: Calculate the intermediate feature of the overflow layer using the second precision range, or calculate the intermediate feature layer by layer starting from the first layer.

[0200] This step can be understood as follows: If the intermediate feature overflows the first precision range, the second precision range can be used to perform calculations for the overflow layer or for all layers.

[0201] Step 607: Perform forward calculation by using the first precision range to obtain the loss.

[0202] If the intermediate feature in step 604 does not overflow, the execution of this step is triggered.

[0203] Step 608: Determine whether the loss overflows. If the loss overflows, step 613 is executed. If the loss does not overflow, the loss is multiplied by the scaling factor and step 609 is executed.

[0204] Step 609: Perform backward calculation by using the first precision range to obtain the weight gradient.

[0205] Step 610: Determine whether the value obtained by dividing the weight gradient by the scaling factor overflows. If the value obtained by dividing the weight gradient by the scaling factor overflows, step 614 is executed. If the value obtained by dividing the weight gradient by the scaling factor does not overflow, step 611 is executed.

[0206] Step 611: Update the weights.

[0207] Step 612: Perform forward and reverse recalculations by using the second accuracy range.

[0208] Step 613: Increment the overflow count by 1 (i.e., num+1).

[0209] Step 614: Update the scaling factor by using the first coefficient.

[0210] Step 615: Determine that the scaling factor is greater than or equal to a preset threshold.

[0211] For steps 607 to 615, refer to the descriptions of steps 503 to 511 in the embodiments shown in FIGS. 5A and 5B. Details are not described herein.

[0212] In the present embodiment of the present application, when the calculation of intermediate features overflows the first accuracy range, the second accuracy range can be used to calculate the intermediate features of the overflow layer or to calculate the intermediate features layer by layer starting from the network structure of the first layer. In other words, the adjustment of the accuracy range can be an adjustment of a specific layer or an adjustment of all layers of the entire network structure. On the other hand, the accuracy range is automatically adjusted in real time by using the overflow information generated in the forward calculation process and / or the backward calculation process. As a result, the memory occupied by model training can be reduced, and the model training efficiency can be improved. In this way, problems such as training stagnation caused by the parameter calculation overflowing the first accuracy range are reduced. In addition, in the present embodiment of the present application, compared with the method in the prior art where the type of the network layer of the model needs to be used to determine whether to use high-precision floating-point numbers or low-precision floating-point numbers, the accuracy range to which the parameters are applicable can be adjusted in real time by using the overflow information of the parameters, and the overflow problem caused by the calculation of low-precision floating-point numbers is reduced. On the other hand, the initial accuracy range can be adjusted. When the first accuracy range affects the training of the model, the initial accuracy range is adjusted from the first accuracy range to the second accuracy range to ensure the accuracy of subsequent model training. On the other hand, in the backward calculation process, the lower limit of the value of the scaling factor is set to reduce the risk of subsequent parameter accuracy underflow.

[0213] FIG. 7A, FIG. 7B, and FIG. 7C show another model training method according to an embodiment of the present application. The execution subject of the method is the same as the execution subject of the embodiment shown in FIG. 3, and the details are not described here. The method may include steps 701 to 717. Hereinafter, steps 701 to 717 will be described in detail.

[0214] Step 701: Obtain training data.

[0215] Step 702: Determine whether the number of overflows (num) is greater than or equal to the second threshold (N). If the number of overflows is greater than or equal to the second threshold, step Pl 712 is executed. If the number of overflows is less than the second threshold, step Pl 703 is executed.

[0216] Step 703: Perform forward calculation by using the first accuracy range to obtain intermediate features.

[0217] Step 704: Determine whether the intermediate features overflow. If the intermediate features overflow, step 705 is executed. If the intermediate features do not overflow, step 707 is executed.

[0218] Step 705: Increment the number of overflows by 1 (i.e., num + 1).

[0219] Step 706: Calculate the intermediate features of the overflow layer by using the second accuracy range, or calculate the intermediate features layer by layer starting from the first layer.

[0220] Step 707: Perform forward calculation by using the first accuracy range to obtain the loss.

[0221] Step 708: Determine whether the loss overflows. If the loss overflows, step 716 is executed. If the loss does not overflow, the loss is multiplied by the scaling factor and step 709 is executed.

[0222] Step 709: Perform backward calculation by using the first accuracy range to obtain the weight gradient.

[0223] Step 710: Determine whether the value obtained by dividing the weight gradient by the scaling coefficient overflows. If the value obtained by dividing the weight gradient by the scaling coefficient overflows, step 714 is executed. If the value obtained by dividing the weight gradient by the scaling coefficient does not overflow, step 711 is executed.

[0224] Step 711: Update the weights.

[0225] Step 712: Perform forward and backward recalculations by using the second precision range.

[0226] Step 713: Increment the overflow count by 1 (i.e., num + 1).

[0227] Step 714: Update the scaling coefficient by using the first coefficient.

[0228] Step 715: Determine that the scaling coefficient is greater than or equal to a preset threshold.

[0229] For steps 701 to 715, refer to the descriptions of steps 601 to 615 in the embodiments shown in FIGS. 6A and 6B. Details are not described herein.

[0230] Step 716: Increment the overflow count by 1 (i.e., num + 1).

[0231] Step 716 is the same as step 705. In other words, the overflow count is the cumulative number of times that intermediate features, loss, and weight gradients overflow the first precision range in the model training process.

[0232] Step 717: Recalculate the loss using the second precision range, or starting from the first layer, perform calculations layer by layer until the loss is obtained. The loss is multiplied by the scaling factor, and step 709 is executed.

[0233] This step can be understood as follows: If the loss overflows the first precision range, the second precision range can be used to perform the final recalculation of the loss for precision range adjustment, or perform calculations for all layers of the model for precision range adjustment.

[0234] In the present embodiment of the present application, when the calculation of the loss overflows the first precision range, the second precision range can be used to recalculate the loss or perform calculations layer by layer starting from the network structure of the first layer until the loss is obtained. In other words, the adjustment of the precision range can be an adjustment of a specific layer or an adjustment of all layers of the entire network structure. On the other hand, the precision range is automatically adjusted in real time by using the overflow information generated in the forward calculation process and / or the backward calculation process. As a result, the memory occupied by model training can be reduced, and the model training efficiency can be improved. In this way, problems such as training stagnation caused by the parameter calculation overflowing the first precision range are reduced. In addition, in the present embodiment of the present application, compared with the method in the prior art that requires the type of the network layer of the model to be used to determine whether to use high-precision floating-point numbers or low-precision floating-point numbers, the precision range to which the parameter is applicable can be adjusted in real time by using the overflow information of the parameter, and the overflow problem caused by the calculation of low-precision floating-point numbers is reduced. On the other hand, the initial precision range can be adjusted. When the first precision range affects the training of the model, the initial precision range is adjusted from the first precision range to the second precision range to ensure the accuracy of subsequent model training. On the other hand, in the backward calculation process, the lower limit of the value of the scaling factor is set to reduce the risk of subsequent parameter precision underflow.

[0235] In the above, the model training method in the embodiments of the present application has been described. Below, the training device in the embodiments of the present application will be described. Please refer to FIG. 8. The embodiment of the training device in the embodiments of the present application is: an acquisition unit 801 configured to acquire training data; and a calculation unit 802 configured to use the training data as the input of the model and calculate parameters by using a first accuracy range in the training process of the model to obtain a calculated value.

[0236] If the calculated value overflows the first accuracy range, the calculation unit 802 is further configured to recalculate the parameters by using a second accuracy range, and perform iterative training on the model one or more times by using the recalculated parameters. Here, the second accuracy range includes the first accuracy range, or the second accuracy range partially overlaps with the first accuracy range.

[0237] Optionally, the model includes a plurality of network structures, and specifically, the calculation unit 802 is configured to recalculate the parameters starting from the first-layer network structure of the model by using the second accuracy range.

[0238] Optionally, the calculation unit 802 is specifically configured to recalculate the parameters starting from the current network structure where the calculated value overflows by using the second accuracy range.

[0239] Optionally, the model includes a plurality of network structures, and the parameters include one or more of the following: intermediate features calculated by the plurality of network structures in the forward propagation process or the value of the loss function of the model, where the intermediate features are the output features of any one of the plurality of network structures; and gradients calculated by the plurality of network structures in the backward propagation process, where the gradients include the gradients of the intermediate features and / or the weight gradients of the model.

[0240] Optionally, when the gradient calculated in the backpropagation process includes parameters, the calculated value is a value obtained by dividing the gradient by a scaling factor, and the scaling factor is used to reduce the probability of the gradient overflowing. The calculation unit 802 further updates the scaling factor by using a first coefficient, where the updated scaling factor replaces the non-updated scaling factor and is used to perform the next iterative training of the model. The first coefficient is a positive number less than 1, and the minimum value of the scaling factor is a preset threshold greater than or equal to 1, and is configured as such.

[0241] Optionally, when the parameters include intermediate features calculated in the forward propagation process, the calculation unit 802 specifically calculates the intermediate features of the overflow layer by using a second precision range, or starts from the network structure of the first layer and calculates the intermediate features layer by layer. Here, the overflow layer is in a plurality of network structures, and is a network structure in which the calculated value of the intermediate features overflows the first precision range, and is configured as such.

[0242] Optionally, when the parameters include the value of the loss function of the model, the calculation unit 802 specifically calculates the value of the loss function by using a second precision range, or starts from the network structure of the first layer and is configured to perform calculations layer by layer until the value of the loss function is obtained.

[0243] Optionally, the calculation unit 802 is specifically configured to obtain the number of overflow times of a plurality of network structures in the model based on the first precision range in the Nth iteration in the training process of the model, where N is a positive integer greater than or equal to 1. The calculation unit 802 is specifically configured to: when the number of overflow times is greater than or equal to a second threshold, determine that the initial precision range in the next iterative training process is changed from the first precision range to the second precision range, and configure to eliminate the number of overflow times and set it to zero.

[0244] Optionally, the number of overflows includes the number of overflows of a plurality of network structures in the forward propagation process and / or the number of overflows of a plurality of network structures in the backward propagation process.

[0245] In this embodiment, the operations performed by the units in the training device are the same as those described in the embodiments shown in FIGS. 1 to 7A, 7B, and 7C. Details are not described herein.

[0246] In this embodiment, when the calculated value of a parameter overflows the first accuracy range in the model training process, the calculation unit Tr 802 is recalculates the parameter by using the second accuracy range. That is, the accuracy range is automatically adjusted in real time by using the overflow information of the calculated value of the parameter. As a result, the memory occupied by model training can be reduced, and the model training efficiency can be improved. In this way, problems such as training stagnation caused by the parameter calculation overflowing the first accuracy range are reduced. In addition, in this embodiment of the present application, compared with the conventional art method that needs to be used to determine whether to use high-precision floating-point numbers or low-precision floating-point numbers for the types of network layers of the model, the accuracy range applicable to the parameter can be adjusted in real time by using the overflow information of the parameter, and the overflow problem caused by the calculation of low-precision floating-point numbers is reduced.

[0247] FIG. 9 is a diagram of the structure of another training device according to the present application. The training device may include a processor 901, a memory 902, and a communication port 903. The processor 901, the memory 902, and the communication port 903 are interconnected by using wiring. The memory 902 stores program instructions and data.

[0248] Memory 902 stores program instructions and data corresponding to steps executed by a training device in corresponding implementations shown in FIGS. 1-7A, 7B, and 7C.

[0249] Processor 901 is configured to execute steps executed by a training device in any one of the embodiments shown in FIGS. 1-7A, 7B, and 7C.

[0250] Communication port 903 can be configured to receive and transmit data and is configured to execute steps related to acquisition and reception in any one of the embodiments shown in FIGS. 1-7A, 7B, and 7C.

[0251] In an implementation, the training device may include more or fewer components than those shown in FIG. 9. This is merely an example for illustration and is not limited in the present application.

[0252] Those skilled in the art can clearly understand the detailed operation processes of the above system, device, and unit for the purpose of convenience and simple explanation by referring to the corresponding processes in the embodiments of the above method, and the details are not described in this specification.

[0253] In some embodiments provided in the present application, it should be understood that the disclosed system, device, and method may be implemented in other ways. For example, the described embodiments of the device are merely examples. For example, the division into units is merely a logical function division, and there may be other division methods in actual implementation. For example, a plurality of units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the indicated mutual coupling, direct coupling, or communication connection may be implemented by using some interfaces. The indirect coupling or communication connection between devices or units may be implemented in an electrical form, a mechanical form, or other forms.

[0254] A unit described as a separate part may or may not be physically separate. Also, a part shown as a unit may or may not be a physical unit, may be located in one position, or may be distributed among multiple network units. Some or all of these units can be selected based on the actual requirements for realizing the object of the solution means of the embodiment.

[0255] In addition, the functional units in the embodiments of the present application can be integrated into one processing unit, or each of the units can physically exist alone, or two or more units can be integrated into one unit. The integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0256] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, the integrated unit can be stored in a computer-readable storage medium. Based on such an understanding, although the technical solution means in the present application is essential, the part contributing to the prior art, or all or part of the technical solution means, may be implemented in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for instructing a computer device (which may be a personal computer, a server, or a network device) to execute all or some of the steps of the method described in the embodiments of the present application. The above storage medium includes any medium such as a USB flash drive, a removable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk that can store program codes. 。 [Other possible items] [Item 1] A model training method, comprising: obtaining training data; using the training data as an input to the model, and in the training process of the model, calculating parameters by using a first accuracy range to obtain a calculated value; and when the calculated value overflows the first accuracy range, recalculating the parameters by using a second accuracy range, and performing iterative training on the model one or more times by using the recalculated parameters, where the second accuracy range includes the first accuracy range, or the second accuracy range partially overlaps with the first accuracy range A method comprising the above steps. [Item 2] The model includes a plurality of network structures, and the step of recalculating the parameters by using the second accuracy range is:[[]] recalculating the parameters starting from the network structure of the first layer of the model by using the second accuracy range The method according to Item 1, comprising the above steps. [Item 3] The model includes a plurality of network structures, and the step of recalculating the parameters by using the second accuracy range is:[[]] recalculating the parameters starting from the current network structure where the calculated value overflows by using the second accuracy range The method according to Item 1, comprising the above steps. [Item 4] The model includes a plurality of network structures, and the parameters are as follows:[[]] intermediate features calculated by the plurality of network structures in the forward propagation process or values of the loss function of the model, where the intermediate features are output features of any one of the plurality of network structures; and gradients calculated by the plurality of network structures in the backpropagation process, where the gradients include gradients of the intermediate features and / or weight gradients of the model The method according to any one of Items 1 to 3, including one or more of the above. [Item 5] When the parameter includes the gradient calculated in the backpropagation process, the calculated value is a value obtained by dividing the gradient by a scaling coefficient, and the scaling coefficient is used to reduce the probability that the gradient overflows. The method further includes: Updating the scaling coefficient by using a first coefficient, where the updated scaling coefficient replaces the non-updated scaling coefficient and is used to perform the next iterative training of the model. The first coefficient is a positive number less than 1, and the minimum value of the scaling coefficient is a preset threshold greater than or equal to 1. The method according to item 4, comprising the above. [Item 6] When the parameter includes the intermediate feature calculated in the forward propagation process, the step of recalculating the parameter by using a second accuracy range includes: Calculating the intermediate feature of the overflow layer by using the second accuracy range, or calculating the intermediate feature layer by layer starting from the network structure of the first layer, where the overflow layer is a network structure in the plurality of network structures where the calculated value of the intermediate feature exceeds the first accuracy range. The method according to item 4, including the above. [Item 7] When the parameter includes the value of the loss function of the model, the step of recalculating the parameter by using a second accuracy range includes: Calculating the value of the loss function by using the second accuracy range, or performing calculations layer by layer starting from the network structure of the first layer until the value of the loss function is obtained. The method according to item 4, including the above. [Item 8] The step of training the model by using the recalculated parameter includes: In the Nth iteration of the training process of the model, obtaining the number of overflow times of the plurality of network structures in the model based on the first accuracy range, where N is a positive integer greater than or equal to 1; and When the number of overflow times is greater than or equal to a second threshold, determining that the initial accuracy range in the next iterative training process is changed from the first accuracy range to the second accuracy range, and clearing the number of overflow times to zero. The method according to any one of items 1 to 7, including [Item 9] The method according to item 8, wherein the number of overflows includes the number of overflows of the plurality of network structures in the forward propagation process and / or the number of overflows of the plurality of network structures in the backward propagation process. [Item 10] An acquisition unit configured to acquire training data; and A calculation unit configured to use the training data as an input to the model and calculate parameters by using a first accuracy range in the training process of the model to obtain a calculated value. Here, the calculation unit further: when the calculated value overflows the first accuracy range, re-calculate the parameters by using a second accuracy range, and perform iterative training on the model one or more times by using the re-calculated parameters. Here, the second accuracy range includes the first accuracy range, or the second accuracy range partially overlaps with the first accuracy range. A training device comprising [Item 11] The training device according to item 10, wherein the model includes a plurality of network structures, and specifically, the calculation unit is configured to re-calculate the parameters starting from the network structure of the first layer of the model by using the second accuracy range. [Item 12] The training device according to item 10, wherein the calculation unit is specifically configured to re-calculate the parameters starting from the current network structure where the calculated value overflows by using the second accuracy range. [Item 13] The model includes a plurality of network structures, and the parameters are as follows: Intermediate features calculated by the plurality of network structures in the forward propagation process or values of the loss function of the model, where the intermediate features are output features of any one of the plurality of network structures; and Gradients calculated by the plurality of network structures in the backward propagation process, where the gradients include gradients of the intermediate features and / or weight gradients of the model. The training device according to any one of items 10 to 12, including one or more of the above. [Item 14] When the parameter includes the gradient calculated in the backpropagation process, the calculated value is a value obtained by dividing the gradient by a scaling coefficient, and the scaling coefficient is used to reduce the probability of the gradient overflowing. The calculation unit further: updates the scaling coefficient by using a first coefficient, where the updated scaling coefficient replaces the non-updated scaling coefficient and is used to perform the next iterative training of the model. The first coefficient is a positive number less than 1, and the minimum value of the scaling coefficient is a preset threshold greater than or equal to 1. The training device according to item 13, configured as such. [Item 15] When the parameter includes the intermediate feature calculated in the forward propagation process, specifically, the calculation unit calculates the intermediate feature of the overflow layer by using the second accuracy range, or starts from the network structure of the first layer and calculates the intermediate feature layer by layer. Here, the overflow layer is a network structure in the plurality of network structures where the calculated value of the intermediate feature exceeds the first accuracy range. The training device according to item 13, configured as such. [Item 16] When the parameter includes the value of the loss function of the model, specifically, the calculation unit calculates the value of the loss function by using the second accuracy range, or is configured to perform calculations layer by layer starting from the network structure of the first layer until the value of the loss function is obtained. The training device according to item 13, configured as such. [Item 17] Specifically, the calculation unit obtains the number of overflows of the plurality of network structures in the model based on the first accuracy range in the Nth iteration in the training process of the model, where N is a positive integer greater than or equal to 1, and is configured as such; and Specifically, when the number of overflows is greater than or equal to a second threshold, the calculation unit determines that the initial accuracy range in the next iterative training process is changed from the first accuracy range to the second accuracy range, and is configured to erase the number of overflows and set it to zero. The training device according to any one of items 10 to 16. [Item 18] The training device according to item 17, wherein the number of overflows includes the number of overflows of the plurality of network structures in the forward propagation process and / or the number of overflows of the plurality of network structures in the backward propagation process. [Item 19] A training device comprising a processor, the processor being coupled to a memory, the memory being configured to store a program or instructions, and when the program or the instructions are executed by the processor, the training device is capable of executing the method according to any one of items 1 to 9. [Item 20] A computer storage medium comprising computer instructions, and when the computer instructions are executed on a training terminal device, the training device is capable of executing the method according to any one of items 1 to 9. [Item 21] A computer program product, wherein when the computer program product is executed on a computer, the computer is capable of executing the method according to any one of items 1 to 9.

Claims

1. A model training method, comprising: obtaining training data; using the training data as input to the model, and in the training process of the model, calculating parameters by using a first accuracy range to obtain a calculated value; and when the calculated value overflows the first accuracy range, recalculating the parameters by using a second accuracy range, and performing iterative training on the model one or more times by using the recalculated parameters, where the second accuracy range includes the first accuracy range, or the second accuracy range partially overlaps with the first accuracy range; A method comprising the above steps.

2. The model includes a plurality of network structures, and the step of recalculating the parameters by using a second accuracy range is: recalculating the parameters starting from the network structure of the first layer of the model by using the second accuracy range. The method according to claim 1, comprising the above step.

3. The model includes a plurality of network structures, and the step of recalculating the parameters by using a second accuracy range is: recalculating the parameters starting from the current network structure where the calculated value overflows by using the second accuracy range. The method according to claim 1, comprising the above step.

4. The model includes a plurality of network structures, and the parameters are as follows: intermediate features calculated by the plurality of network structures in the forward propagation process or values of the loss function of the model, where the intermediate features are output features of any one of the plurality of network structures; and gradients calculated by the plurality of network structures in the backpropagation process, where the gradients include gradients of the intermediate features and / or weight gradients of the model. The method according to any one of claims 1 to 3, including one or more of the above.

5. When the parameter includes the gradient calculated in the backpropagation process, the calculated value is a value obtained by dividing the gradient by a scaling factor, and the scaling factor is used to reduce the probability that the gradient overflows. The method further includes: Updating the scaling factor by using the first coefficient, wherein the updated scaling factor replaces the non-updated scaling factor and is used to perform the next iterative training of the model, the first coefficient is a positive number less than 1, and the minimum value of the scaling factor is a preset threshold greater than or equal to 1. The method according to claim 4, comprising the above. **Claim 6** When the parameter includes the intermediate feature calculated in the forward propagation process, the step of recalculating the parameter by using the second accuracy range is: Calculating the intermediate feature of the overflow layer by using the second accuracy range, or calculating the intermediate feature layer by layer starting from the network structure of the first layer, wherein the overflow layer is the network structure in the plurality of network structures where the calculated value of the intermediate feature overflows the first accuracy range. The method according to claim 4, including the above. **Claim 7** When the parameter includes the value of the loss function of the model, the step of recalculating the parameter by using the second accuracy range is: Calculating the value of the loss function by using the second accuracy range, or starting from the network structure of the first layer and performing calculations layer by layer until the value of the loss function is obtained. The method according to claim 4, including the above. **Claim 8** The step of training the model by using the recalculated parameter is: In the Nth iteration of the training process of the model, obtaining the number of overflows of the plurality of network structures in the model based on the first accuracy range, where N is a positive integer greater than or equal to 1; and When the number of overflows is greater than or equal to the second threshold, determining that the initial accuracy range in the next iterative training process is changed from the first accuracy range to the second accuracy range, and clearing the number of overflows to zero. The method according to any one of claims 1 to 7, including the above. **Claim 9** The method according to claim 8, wherein the number of overflows includes the number of overflows of the plurality of network structures in the forward propagation process and / or the number of overflows of the plurality of network structures in the backward propagation process.

10. An acquisition unit configured to acquire training data; and A calculation unit configured to use the training data as an input to a model and calculate parameters by using a first accuracy range to obtain a calculated value in a training process of the model, wherein the calculation unit further: when the calculated value overflows the first accuracy range, recalculate the parameters by using a second accuracy range, and perform iterative training on the model one or more times by using the recalculated parameters, wherein the second accuracy range includes the first accuracy range, or the second accuracy range partially overlaps with the first accuracy range. A training device comprising the above.

11. The model includes a plurality of network structures, and specifically, the calculation unit is configured to recalculate the parameters starting from the network structure of the first layer of the model by using the second accuracy range. The training device according to claim 10.

12. Specifically, the calculation unit is configured to recalculate the parameters starting from the current network structure where the calculated value overflows by using the second accuracy range. The training device according to claim 10.

13. The model includes a plurality of network structures, and the parameters are as follows: Intermediate features calculated by the plurality of network structures in a forward propagation process or values of a loss function of the model, where the intermediate features are output features of any one of the plurality of network structures; and Gradients calculated by the plurality of network structures in a backward propagation process, where the gradients include gradients of the intermediate features and / or weight gradients of the model. The training device according to any one of claims 10 to 12, including one or more of the above.

14. When the parameter includes the gradient calculated in the backpropagation process, the calculated value is a value obtained by dividing the gradient by a scaling coefficient, and the scaling coefficient is used to reduce the probability of the gradient overflowing. The calculation unit further: updates the scaling coefficient by using a first coefficient, where the updated scaling coefficient replaces the non-updated scaling coefficient and is used to perform the next iterative training of the model, the first coefficient is a positive number less than 1, and the minimum value of the scaling coefficient is a preset threshold greater than or equal to 1. The training device according to claim 13, configured as such.

15. When the parameter includes the intermediate features calculated in the forward propagation process, the calculation unit specifically calculates the intermediate features of the overflow layer by using the second accuracy range, or starts from the network structure of the first layer and calculates the intermediate features layer by layer, where the overflow layer is a network structure in the plurality of network structures where the calculated value of the intermediate features exceeds the first accuracy range. The training device according to claim 13, configured as such.

16. When the parameter includes the value of the loss function of the model, the calculation unit specifically calculates the value of the loss function by using the second accuracy range, or is configured to perform calculations layer by layer starting from the network structure of the first layer until the value of the loss function is obtained. The training device according to claim 13, configured as such.

17. The calculation unit specifically obtains the number of overflows of the plurality of network structures in the model based on the first accuracy range in the Nth iteration of the training process of the model, where N is a positive integer greater than or equal to 1, and is configured as such; and The calculation unit specifically determines that when the number of overflows is greater than or equal to a second threshold, the initial accuracy range in the next iterative training process is changed from the first accuracy range to the second accuracy range, and is configured to eliminate the number of overflows and set it to zero. The training device according to any one of claims 10 to 16.

18. The training device according to claim 17, wherein the number of overflows includes the number of overflows of the plurality of network structures in the forward propagation process and / or the number of overflows of the plurality of network structures in the backward propagation process.

19. A training device comprising a processor, wherein the processor is coupled to a memory, the memory is configured to store a program or instructions, and when the program or the instructions are executed by the processor, the training device is capable of executing the method according to any one of claims 1 to 9.

20. A computer storage medium comprising computer instructions, wherein when the computer instructions are executed on a training terminal device, the training device is capable of executing the method according to any one of claims 1 to 9.

21. A computer program product, wherein when the computer program product is executed on a computer, the computer is capable of executing the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Dynamic self-adaptive data truncation method for convolutional neural network calculation

    CN110210611A

  • Arithmetic processing apparatus, arithmetic processing method, and arithmetic processing program

    JP2022094508A

  • Microprocessor with dynamically adjustable bit width for processing data

    US20190227799A1