Data processing method and apparatus
By recording and analyzing the adjustment operations of the loss scale and dynamically adjusting the loss scale, the problem of insufficient dynamic range or insufficient accuracy of data caused by low-precision floating point numbers is solved, and the convergence speed and training efficiency of the machine learning model are improved.
Patent Information
- Application Number
- PCT/CN2024/106531
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-07-19
- Publication Date
- 2025-05-30
AI Technical Summary
During the training of machine learning models, low-precision floating-point numbers are prone to encounter problems such as insufficient dynamic range or insufficient accuracy of data, resulting in poor convergence of the model and impaired convergence rate.
By recording and analyzing the adjustment operations of the loss scale, dynamically adjust the loss scale to achieve adaptive adjustment of the loss scale. The specific method includes determining the scaling window according to the previous adjustment operation and adjusting the loss scale during subsequent training to ensure that it matches the model.
By adaptively adjusting the loss scale, the loss scale value matching the model can be quickly and accurately found, thereby improving the convergence speed and training efficiency of the model.
Smart Images

Figure CN2024106531_30052025_PF_FP_ABST
Abstract
Description
Data processing method and device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 22, 2023, with application number 202311577262.X and application name “Data Processing Method and Device”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of computers and in particular to a data processing method and device. Background Art
[0003] With the development of computer technology, machine learning models are increasingly being applied in various scenarios. To speed up model training, existing technologies can introduce lower-precision data formats and employ a mixed-precision approach to train models.
[0004] For example, you can mix single-precision floating point numbers (i.e., 32-bit floating point numbers (float-point 32, FP32)), half-precision floating point numbers (i.e., 16-bit floating point numbers (float-point 16, FP16)), and FP8 to train the model.
[0005] Among them, in the process of training the model by mixing multiple precision data formats, compared with using high-precision floating-point numbers for model training, when using low-precision floating-point numbers for model training, there is a problem of insufficient data dynamic range or insufficient accuracy. This will affect the convergence of the model and cause the convergence rate to be damaged.
[0006] Summary of the Invention
[0007] The present application provides a data processing method and apparatus for determining loss scale during machine learning model training.
[0008] In a first aspect, a data processing method is provided. The method is applied to an electronic device used to train a machine learning model. The method includes: recording loss scale adjustment operations during machine learning model training; and adjusting the loss scale based on previous loss scale adjustment operations.
[0009] The above method allows the electronic device to adjust the loss scale during subsequent model training based on the changing trend of the loss scale value during previous training. This allows for adaptive adjustment of the loss scale, allowing for quick and accurate identification of a loss scale value that matches the model. This, in turn, improves model convergence speed.
[0010] In one implementation, adjusting the loss scale according to a previous loss scale adjustment operation includes: determining a scale window to be subsequently used according to the previous loss scale adjustment operation, and adjusting the loss scale according to the scale window to be subsequently used.
[0011] In the above implementation, the loss scale is adjusted by adjusting the scale window based on the previous loss scale adjustment operation. For example, when the loss scale value is determined to be stable based on the previous loss scale adjustment operation, the scale window is adjusted upward to reduce the rate of increase in the loss scale during subsequent training. Alternatively, when the loss scale value is determined to have dropped sharply based on the previous loss scale adjustment operation, the scale window is adjusted downward to ensure that the loss scale can be adjusted upward in a timely manner in the subsequent process, avoiding the situation where the loss scale is continuously too low and causes the training process to fail.
[0012] In one implementation, determining a subsequent scale window based on a previous loss scale adjustment includes: when the previous loss scale adjustment satisfies a first preset condition, increasing the scale window to obtain the subsequent scale window. The first preset condition includes: the number of loss scale increases since the last scale window increase reaches a first threshold.
[0013] In the above implementation, the number of loss scale increases can be used to determine whether the loss scale value is stable. When the number of loss scale increases reaches a first threshold (i.e., when the loss scale value is stable), the scale window is adjusted upward, thereby reducing the rate of loss scale increases during subsequent training.
[0014] In one implementation, determining a subsequent scale window based on a previous loss scale adjustment operation includes: when the previous loss scale adjustment operation meets a second preset condition, lowering the scale window to obtain the subsequent scale window; the second preset condition includes: the number of consecutive loss scale reduction operations reaches a second threshold.
[0015] In the above implementation, the number of consecutive loss scale reductions is used to determine whether the loss scale has dropped dramatically. When the number of consecutive loss scale reductions reaches a second threshold (i.e., a sharp drop in the loss scale), the scale window is adjusted downward to ensure that the loss scale can be adjusted upward in a timely manner in the subsequent process, avoiding training failures caused by a persistently low loss scale.
[0016] In one implementation, during machine learning model training, loss scale adjustment operations are recorded, including: during machine learning model training, recording the value of the loss scale after the adjustment operation. Determining a subsequent scale window based on the previous loss scale adjustment operation includes: determining a subsequent scale window based on the value of the currently used loss scale; wherein different loss scales correspond to different preset levels, and different preset levels correspond to different scale windows.
[0017] In the above implementation, multiple preset levels can be set according to the loss scale value, where different loss scales correspond to different preset levels. In addition, the multiple preset levels correspond to different scale windows. In this way, when determining the scale window to be adopted later (i.e., S2021) based on the previous adjustment operation of the loss scale, the corresponding scale window can be determined according to the value of the loss scale. This can achieve adaptive adjustment of the loss scale, so as to quickly and accurately find the loss scale value that matches the model. This can also improve the model convergence speed.
[0018] In one implementation, during the training of a machine learning model, recording loss scale adjustment operations includes: during the training of the machine learning model, recording the maximum absolute value of the gradient (grad) after the loss scale is adjusted. Determining a subsequent scale window based on the previous loss scale adjustment operation includes: when the maximum absolute value of the grad is greater than a third threshold, increasing the scale window used in the previous iteration to obtain the subsequent scale window used.
[0019] In the above implementation, the scale window can be adjusted based on the maximum absolute value of the model's gradient (hereinafter referred to as Amax). When Amax is large enough, it indicates that the loss scale is stabilizing, and the scale window can be adjusted upward. This reduces the rate at which the loss scale is adjusted upward during subsequent training.
[0020] In one implementation, the scale window used later is one of a plurality of preset scale windows; the plurality of preset scale windows respectively correspond to a different number of model iterations.
[0021] In one implementation, adjusting the loss scale based on a previous loss scale adjustment operation includes: determining a ratio of a loss scale used in a previous iteration to a loss scale used thereafter based on the previous loss scale adjustment operation, and adjusting the loss scale based on the ratio of the loss scale used in the previous iteration to the loss scale used thereafter.
[0022] In the above implementation, adaptive adjustment of the loss scale can be achieved by adjusting the amplitude of each adjustment of the loss scale.
[0023] In one implementation, the machine learning model is a model for processing at least one of text data, image data, or audio data.
[0024] In a second aspect, a data processing device is provided for use in an electronic device for training a machine learning model. The data processing device includes: a recording unit for recording loss scale adjustment operations during machine learning model training; and an adjustment unit for adjusting the loss scale based on previous loss scale adjustment operations.
[0025] In one implementation, an adjustment unit, configured to adjust the loss scale based on a previous loss scale adjustment operation, includes: determining a scale window to be subsequently used based on the previous loss scale adjustment operation; and further adjusting the loss scale based on the subsequently used scale window.
[0026] In one implementation, an adjustment unit, configured to determine a subsequent scale window based on a previous loss scale adjustment operation, includes: an adjustment unit configured to increase the scale window to obtain the subsequent scale window when the previous loss scale adjustment operation satisfies a first preset condition. The first preset condition includes: the number of loss scale increases since the last scale window increase reaches a first threshold.
[0027] In one implementation, an adjustment unit, configured to determine a subsequent scale window based on a previous loss scale adjustment operation, includes an adjustment unit configured to adjust the scale window downward to obtain the subsequent scale window when the previous loss scale adjustment operation satisfies a second preset condition. The second preset condition includes: the number of consecutive loss scale downward adjustment operations reaching a second threshold.
[0028] In one implementation, a recording unit is used to record the adjustment operation of the loss scale during the machine learning model training process, including: a recording unit is used to record the value of the loss scale after the adjustment operation during the machine learning model training process; an adjustment unit is used to determine the scale window used subsequently based on the previous adjustment operation of the loss scale, including: an adjustment unit is used to determine the scale window used subsequently based on the value of the currently used loss scale; wherein, loss scales of different sizes correspond to different preset levels, and different preset levels correspond to different scale windows.
[0029] In one implementation, a recording unit is used to record the adjustment operation of the loss scale during the machine learning model training process, including: a recording unit is used to record the maximum absolute value of the gradient grad after the loss scale is adjusted during the machine learning model training process; an adjustment unit is used to determine the scale window used subsequently based on the previous adjustment operation of the loss scale, including: an adjustment unit is used to increase the scale window used in the previous iteration when the maximum absolute value of the grad is greater than a third threshold, to obtain the scale window used subsequently.
[0030] In one implementation, the scale window used later is one of a plurality of preset scale windows; the plurality of preset scale windows respectively correspond to a different number of model iterations.
[0031] In one implementation, an adjustment unit, configured to adjust the loss scale based on a previous loss scale adjustment operation, includes: an adjustment unit, configured to determine, based on the previous loss scale adjustment operation, a ratio of a loss scale used in a previous iteration to a loss scale used thereafter; and an adjustment unit, further configured to adjust the loss scale based on the ratio of the loss scale used in the previous iteration to the loss scale used thereafter.
[0032] In one implementation, the machine learning model is a model for processing at least one of text data, image data, or audio data.
[0033] In a third aspect, a data processing device is provided, comprising a memory and a processor, wherein the memory is used to store computer instructions, and the processor is used to call and execute computer instructions from the memory to implement a method as in the first aspect or any implementation method in the first aspect.
[0034] In a fourth aspect, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the instructions are executed on a processor, the method of the first aspect or any implementation method of the first aspect is implemented.
[0035] In a fifth aspect, a computer program product is provided, which includes instructions. When the instructions are executed on a processor, the method of the first aspect or any implementation of the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] FIG1 is a flow chart of a model training process according to an embodiment of the present invention;
[0037] FIG2 is a flow chart of a data processing method according to an embodiment of the present application;
[0038] FIG3 is a second flow chart of a data processing method provided in an embodiment of the present application;
[0039] FIG4 is a third flow chart of a data processing method provided in an embodiment of the present application;
[0040] FIG5 is a second flow chart of a model training process provided in an embodiment of the present application;
[0041] FIG6 is a third flow chart of a model training process provided in an embodiment of the present application;
[0042] FIG7 is a fourth flow chart of a data processing method provided in an embodiment of the present application;
[0043] FIG8 is a fifth flow chart of a data processing method provided in an embodiment of the present application;
[0044] FIG9 is a fourth flow chart of a model training process provided in an embodiment of the present application;
[0045] FIG10 is a sixth flow chart of a data processing method provided in an embodiment of the present application;
[0046] FIG11 is a structural diagram of a data processing device according to an embodiment of the present application;
[0047] FIG12 is a second structural diagram of a data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0048] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0049] To facilitate understanding of the technical solutions provided in the embodiments of the present application, the relevant technologies involved in the embodiments of the present application are first introduced:
[0050] At present, in the process of training models using a mixture of multiple precision data formats, in order to avoid the problems of insufficient data dynamic range or insufficient accuracy when using low-precision floating-point numbers for model training, the corresponding loss scale can be used to scale the data during the model training process, thereby preserving the accuracy of the gradient and reducing the occurrence of underflow.
[0051] For example, when the data value is very small, when the data is converted from a high-precision floating point number (such as FP32) to a low-precision floating point number (such as FP16), underflow may occur, causing the data to become 0. In this case, the data can be amplified (for example, by 2 times, 4 times, or 8 times, etc.) by multiplying the data by the corresponding loss scale. In this way, when the amplified data is converted from a high-precision floating point number (such as FP32) to a low-precision floating point number (such as FP16), the possibility of underflow is reduced.
[0052] The following examples introduce the techniques for scaling data using loss scale:
[0053] In the first related technology, a fixed loss scale can be used to scale the data during model training.
[0054] However, the first related technology mentioned above has at least the following problems: (1) The value of the loss scale is a hyperparameter that requires multiple training and optimization of the model. In most cases, an inappropriate loss scale will not be exposed until the later stages of training. Therefore, it takes a lot of resources to determine the appropriate value of the loss scale. (2) After modifying other hyperparameters in the model, the loss scale needs to be re-optimized. (3) Using only one fixed loss scale may not be able to handle the differences between different types of gradient tensors. When the numerical differences between tensors are large, using a fixed loss scale cannot take into account all tensors, which will lead to training failure.
[0055] In a second related technique, the loss scale can be dynamically adjusted during model training. Specifically, the loss scale can be adjusted downward based on whether the backpropagation gradient overflows (for example, if a non-number (NaN) or infinite (inf) prompt appears in the backpropagation gradient) or upward based on a preset scale window.
[0056] For example, when training a machine learning model according to the above-mentioned related technologies, as shown in FIG1 , the model training process includes:
[0057] S101. Initialize the FP32 model (for ease of description in the embodiments of this application, the model using FP32 is referred to as the "FP32 model," and the model using FP16 is referred to as the "FP16 model"). This model is built based on our needs and problems, such as a deep neural network.
[0058] S102: Start iterative training of the model.
[0059] S103. At the beginning of each iteration, copy the weights in the FP32 model to the FP16 model.
[0060] S104: Perform a forward pass on the FP16 model. During this process, the input data is passed through the model to obtain the predicted results. Because FP16 is used, this process is faster than using FP32.
[0061] S105. Calculate the loss function (loss).
[0062] S106. Use loss scale to scale the loss. Simply put, this is to enlarge or reduce the loss by a certain ratio. This helps to preserve the accuracy of the gradient and reduce the occurrence of underflow, thereby solving the problem of insufficient data dynamic range or insufficient accuracy when using low-precision floating-point numbers for model training.
[0063] S107: Perform back propagation on the FP16 model. In this process, the model loss can be calculated based on the prediction results and the true labels, and then the model parameters can be adjusted based on the loss.
[0064] S108. Determine whether NaN or inf prompts appear in the gradient. If NaN or inf prompts appear, it indicates that the data has overflowed, and execute S109 (downgrade the loss scale and retrain the model). If no NaN or inf prompts appear, it indicates that the training process is normal, and execute S110 (unscale the gradient using the loss scale and update the weights of the FP32 model).
[0065] S109. Lower the loss scale and retrain the model.
[0066] S110. Use the loss scale to unscale the gradient and update the weights of the FP32 model to train the FP32 model.
[0067] S111. Determine whether the scale window is hit. If the scale window is hit, execute S112 (increase the loss scale and start a new iteration); if the scale window is not hit, execute S102 (start a new iteration). The value of the scale window can be set according to actual needs. For example, if the scale window is set to 20 (i.e., 20 iterations), on the one hand, if it is determined according to S108 that no NaN or inf prompts appear during the training process of 20 consecutive iterations, the scale window is hit, and then S112 is executed (increase the loss scale and start a new iteration); on the other hand, if the number of iterations without NaN or inf prompts does not reach 20, that is, the scale window is not hit, then S102 (start a new iteration).
[0068] S112: Increase the loss scale. After increasing the loss scale, start a new iterative training (i.e., S102), or end the model training process.
[0069] It can be seen that although the second related technology mentioned above can dynamically adjust the loss scale, this method still cannot quickly and accurately find a loss scale value that matches the model. For example, when the loss window value is too large, the loss scale may be adjusted too slowly, resulting in a mismatch between the used loss scale and the model. When the loss window value is too small, the loss scale may be adjusted too quickly, resulting in the loss of relevant training data.
[0070] In addition, in addition to the two related technologies mentioned above that use loss scale to scale the global data in the machine learning model, the third related technology can also avoid the problems of insufficient data dynamic range or insufficient accuracy when using low-precision floating-point numbers for model training by scaling the tensors used in the machine learning model separately.
[0071] Specifically, in the third related technology, the maximum value in the tensor (called Amax corresponding to the tensor) can be calculated for each tensor used in the machine learning model, and then each tensor can be scaled according to the ratio of the maximum value of the currently used low-precision floating-point number (such as FP16) to the maximum value in the tensor (hereinafter referred to as scaling).
[0072] The third related technology mentioned above has the following problems: (1) This method requires calculating the corresponding Amax for each tensor separately. It increases vector operations and computing tasks of the compute unified device architecture core (CUDA core), and also increases memory movement activities. (2) The transformer engine application programming interface (Transformer Engine API) in this method opens the Amax pool size and scaling value algorithm for users to modify and design according to network characteristics, which increases the complexity of use. (3) When using the third related technology, it still needs to be used in conjunction with the global loss scale.
[0073] To address the issues with the aforementioned related technologies, embodiments of the present application consider that, during the process of training a machine learning model on an electronic device, when using a loss scale to scale the data in the machine learning model, the subsequent loss scale can be adjusted based on the changing trend of the loss scale during the previous training process. This allows for adaptive adjustment of the loss scale, allowing for quick and accurate identification of a loss scale value that matches the model.
[0074] For example, when the loss scale value is determined to be stable based on the changing trend of the loss scale value in the previous training process, it means that the loss scale at this time is highly matched with the model. At this time, the scale window can be increased (for example, in the previous training process, the scale window is 20. Then the scale window can be increased to 40 in the subsequent training process), or the increase amplitude of the loss scale can be reduced (for example, each time the loss scale is increased, the original loss scale value is multiplied by 4. Then each time the loss scale is increased in the subsequent training process, the original loss scale value can be multiplied by 2) to avoid a large increase in the loss scale in the subsequent training process.
[0075] For another example, when it is determined that the loss scale value has dropped sharply based on the changing trend of the loss scale value in the previous training process, the scale window can be lowered (for example, in the previous training process, the scale window is 40. Then the scale window can be increased to 20 in the subsequent training process), or the increase amplitude of the loss scale can be increased (for example, each time the loss scale was increased, the original loss scale value was multiplied by 2. Then, each time the loss scale was increased in the subsequent training process, the original loss scale value can be multiplied by 4) to ensure that the loss scale can be adjusted in time in the subsequent process to avoid the failure of the training process due to the loss scale being too low.
[0076] Based on the above considerations, an embodiment of the present application provides a data processing method. The data processing method can be applied to various electronic devices that can be used to train machine learning models, so that the electronic device can perform model training according to the data processing method provided in the embodiment of the present application during operation. Among them, the above-mentioned electronic device can be a desktop computer, a tablet computer, a desktop type, a laptop, a handheld computer, a notebook computer, a super mobile personal computer, a netbook, as well as a cellular phone, a personal digital assistant, an augmented reality\virtual reality device and various artificial intelligence hardware accelerators and other devices. In actual application, the data processing method can be executed by the above-mentioned electronic device, or it can be executed by some hardware / software devices in the electronic device, and this embodiment of the present application may not be limited to this.
[0077] Specifically, in the process of training a machine learning model by an electronic device, as shown in FIG2 , the data processing method may include:
[0078] S201. The electronic device records the adjustment operation of the loss scale during the process of training the machine learning model.
[0079] The machine learning model is a model for processing at least one of text data, image data or audio data.
[0080] For example, during the training of a machine learning model, each time the loss scale is adjusted up or down, the electronic device records the data accordingly.
[0081] The specific content recorded by the electronic device may include one or more of the following: the specific content of the loss scale adjustment operation (including an increase or decrease operation), the value of the loss scale after the adjustment, or the maximum absolute value of the gradient grad when the gradient grad is detected through backpropagation after the loss scale is adjusted. In addition, after the loss scale adjustment operation, the electronic device may also record other content that can reflect the trend of loss scale changes. The specific content recorded by the electronic device can be determined based on the specific implementation needs of the subsequent S202.
[0082] S202: The electronic device adjusts the loss scale according to the previous recorded adjustment operation on the loss scale.
[0083] It is understood that in the method provided in the embodiments of the present application, the previous adjustment of the loss scale can reflect the changing trend of the loss scale value during the previous training process. Therefore, the above method can adjust the loss scale during the subsequent model training process based on the changing trend of the loss scale value during the previous training process, thereby achieving adaptive adjustment of the loss scale to quickly and accurately find the loss scale value that matches the model.
[0084] In one implementation, it is considered that: in S202, the electronic device can adjust the loss scale by adjusting the scale window according to the previous loss scale adjustment operation. For example, when it is determined that the value of the loss scale tends to be stable according to the previous loss scale adjustment operation, the scale window is increased to reduce the speed of increasing the loss scale in the subsequent training process. Alternatively, when it is determined that the value of the loss scale drops sharply according to the previous loss scale adjustment operation, the scale window is lowered to ensure that the loss scale can be adjusted up in time in the subsequent process, thereby avoiding the situation where the training process fails due to the loss scale being too low.
[0085] Based on the above considerations, as shown in FIG3 , S202 may specifically include the following contents of S2021-S2022:
[0086] S2021. The electronic device determines a scale window to be used subsequently based on a previous adjustment operation on the loss scale.
[0087] In one possible design, S2021 may include the following contents of S20211:
[0088] S20211: When the previous adjustment operation on the loss scale meets a first preset condition, the electronic device increases the scale window to obtain a scale window used subsequently.
[0089] The first preset condition includes: after the last increase of the scale window, the number of times the loss scale is increased reaches a first threshold m.
[0090] In an optional manner, multiple gears may be pre-set for the scale window, wherein different gears correspond to different sizes of the scale window.
[0091] For example, electronic devices have nine pre-set scale window settings: [1, 20, 30, 40, 50, 100, 200, 500, 1000]. For example, if the scale window is 20, then if no NaN or inf values appear within 20 model iterations, the loss scale is adjusted upward. If NaN or inf values appear within fewer than 20 model iterations, the loss scale is adjusted downward and the scale window is reset. It is understood that for other scale window settings, the loss scale is adjusted accordingly.
[0092] Each time the scale window is adjusted upward or downward, the scale window is adjusted upward or downward by one gear from the original gear.
[0093] In addition, in an optional manner, an increase count (hereinafter referred to as growth_num) may be maintained in the electronic device.
[0094] The initial value of growth_num is 0. Each time the loss scale is increased, growth_num+1 is added. When growth_num=m (ie, the number of times the loss scale is increased reaches the first threshold m), the electronic device increases the scale window and resets growth_num to 0.
[0095] For example, when m is 3 (ie, the first threshold is 3), the electronic device is triggered to increase the scale window every time the loss scale is increased three times.
[0096] Through the above S20211 process, after the loss scale is increased multiple times, the scale window can be increased to reduce the increase speed of the loss scale in the subsequent training process.
[0097] In another possible design, S2021 may further include the following content of S20212:
[0098] S20212: When the previous adjustment operation on the loss scale meets a second preset condition, the electronic device lowers the scale window to obtain a scale window used subsequently.
[0099] The second preset condition includes: the number of times the loss scale is continuously reduced reaches a second threshold value n.
[0100] In an optional manner, a down count (hereinafter referred to as down_num) may be maintained in the electronic device.
[0101] The initial value of down_num is 0. Each time the loss scale is lowered, down_num is increased by 1; when the loss scale is raised, down_num is reset to 0. In addition, when down_num=n (that is, the number of consecutive reductions in the loss scale reaches the second threshold n), the electronic device lowers the scale window.
[0102] For example, when n is 3 (ie, the second threshold is 3), when the loss scale is lowered three times in succession, the electronic device is triggered to lower the scale window.
[0103] Through the above S20212 process, after the loss scale is lowered multiple times continuously, the scale window can be lowered to ensure that the loss scale can be adjusted up in time in the subsequent process, thereby avoiding the situation where the training process fails due to the loss scale being continuously too low.
[0104] In one optional embodiment, S20212 specifically includes:
[0105] When the previous adjustment operation of the loss scale meets the second preset condition, the electronic device adjusts the scale window down to the lowest gear to obtain the scale window used subsequently.
[0106] That is, in the above approach, when it is detected that the number of consecutive loss scale reduction operations reaches the second threshold n, the scale window is directly lowered to the lowest gear. In this way, when a large amount of abnormal data appears in the data, causing the gradient to continuously display NaN or inf prompts and causing the loss scale to be repeatedly reduced, the scale window can be directly lowered to the lowest gear, thereby ensuring that the loss scale can be adjusted up in time in the subsequent process, avoiding the situation where the loss scale is continuously too low and causing the training process to fail.
[0107] After adjusting the scale window according to the above S2021, the method further includes:
[0108] S2022. The electronic device adjusts the loss scale according to the scale window determined in S2021.
[0109] For example, taking the scale window determined by S2021 as 20 as an example, if no NaN or inf prompt appears during the 20 model iterations corresponding to a scale window of the electronic device, the loss scale is adjusted upward (for example, the loss scale is multiplied by 2) and the contents of S201-S202 are re-executed; in addition, if a NaN or inf prompt appears in a scale window, the loss scale is adjusted downward (for example, the loss scale is divided by 2) and the contents of S201-S202 are re-executed.
[0110] The following is an example of the practical application process of the data processing method provided in the embodiment of the present application, combined with the workflow of the electronic device in the process of training the machine learning model. As shown in Figure 5, the workflow of the electronic device includes:
[0111] S301. The electronic device initializes the machine learning model.
[0112] In S301, the electronic device may initialize various parameters corresponding to the machine learning model. For example, in S301, the electronic device may set the corresponding scale window. For example, nine scale window gears are pre-set in the electronic device: [1, 20, 30, 40, 50, 100, 200, 500, 1000].
[0113] S302. During the first stage of model training, the electronic device uses relevant technologies to iterate the machine learning model a preset number of times (assuming it is p times).
[0114] For example, the electronic device may first iterate the machine learning model for a number p (for example, p may be 20 or 30, etc.) according to the process corresponding to FIG1 . The specific implementation process of S302 may refer to the corresponding description of FIG1 .
[0115] When training the machine learning model according to the process corresponding to Figure 1, you can first let the loss scale drop freely so that the loss scale drops to a value that does not cause overflow.
[0116] S303 : During the second stage of model training, the electronic device adjusts the loss window and loss scale according to the contents of FIG. 4 .
[0117] For example, as shown in FIG6 , S303 specifically includes the contents of S401 to S414:
[0118] S401. The electronic device initializes the FP32 model.
[0119] It is understandable that the FP32 model is used as an example here for illustration, and other high-precision floating-point format models can also be used in actual applications.
[0120] S402: The electronic device performs a new iterative training on the model.
[0121] S403. At the beginning of each iteration, the electronic device copies the weights in the FP32 model to the FP16 model.
[0122] S404. The electronic device performs forward propagation on the FP16 model.
[0123] S405. The electronic device calculates loss.
[0124] S406: The electronic device scales the loss using a loss scale.
[0125] S407. The electronic device performs backpropagation on the FP16 model.
[0126] S408: The electronic device determines whether NaN or inf appears in the gradient. If NaN or inf appears, execute S409 (downgrade the loss scale); if not, execute S411 (unscale the gradient using the loss scale and update the FP32 model weights).
[0127] For the specific contents of the above S401-S409 and S411, reference may be made to the corresponding contents of S101-S109 and S110 in FIG1 above, and the repeated parts will not be repeated here.
[0128] In addition, after S409, the method further includes:
[0129] S410 : Record the adjustment operation of the loss scale according to the content of S201 above and adjust the scale window according to the content of S2021 .
[0130] For example, in Figure 6, after the electronic device lowers the loss scale (i.e., S409), it first increases the lowering count down_num by 1 (the process of adjusting the count down_num can be understood as recording the adjustment operation on the loss scale). It then determines whether down_num reaches n. If down_num = n, the scale window is lowered and model training is re-performed; if down_num ≠ n, the scale window is not adjusted and model training is directly re-performed.
[0131] After S411, the method further includes:
[0132] S412: The electronic device determines whether the scale window is hit. If the scale window is hit, S413 is executed (loss scale is increased); if the scale window is not hit, S402 is executed (a new iteration is started).
[0133] For the specific contents of S412 and S413, please refer to the corresponding contents of S111 and S112 in Figure 1 above, and the repeated parts will not be repeated here.
[0134] After S413, the method further includes:
[0135] S414: Record the adjustment operation of the loss scale according to the content of S201 above and adjust the scale window according to the content of S2021.
[0136] For example, in Figure 6, after the electronic device increases the loss scale (i.e., S413), it increases the count growth_num+1 and resets the count down_num to 0 (the processing of increasing the count growth_num and decreasing the count down_num can be understood as recording the adjustment operation on the loss scale). It then determines whether growth_num reaches m. When growth_num = m, the scale window is increased and the model training is re-performed; when growth_num ≠ m, the scale window is not adjusted and the model training is directly re-performed.
[0137] In another implementation, it is considered that: multiple preset levels can be set according to the value of the loss scale, where different loss scales correspond to different preset levels. In addition, multiple preset levels correspond to different scale windows. In this way, in the process of determining the scale window to be adopted later (i.e., S2021) based on the previous adjustment operation of the loss scale, the corresponding scale window can be determined according to the value of the loss scale. Therefore, in this method, as shown in Figure 7, the above S201 may specifically include:
[0138] S201a. During the machine learning model training process, the electronic device records the value of the loss scale after the adjustment operation.
[0139] For example, a scaling counter scale_num is set in the electronic device. Each time the loss scale is adjusted, the value of the loss scale after the adjustment operation is recorded in scale_num.
[0140] In addition, the above S2021 specifically includes:
[0141] S2021a: The electronic device determines a scale window to be used later according to the value of the currently used loss scale.
[0142] Different loss scales correspond to different preset levels, and different preset levels correspond to different scale windows. Specifically, a larger loss scale corresponds to a larger scale window; conversely, a smaller loss scale corresponds to a smaller scale window.
[0143] After determining the scale window to be adopted subsequently according to the content of S2021a, the electronic device may adjust the loss scale according to the scale window to be adopted subsequently (ie, S2022).
[0144] For example, when the process of determining the scale window shown in FIG. 7 is applied to the process shown in FIG. 6 , S410 and S414 may specifically include: the electronic device recording the value of the loss scale after the adjustment operation (i.e., S201a), and determining the scale window to be subsequently adopted based on the value of the currently adopted loss scale (i.e., S2021a). For other processes other than S410 and S414, refer to the corresponding description of FIG. 6 above, and repeated content is not repeated here.
[0145] In another implementation, it is considered that the scale window can be adjusted according to the maximum absolute value of the gradient of the model (hereinafter referred to as Amax). When Amax is large enough, it means that the loss scale tends to be stable, and the scale window is then adjusted upward. This reduces the speed of increasing the loss scale during subsequent training. Therefore, in this method, as shown in Figure 8, the above S201 may specifically include:
[0146] S201b. During the training process of the machine learning model, the electronic device records the maximum value Amax of the absolute value of grad after adjusting the loss scale.
[0147] For example, during model training, after adjusting the loss scale, the electronic device can detect the absolute value of grad during backpropagation in the next model training cycle and then select the maximum value, Amax, from the absolute values of grad.
[0148] In addition, the above S2021 specifically includes:
[0149] S2021b: When the maximum absolute value of grad is greater than a third threshold q, the electronic device increases the scale window used in the previous iteration to obtain a scale window used subsequently.
[0150] Among them, the value of the third threshold can be determined according to actual needs, and there is no limitation on this in the embodiment of the present application.
[0151] After determining the scale window to be adopted subsequently according to the content of S2021b, the electronic device may adjust the loss scale according to the scale window to be adopted subsequently (ie, S2022).
[0152] Exemplarily, when the scale window is determined using the content shown in FIG8 , as shown in FIG9 , the model training process of the electronic device includes the content of S401 to S414:
[0153] S501: The electronic device initializes the FP32 model.
[0154] S502: The electronic device performs a new iterative training on the model.
[0155] S503. At the beginning of each iteration, the electronic device copies the weights in the FP32 model to the FP16 model.
[0156] S504. The electronic device performs forward propagation on the FP16 model.
[0157] S505. The electronic device calculates loss.
[0158] S506: The electronic device scales the loss using a loss scale.
[0159] S507. The electronic device performs backpropagation on the FP16 model.
[0160] For the specific contents of S501-S407 above, reference may be made to the corresponding contents of S101-S107 in FIG1 above, and the repeated parts will not be repeated here.
[0161] S508: The electronic device records the maximum value Amax of the absolute value of grad according to the content of S201b above and adjusts the scale window according to the content of S2021b.
[0162] For example, in Figure 8 , after each backpropagation, the electronic device records the maximum absolute value Amax of grad (i.e., S201b ). It then determines whether Amax>q. If so, the scale window is adjusted upward; if not, the scale window is not adjusted (i.e., S2021b ).
[0163] S509: The electronic device determines whether NaN or inf appears in the gradient. If NaN or inf appears, execute S510 (downgrade the loss scale); if not, execute S512 (unscale the gradient using the loss scale and update the FP32 model weights).
[0164] In addition, after S510, the method further includes:
[0165] S511 . Record the adjustment operation of the loss scale according to the content of S201 above and adjust the scale window according to the content of S2021 .
[0166] For example, in Figure 8, after the electronic device lowers the loss scale (i.e., S510), it first increases the lowering count down_num by 1 (the process of adjusting the count down_num can be understood as recording the adjustment operation on the loss scale). It then determines whether down_num reaches n. If down_num = n, the scale window is lowered and model training is re-performed; if down_num ≠ n, the scale window is not adjusted and model training is directly re-performed.
[0167] After S512, the method further includes:
[0168] S513: The electronic device determines whether the scale window is hit. If the scale window is hit, S514 is executed (loss scale is increased); if the scale window is not hit, S502 is executed (a new iteration is started).
[0169] For the specific contents of S509-S514 above, reference can be made to the corresponding contents of S408-S413 above, and the repeated parts will not be repeated here.
[0170] 3-9 above mainly introduce the implementation method of adaptively adjusting the loss scale by adjusting the scale window. In another implementation method, in the embodiment of the present application, adaptive adjustment of the loss scale can also be achieved by adjusting the amplitude of each adjustment of the loss scale.
[0171] For example, when the loss scale tends to be stable, you can increase the loss scale slightly later (for example, each time you increase the loss scale, you multiply the original loss scale by 4. Then, each time you increase the loss scale in the subsequent training process, you can multiply the original loss scale by 2) to avoid a large increase in the loss scale in the subsequent training process.
[0172] For another example, when the loss scale value drops sharply, the subsequent increase in loss scale can be achieved by significantly increasing the loss scale (for example, each time the loss scale is increased, the original loss scale value is multiplied by 2. Then, each time the loss scale is increased during subsequent training, the original loss scale value can be multiplied by 4). This ensures that the loss scale can be increased in a timely manner in the subsequent process, avoiding the failure of the training process due to a continuously low loss scale.
[0173] Based on the above considerations, as shown in FIG10 , S202 may specifically include the following contents of S2023-S2024:
[0174] S2023: The electronic device determines a ratio (hereinafter referred to as a scale ratio for ease of description) between the loss scale used in the previous iteration and the loss scale used thereafter, based on the previous adjustment operation on the loss scale.
[0175] For example, after the last increase in the scale ratio, when the number of increases in the loss scale reaches a first threshold m, the scale ratio is increased, thereby achieving the purpose of slightly increasing the loss scale in the subsequent model training process.
[0176] For another example, when the number of consecutive reduction operations on the loss scale reaches a second threshold n, the scale ratio is reduced, thereby achieving the purpose of significantly increasing the loss scale in the subsequent model training process.
[0177] S2024: The electronic device adjusts the loss scale according to the ratio of the loss scale used in the previous iteration to the loss scale used subsequently.
[0178] In addition, in the data processing method provided in the embodiments of the present application, when the electronic device performs forward propagation on the model, overflow conditions that occur during the forward propagation (for example, recording Nan / inf prompts that appear during the forward propagation) are recorded. If a large amount of overflow occurs during the forward propagation, the scale window is lowered. This ensures that the loss scale can be adjusted up in a timely manner in the subsequent process, avoiding the situation where the training process fails due to a continuously low loss scale.
[0179] The data processing method provided in accordance with this embodiment is described in detail above with reference to FIG. 2 to FIG. 10 . Various devices corresponding to the data processing method provided in this embodiment will be described below.
[0180] As shown in Figure 11, a schematic diagram of the structure of a data processing device provided in this embodiment is provided. The data processing device 60 can be a software / hardware device for model training. Specifically, the data processing device 60 can include a desktop computer, a tablet computer, a desktop, a laptop, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, as well as a cellular phone, a personal digital assistant (PDA), an augmented reality (AR)\virtual reality (VR) device, and various artificial intelligence (AI) hardware accelerators and other electronic devices that can perform model training, or the data processing device 60 can also be a software device or hardware device running in the above electronic devices.
[0181] The data processing device 60 can be used to execute all or part of the steps executed by the electronic device in Figures 2 to 10 above. Specifically, the data processing device 60 includes:
[0182] The recording unit 601 is used to record the adjustment operation of the loss scale during the machine learning model training process.
[0183] The adjusting unit 602 is configured to adjust the loss scale according to the previous adjustment operation on the loss scale.
[0184] In one implementation, the adjusting unit 602 is configured to adjust the loss scale according to a previous loss scale adjustment operation, including:
[0185] The adjusting unit 602 is configured to determine a scale window to be used subsequently based on a previous adjustment operation on the loss scale.
[0186] The adjustment unit 602 is further configured to adjust the loss scale according to the scale window used later.
[0187] In one implementation, the adjustment unit 602 is configured to determine a subsequent scale window based on a previous adjustment operation on the loss scale, including:
[0188] The adjustment unit 602 is configured to increase the scale window to obtain the scale window to be subsequently adopted when the previous loss scale adjustment operation meets a first preset condition. The first preset condition includes: the number of loss scale increases since the last scale window increase reaches a first threshold.
[0189] In one implementation, the adjustment unit 602 is configured to determine a subsequent scale window based on a previous adjustment operation on the loss scale, including:
[0190] The adjustment unit 602 is configured to adjust the scale window downward to obtain the scale window used subsequently when the previous loss scale adjustment operation meets a second preset condition, wherein the second preset condition includes: the number of consecutive loss scale adjustment operations reaches a second threshold.
[0191] In one implementation, the recording unit 601 is configured to record, during the machine learning model training process, an adjustment operation on the loss scale, including:
[0192] A recording unit 601 is used to record the value of the loss scale after the adjustment operation during the machine learning model training process;
[0193] The adjustment unit 602 is configured to determine a subsequent scale window based on the previous adjustment operation on the loss scale, including:
[0194] The adjusting unit 602 is configured to determine a scale window to be used subsequently according to the value of the currently used loss scale. Loss scales of different sizes correspond to different preset levels, and different preset levels correspond to different scale windows.
[0195] In one implementation, the recording unit 601 is configured to record the adjustment operation of the loss scale during the machine learning model training process, including:
[0196] A recording unit 601 is used to record the maximum absolute value of the gradient grad after adjusting the loss scale during the machine learning model training process;
[0197] The adjustment unit 602 is configured to determine a subsequent scale window based on the previous adjustment operation on the loss scale, including:
[0198] The adjusting unit 602 is configured to, when the maximum absolute value of the grad is greater than a third threshold, increase the scale window used in the previous iteration to obtain the scale window used subsequently.
[0199] In one implementation, the scale window used later is one of a plurality of preset scale windows; the plurality of preset scale windows respectively correspond to a different number of model iterations.
[0200] In one implementation, the adjusting unit 602 is configured to adjust the loss scale according to a previous loss scale adjustment operation, including:
[0201] An adjustment unit 602 is configured to determine a ratio of a loss scale used in a previous iteration to a loss scale used thereafter based on a previous adjustment operation on the loss scale;
[0202] The adjusting unit 602 is further configured to adjust the loss scale according to a ratio of the loss scale used in the previous iteration to the loss scale used subsequently.
[0203] In one implementation, the machine learning model is a model for processing at least one of text data, image data, or audio data.
[0204] Figure 12 is a schematic diagram of the structure of another data processing device provided in this embodiment. The data processing device 70 may be a chip or system-on-chip. Specifically, the data processing device 70 may include all or part of the hardware in a desktop computer, tablet computer, desktop, laptop, handheld computer, notebook computer, ultra-mobile personal computer, netbook, as well as a cellular phone, personal digital assistant, augmented reality / virtual reality device, various artificial intelligence hardware accelerators, and other devices capable of model training.
[0205] The data processing device 70 may include: a processor 701 , a communication line 702 , a memory 703 , and part or all of at least one communication interface 704 .
[0206] The processor 701 is used to execute all or part of the steps executed by the electronic device in the method provided in FIG. 2 to FIG. 10 of this embodiment.
[0207] Specifically, the processor 701 may include a general-purpose central processing unit (CPU), and the processor 701 may also include a microprocessor, a field programmable gate array (FPGA), a digital signal processor (DSP) or an application-specific integrated circuit (ASIC), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0208] In a specific implementation, as an embodiment, the processor 701 may include one or more CPUs, such as CPU0 and CPU1 in FIG12 .
[0209] In a specific implementation, as an embodiment, the data processing device 70 may include multiple processors, such as the processor 701 and the processor 708 in FIG12 . Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. The processor herein may refer to one or more devices, circuits, and / or processing cores for processing, for example, computer data (computer program instructions).
[0210] In addition, the memory 703 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). Memory 703 can be independent and connected to processor 701 via communication line 702. Memory 703 can also be integrated with processor 701.
[0211] The memory 703 stores computer instructions. The processor 701 can be used to execute all or part of the steps in the model training method provided in this embodiment by executing the computer instructions stored in the memory 703. As shown in Figure 12, the computer instructions stored in the memory 703 may include software modules for implementing the functions of each functional unit in the above-mentioned data processing device 60. The processor 701 can be used to execute the model training method provided in this embodiment by executing the computer instructions stored in the memory 703.
[0212] Optionally, the computer-executable instructions in this embodiment may also be referred to as application code, which is not specifically limited in this embodiment.
[0213] In addition, the communication interface 704 uses any transceiver or other device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc.
[0214] In addition, the communication line 702 is used to connect the various components in the data processing device 70. Specifically, the communication line 702 may include a data bus, a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as communication lines 702 in the figure.
[0215] In a specific implementation, as an embodiment, the data processing device 70 may further include an output device 707 and an input device 706. The output device 707 may communicate with the processor 701 and may display information in a variety of ways. For example, the output device 707 may be a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector. The input device 706 may communicate with the processor 701 and may receive user input in a variety of ways. For example, the input device 706 may be a mouse, a keyboard, a touch screen device, or a sensor device.
[0216] In addition, the data processing device 70 may further include a storage medium 705. The storage medium 705 is used to store computer instructions and various data for implementing the technical solution of this embodiment. When the data processing device 70 executes the above-mentioned model training method of this embodiment, the computer instructions and various data stored in the storage medium 705 are loaded into the memory 703, so that the processor 701 can execute the computer instructions stored in the memory 703 to execute the model training method provided by this embodiment.
[0217] The method steps in this embodiment can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in RAM, flash memory, ROM, PROM, EPROM, EEPROM, registers, hard disk, mobile hard disk, CD-ROM or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a data processing device. Of course, the processor and the storage medium can also exist in the data processing device as discrete components.
[0218] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in this embodiment are performed in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a communication device, a user device, or other programmable device. The computer program or instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions may be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; an optical medium, such as a digital video disc (DVD); or a semiconductor medium, such as an SSD.
[0219] In this embodiment, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different implementations are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form a new embodiment based on their internal logical relationships.
[0220] In this embodiment, "at least one" means one or more, "more than one" means two or more, and other quantifiers are similar. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, for elements (element) in the singular form "a", "an" and "the", unless the context clearly stipulates otherwise, it does not mean "one or only one", but means "one or more than one". For example, "a device" means one or more such devices. Furthermore, at least one (at least one of)..." means one or any combination of the subsequent associated objects, for example, "at least one of A, B and C" includes A, B, C, AB, AC, BC, or ABC. In the text description of this embodiment, the character " / " generally indicates that the previous and next associated objects are in an "or" relationship; in the formula of this embodiment, the character " / " indicates that the previous and next associated objects are in a "division" relationship.
Claims
1. A data processing method, the method being applied to an electronic device, the electronic device being used to train a machine learning model, characterized in that: The method comprises: During the training process of the machine learning model, recording the adjustment operation of the loss scale; Adjust the loss scale based on the previous loss scale adjustment operation.
2. The method according to claim 1, characterized in that The adjusting the loss scale according to the previous adjustment operation on the loss scale includes: According to the previous adjustment operation of the loss scale, the scale window used later is determined; Adjust the loss scale according to the scale window adopted later.
3. The method according to claim 2, characterized in that The step of determining the scale window to be used later based on the previous adjustment operation of the loss scale includes: When the previous adjustment operation of the loss scale meets a first preset condition, the scale window is adjusted upward to obtain the scale window used later; the first preset condition includes: after the last adjustment of the scale window, the number of times the loss scale is adjusted upward reaches a first threshold.
4. The method according to claim 2 or 3, characterized in that: The step of determining the scale window to be used later based on the previous adjustment operation of the loss scale includes: When the previous adjustment operation on the loss scale meets a second preset condition, the scale window is adjusted downward to obtain the scale window used later; the second preset condition includes: the number of consecutive downward adjustment operations on the loss scale reaches a second threshold.
5. The method according to claim 2, characterized in that: The operation of adjusting the loss scale during the machine learning model training process is recorded, including: During the training process of the machine learning model, recording the value of the loss scale after the adjustment operation; The step of determining the scale window to be used later based on the previous adjustment operation of the loss scale includes: According to the value of the currently used loss scale, the scale window used later is determined; wherein, loss scales of different sizes correspond to different preset levels, and different preset levels correspond to different scale windows.
6. The method according to claim 2, characterized in that The operation of adjusting the loss scale during the machine learning model training process is recorded, including: During the training process of the machine learning model, the maximum absolute value of the gradient grad is recorded after the loss scale is adjusted; The step of determining the scale window to be used later based on the previous adjustment operation of the loss scale includes: When the maximum absolute value of the grad is greater than a third threshold, the scale window used in the previous iteration is adjusted upward to obtain the scale window used subsequently.
7. The method according to any one of claims 2 to 6, characterized in that: The scale window used later is one of a plurality of preset scale windows; the plurality of preset scale windows respectively correspond to a different number of model iterations.
8. The method according to claim 1, characterized in that: The adjusting the loss scale according to the previous adjustment operation on the loss scale includes: According to the previous adjustment operation of the loss scale, determine the ratio of the loss scale used in the previous iteration to the loss scale used later; The loss scale is adjusted according to the ratio of the loss scale used in the previous iteration to the loss scale used thereafter.
9. The method according to any one of claims 1 to 8, characterized in that: The machine learning model is a model used to process at least one of text data, image data or audio data.
10. A data processing device, characterized in that: The data processing device is applied to an electronic device, and the electronic device is used to train a machine learning model. The data processing device includes: A recording unit, used to record the adjustment operation of the loss scale during the training process of the machine learning model; The adjustment unit is used to adjust the loss scale according to the previous adjustment operation on the loss scale.
11. A resource changing device, characterized in that: The system comprises a memory and a processor, wherein the memory is used to store computer instructions, and the processor is used to call and execute the computer instructions from the memory to implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer program product, characterized in that The computer program product comprises instructions which, when executed on a processor, implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Loss-scaling for deep neural network training with reduced precision
CN110073371A
Deep learning model training method, device and system based on mixed precision
CN110163368A
Predicting deep learning scaling
CN111260021A
Neural network model training method and device
CN113762502A
Method and device for training neural network, and computer readable storage medium
CN114580625A