Task Processing Method, Device, Electronic Device and Storage Medium
By sliding average smoothing the statistical values of the normalized layer during the Transformer model training process, and fixed as the third statistical value after training, the problems of high memory consumption and batch normalized operator performance degradation caused by the layer normalized operator are solved, and the effect of improving model performance under low memory consumption is achieved.
Patent Information
- Application Number
- CN202210709171.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-06-21
AI Technical Summary
The layer normalization operator in the native Transformer model results in high computational memory consumption, while replacing it with a batch normalization operator will lead to performance degradation.
During the Transformer model training process, the sliding average strategy is used to smooth the statistical value of the normalized layer. After the training is completed, it is fixed as the third statistical value for inference calculation, and the third statistical value is determined using the exponential average strategy.
While reducing computational memory consumption, the processing performance of the model is improved, the oscillation of batch statistics is alleviated, and the stability and efficiency of task processing are ensured.
Smart Images

Figure CN115062765B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of chips and neural networks, and particularly to a task processing method, apparatus, electronic device, and storage medium. Background Art
[0002] In recent years, models with Transformer architectures have achieved surprising performance in natural language processing and computer vision tasks, becoming a research hotspot and focus in the current field of artificial intelligence.
[0003] However, the layer normalization operator (Layer Normalization) included in the native Transformer brings high computational memory consumption. At the same time, the layer normalization operator is not supported on many edge devices. If the batch normalization operator (Batch Normalization) with low computational memory is used to replace the layer normalization operator in the Transformer architecture, it will lead to a significant performance degradation. Summary of the Invention
[0004] In view of this, this application provides a task processing method, apparatus, electronic device, and storage medium to optimize the task processing performance of the Transformer model.
[0005] Specifically, this application is implemented through the following technical solutions:
[0006] According to the first aspect of the embodiments of this application, a task processing method is provided, including:
[0007] During the training process of the Transformer model, for any normalization layer in the Transformer model, determine the first statistical value of the current batch of this normalization layer. Based on this first statistical value and the statistical values of the historical batches of this normalization layer, use the moving average strategy to smooth this first statistical value to obtain the second statistical value, and use the second statistical value of this normalization layer for forward or backward propagation;
[0008] During the process of using the trained Transformer model to process tasks, for any normalization layer in the Transformer model, fix the statistical value of this normalization layer as the third statistical value for inference calculation; where the third statistical value is determined based on the second statistical value using the exponential average strategy during the training process of the Transformer model.
[0009] According to the second aspect of the embodiments of this application, a task processing apparatus is provided, including:
[0010] A training unit, which is used to determine a first statistical value of any normalization layer in the Transformer model for the current batch during the training of the Transformer model, smooth the first statistical value by using an exponential moving average strategy based on the first statistical value and the statistical values of historical batches of the normalization layer to obtain a second statistical value, and perform forward or backward propagation by using the second statistical value of the normalization layer;
[0011] An inference unit, which is used to fix the statistical value of any normalization layer in the Transformer model to a third statistical value for inference calculation during the process of using the trained Transformer model to process tasks; wherein, the third statistical value is determined by using an exponential moving average strategy based on the second statistical value during the training of the Transformer model.
[0012] According to a third aspect of the embodiments of the present application, there is provided an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor is configured to execute the machine-executable instructions to implement the method provided in the first aspect.
[0013] According to a fourth aspect of the embodiments of the present application, there is provided a machine-readable storage medium, where machine-executable instructions are stored in the machine-readable storage medium, and when the machine-executable instructions are executed by a processor, the method provided in the first aspect is implemented.
[0014] The technical solution provided by the present application can at least bring the following beneficial effects:
[0015] During the training of the Transformer model, for any normalization layer in the Transformer model, a first statistical value of the normalization layer for the current batch is determined, and based on the first statistical value and the statistical values of historical batches of the normalization layer, the first statistical value is smoothed by using an exponential moving average strategy to obtain a second statistical value, and forward or backward propagation is performed by using the second statistical value of the normalization layer, which alleviates the oscillation of batch statistical values and improves the model performance. Furthermore, when using the trained Transformer model to process tasks, the statistical values of each normalization layer are fixed to the third statistical value for inference calculation, which ensures the processing performance while reducing the computational memory consumption of the Transformer model. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a schematic flowchart of a task processing method shown in an exemplary embodiment of the present application;
[0017] Figure 2It is a schematic diagram of the architecture of a feature extraction network in a neural network with a Transformer structure shown in an exemplary embodiment of the present application;
[0018] Figure 3A It is a schematic diagram of the calculation dimension of an online normalization method shown in an exemplary embodiment of the present application;
[0019] Figure 3B It is a schematic diagram of the calculation dimension of an offline normalization method shown in an exemplary embodiment of the present application;
[0020] Figure 4 It is a schematic diagram of the structure of a task processing device shown in an exemplary embodiment of the present application;
[0021] Figure 5 It is a schematic diagram of the hardware structure of an electronic device shown in an exemplary embodiment of the present application. Detailed implementation manners
[0022] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0023] The terms used in the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0024] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of the present application and to make the above objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the drawings.
[0025] Please refer to Figure 1 , which is a schematic flowchart of a task processing method provided by an embodiment of the present application. As Figure 1 shown, the task processing method may include the following steps:
[0026] It should be noted that the sequence numbers of the steps in the embodiments of the present application do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0027] Step S100. During the training process of the Transformer model, for any normalization layer in the Transformer model, determine the first statistical value of the current batch of this normalization layer. Based on this first statistical value and the statistical values of the historical batches of this normalization layer, use the moving average strategy to smooth the first statistical value to obtain the second statistical value, and use the second statistical value of this normalization layer for forward or backward propagation.
[0028] In the embodiments of the present application, in order to reduce the computational memory consumption of the Transformer model, the normalization operator of the Transformer model can be replaced with a batch normalization operator. And in order to reduce the performance degradation caused by the batch normalization operator, during the training process of the Transformer model, when determining the statistical value of the normalization layer, instead of using the statistical value of a single batch as the final statistical value, the moving average strategy is adopted to smooth the statistical value of a single batch to alleviate the oscillation of the batch statistical value and improve the model performance.
[0029] Correspondingly, during the training process of the Transformer model, for any normalization layer in the Transformer model, the statistical value of the current batch of this normalization layer (which can also be called the statistical value at the current moment, denoted as the first statistical value in this article) can be determined, and based on this first statistical value and the statistical values of the historical batches of this normalization layer (which can also be called the statistical values at historical moments), use the moving average strategy to smooth the first statistical value to obtain the smoothed statistical value (referred to as the second statistical value in this article), and use the second statistical value of this normalization layer for forward or backward propagation.
[0030] Exemplarily, for any normalization layer in the Transformer model, on the one hand, the input data can be normalized using the second statistical value determined for this batch during the training process of this normalization layer; on the other hand, based on the second statistical value determined during the training process of this normalization layer, the third statistical value can be determined through an iterative optimization method. For example, the third statistical value can be obtained by performing exponential smoothing on the second statistical value.
[0031] Step S110. During the process of using the trained Transformer model to process tasks, for any normalization layer in the Transformer model, fix the statistical value of this normalization layer as the third statistical value for inference calculation; where the third statistical value is determined using the exponential average strategy based on the second statistical value during the training process of the Transformer model.
[0032] In the embodiments of the present application, when the training of the Transformer model is completed in the manner described in step S100, the trained Transformer model can be used for task processing.
[0033] Exemplarily, for any normalization layer of the Transformer model, the third statistical value of the normalization layer can be determined by using an exponential moving average strategy according to the second statistical value of the normalization layer during the training process of the Transformer model.
[0034] Exemplarily, the third statistical value of any normalization layer of the Transformer model is finally determined and fixed at the end of training. During the inference process, for any normalization layer of the trained Transformer model, the statistical value of the normalization layer can be fixed as the third statistical value for inference calculation. For example, the input data is normalized by using the third statistical value of the normalization layer.
[0035] Exemplarily, the above-mentioned task processing includes but is not limited to NLP (Natural Language Processing) tasks (such as machine translation), or Speech tasks (such as speech recognition), or CV (Computer Vision) tasks (such as classification, object detection, or instance segmentation tasks, etc.).
[0036] It can be seen that in Figure 1 In the shown method flow, during the training of the Transformer model, for any normalization layer in the Transformer model, the first statistical value of the current batch of the normalization layer is determined. According to the first statistical value and the statistical values of the historical batches of the normalization layer, the first statistical value is smoothed by using a moving average strategy to obtain a second statistical value, and the second statistical value of the normalization layer is used for forward or backward propagation, which alleviates the oscillation of the batch statistical value and improves the model performance. Furthermore, when using the trained Transformer model for task processing, the statistical values of each normalization layer are fixed as the third statistical value for inference calculation, which ensures the processing performance while reducing the computational memory consumption of the Transformer model.
[0037] In some embodiments, the determining the first statistical value of the current batch of the normalization layer, according to the first statistical value and the statistical values of the historical batches of the normalization layer, and smoothing the first statistical value by using a moving average strategy to obtain a second statistical value includes:
[0038] During the forward propagation process, the squared mean operation is performed on the input activation values of the current batch of this normalization layer to obtain the first activation value statistic of the current batch of this normalization layer;
[0039] Based on the first activation value statistic of the current batch of this normalization layer and the activation value statistics of the historical batches of this normalization layer, using the geometric mean strategy, the first activation value statistic of the current batch of this normalization layer is smoothed to obtain the second activation value statistic.
[0040] Exemplarily, in order to improve the training stability when using the smoothing strategy, during the forward propagation process of training the Transformer model, for any normalization layer, when calculating the activation value statistic of the current batch of this normalization layer, the zero-mean constraint can be removed, and the squared mean operation is performed on the input activation values of the current batch of this normalization layer to obtain the first statistic (which can be called the first activation value statistic) of the current batch of this normalization layer.
[0041] Exemplarily, in order to suppress the oscillation of the batch statistic, the first activation value statistic determined in the above manner can be smoothed using the moving average strategy.
[0042] Exemplarily, during the forward propagation process, based on the first activation value statistic of the current batch of this normalization layer and the activation value statistics of the historical batches of this normalization layer, using the geometric mean strategy, the geometric mean of the first activation value statistic of the current batch of this normalization layer and the activation value statistics of the historical batches of this normalization layer is calculated to obtain the second activation value statistic.
[0043] Exemplarily, the number of activation value statistics of the historical batches participating in the geometric mean calculation in the above process can be determined by the size of the preset sliding window.
[0044] For example, assuming the size of the sliding window is M (M≥2), then (M - 1) activation value statistics of the historical batches need to be used to participate in the geometric mean calculation.
[0045] It should be noted that, in the embodiments of the present application, the above third statistical value may be a third activation value statistical value, and the third activation value statistical value may be smoothed from the second activation value statistical values of each batch during the training process of the Transformer model. For example, during the training process of the Transformer model, for any normalization layer, during the forward propagation process of any batch, when the second activation value statistical value of the current batch of this normalization layer is determined, the third activation value statistical value of this batch of this normalization layer can be determined according to the second activation value statistical values of each batch of this normalization layer by using the exponential smoothing strategy; furthermore, when the training end condition is reached, the third activation value statistical value determined for the last batch of this normalization layer can be used as the third activation value statistical value of this normalization layer for final inference.
[0046] In some embodiments, the above determining the first statistical value of the current batch of this normalization layer, and smoothing the first statistical value by using the moving average strategy according to the first statistical value and the statistical values of the historical batches of this normalization layer to obtain the second statistical value includes:
[0047] During the backpropagation process, determining the first gradient statistical value of the current batch of this normalization layer according to the second activation value statistical value of the current batch of this normalization layer;
[0048] Smoothing the first gradient statistical value of the current batch of this normalization layer by using the algorithm average and exponential average strategies according to the first gradient statistical value of the current batch of this normalization layer and the first gradient statistical values of the historical batches of this normalization layer to obtain the second gradient statistical value.
[0049] Exemplarily, during the backpropagation process of training the Transformer model, for any normalization layer, the derivative of the second activation value statistical value of the current batch of this normalization layer can be obtained through the chain rule to obtain the first gradient statistical value of the current batch of this normalization layer.
[0050] Exemplarily, since during the forward propagation, the calculation of the second activation value statistical value incorporates the historical activation value statistical value information, and the gradient brought in by the historical activation value statistical value cannot be accurately derived according to the chain rule. Therefore, in order to more accurately estimate the gradient, the determined first gradient statistical value can be smoothed by using the moving average strategy to obtain the second gradient statistical value, and the gradient from the historical activation value statistical value is approximately determined by smoothing.
[0051] Exemplarily, during the backpropagation process, based on the first gradient statistic value of the current batch of this normalization layer and the gradient statistic values of the historical batches of this normalization layer, using the arithmetic mean and exponential mean strategies, calculate the arithmetic mean and exponential mean of the first gradient statistic value of the current batch of this normalization layer and the gradient statistic values of the historical batches of this normalization layer to obtain a second gradient statistic value.
[0052] Exemplarily, the number of gradient statistic values of the historical batches participating in the calculation of the arithmetic mean and exponential mean in the above process can be determined by the size of a preset sliding window.
[0053] For example, assuming the size of the sliding window is M (M ≥ 2), then (M - 1) gradient statistic values of historical batches need to be used to participate in the calculation of the arithmetic mean and exponential mean.
[0054] In some embodiments, for any normalization layer in the Transformer model, after determining the first statistic value of the current batch of this normalization layer, before smoothing the first statistic value using the moving average strategy based on the first statistic value and the statistic values of the historical batches of this normalization layer to obtain a second statistic value, it may further include:
[0055] Determine whether the current batch of this normalization layer meets the smoothing condition based on the first statistic value of the current batch of this normalization layer;
[0056] In the case of determining that the current batch of this normalization layer meets the smoothing condition, determine to perform the operation of smoothing the first statistic value using the moving average strategy based on the first statistic value and the statistic values of the historical batches of this normalization layer to obtain a second statistic value.
[0057] Exemplarily, considering that if there are large outliers in the batch statistic values, using the moving average strategy will make the training of the Transformer model unstable (the gradient approximation error becomes larger). Therefore, in order to improve the stability of the Transformer model training and reduce the gradient approximation error introduced by the moving average strategy in the process of improving the performance of the Transformer model, an adaptive judgment strategy can be introduced to determine whether to adopt the moving average strategy to reduce the influence of outliers on the moving average strategy.
[0058] Accordingly, during the training of the Transformer model, for any normalization layer, after determining the first statistical value of the current batch of this normalization layer (such as the above-mentioned first activation value statistical value or the first gradient statistical value), it is possible to determine whether the current batch of this normalization layer meets the smoothing processing condition based on the first statistical value of the current batch of this normalization layer, and when it is determined that the current batch of this normalization layer meets the smoothing processing condition, based on this first statistical value and the statistical values of the historical batches of this normalization layer, using the moving average strategy, perform smoothing processing on this first statistical value to obtain the second statistical value.
[0059] It should be noted that when it is determined that the current batch of this normalization layer does not meet the smoothing processing condition, it is not necessary to perform smoothing processing on the first statistical value using the moving average strategy, but it is possible to determine the first statistical value of the current batch of this normalization layer as the second statistical value, that is, when it is determined that the current batch of this normalization layer does not meet the smoothing processing condition, the statistical value of this normalization layer uses the first statistical value.
[0060] In one example, the current batch of this normalization layer meeting the smoothing processing condition may include:
[0061] The absolute value of the difference between the arithmetic mean of the first statistical value of the current batch of this normalization layer and the first statistical value of the historical batches and the geometric mean of the first statistical value of the current batch of this normalization layer and the first statistical value of the historical batches is less than or equal to the first threshold;
[0062] Among them, the first threshold is determined based on the variance of the historical statistical values within the previous batch record window.
[0063] Exemplarily, considering that when performing smoothing processing on the first statistical value of the current batch of this normalization layer using the moving average strategy, if the absolute value of the difference between the arithmetic mean of the first statistical values of each batch within the sliding window and the geometric mean of the first statistical values of each batch within the sliding window is too large, it indicates that there are outliers in the first statistical value of the current batch. In this case, if the moving average strategy is used for smoothing processing, it will cause the training of the Transformer model to be unstable.
[0064] For example, taking the sliding window size as M, the arithmetic mean of the first statistical value of the current batch and the first statistical values of the previous M - 1 batches can be calculated, and the geometric mean of the first statistical value of the current batch and the first statistical values of the previous M - 1 batches can be calculated. Based on the absolute value of the difference between this arithmetic mean and geometric mean, it is determined whether the current batch of this normalization layer meets the smoothing processing condition.
[0065] Exemplarily, when determining whether the current batch of the normalization layer meets the smoothing condition, a determination threshold (which can be referred to as the first threshold) can be determined based on the variance of the first statistical value of the historical batches of the normalization layer, and is used to determine whether the absolute value of the difference between the above arithmetic mean and the geometric mean is too large. When the absolute value of the difference between the arithmetic mean and the geometric mean is less than or equal to the first threshold, it is determined that the current batch of the normalization layer meets the smoothing condition; otherwise, it is determined that the current batch of the normalization layer does not meet the smoothing condition.
[0066] For example, assuming that the current time is time t and the sliding window size is M, the above first threshold can be determined based on the first statistical value of the normalization layer at time t - 1 and the variances of the first statistical values of the previous M - 1 batches.
[0067] In another example, the current batch of the normalization layer meeting the smoothing condition may include:
[0068] The absolute value of the first statistical value of the current batch of the normalization layer is less than or equal to a second threshold;
[0069] Among them, the second threshold is determined based on the standard deviation of the statistical value distribution of the historical batches of the normalization layer.
[0070] Exemplarily, considering that in the case of smoothing the first statistical value of the current batch of the normalization layer using the moving average strategy, if the first statistical value of the current batch of the normalization layer differs too much from the standard deviation of the distribution of the first statistical values of the historical batches of the normalization layer, it indicates that there are outliers in the first statistical value of the current batch of the normalization layer. At this time, if the moving average strategy is used for smoothing in this case, it will cause the training of the Transformer model to be unstable.
[0071] Exemplarily, since the normalization calculation is performed using the batch normalization operator according to the channel dimension, the calculated statistical value (such as the above first statistical value) can be in the form of a C - element one - dimensional vector, that is, there are a total of C elements (one element corresponds to one channel, and C is the number of channels).
[0072] Correspondingly, the absolute value of the above first statistical value being less than or equal to the second threshold can mean that the absolute value of any element of the first statistical value is less than or equal to the corresponding second threshold.
[0073] Exemplarily, the second threshold can be determined based on the standard deviation of the statistical value distribution of the historical batches of the normalization layer.
[0074] For example, assuming that the sliding window size is M, for the second threshold corresponding to the first statistical value of the current batch of the normalization layer, it can be determined based on the standard deviation of the distribution of the first statistical values of the previous M - 1 batches of the current batch.
[0075] It should be noted that in the embodiments of the present application, the determination of whether the current batch of the normalization layer meets the smoothing processing condition emphasizes whether, during the forward / backward propagation process, it is necessary to determine the second statistical value according to the first statistical value by using the moving average strategy.
[0076] Exemplarily, when the current batch of the normalization layer meets the smoothing processing condition, it is necessary to determine the second statistical value according to the first statistical value by using the moving average strategy; when the current batch of the normalization layer does not meet the smoothing processing condition, the moving average strategy may not be used to determine the second statistical value. For example, the first statistical value of the current batch of the normalization layer may be determined as the second statistical value of the current batch of the normalization layer.
[0077] Exemplarily, during the forward propagation process of any batch, it is necessary to update the third activation value statistic determined in the previous batch according to the second activation value statistic of the current batch by using the exponential smoothing strategy to obtain the third activation value statistic of the current batch. That is, regardless of whether it is necessary to use the moving average strategy to determine the second statistical value, it is necessary to use the exponential average strategy to determine the third statistical value. The above smoothing processing condition is a constraint condition for the moving average strategy, but does not constrain the exponential average strategy.
[0078] To enable those skilled in the art to better understand the technical solutions provided in the embodiments of the present application, the technical solutions provided in the embodiments of the present application will be described below with reference to specific examples.
[0079] First, a simple description of the neural network with the Transformer structure will be given below.
[0080] The neural network with the Transformer structure may include an encoder part (i.e., the feature extraction network in the neural network with the Transformer structure) and a decoder part.
[0081] Please refer to Figure 2 , which is a schematic diagram of an architecture of the feature extraction network in the neural network with the Transformer structure provided in the embodiments of the present application. As Figure 2As shown in the figure, the feature extraction network in the neural network of the Transformer structure includes an embedding layer and at least one Transformer layer. One Transformer layer includes two basic modules, where one basic module includes a normalization layer, a multi-head attention layer, and an add layer; the other basic module includes a normalization layer, a feed forward neural network layer, and an add layer. That is, after the original sequence information to be processed is processed by the feature extraction network in the neural network of the Transformer structure, the feature information of the entire original sequence information to be processed can be obtained.
[0082] Among them, this feature information is a kind of feature information suitable for computer processing of the sequence information to be processed.
[0083] Exemplarily, the input to be processed by the Transformer is sequence information, which can include but is not limited to text, speech, or image blocks. Therefore, the Transformer structure can be used for natural language processing tasks, such as text similarity, text classification, machine translation, etc.; it can also be applied to speech tasks, such as speech recognition; it can also be used for visual tasks, such as image classification, object detection, and semantic segmentation, etc.
[0084] Next, the above-mentioned embedding layer and multi-head attention layer will be specifically introduced in combination with specific examples.
[0085] Taking text as the input sequence as an example, after obtaining the text to be processed, the embedding layer can perform embedding processing on each word in the text to be processed to obtain the initial feature information of each word. The text to be processed can be a paragraph of text or a sentence. The text can be Chinese text, English text, or text in other languages.
[0086] Specifically, in some embodiments, as Figure 2 shown, the embedding layer includes an input embedding layer and a positional encoding layer.
[0087] Exemplarily, in the input embedding layer, word embedding processing can be performed on each word in the text to be processed, so as to obtain the word embedding tensors of each word. The tensors can specifically be represented as one-dimensional vectors, two-dimensional matrices, three-dimensional or more-dimensional data, etc.
[0088] Exemplarily, in the positional encoding layer, the positions of each word in the text to be processed can be obtained, and then, position tensors are generated for the positions of each word.
[0089] In some examples, the positions of each word can be the absolute positions of each word in the text to be processed.
[0090] Taking the text to be processed as "The weather is really good today" as an example, the position of "今" can be expressed as the first position, the position of "天" can be expressed as the second position, and so on.
[0091] In some examples, the positions of the words may be relative positions between the words.
[0092] Still taking the to-be-processed text "The weather is really nice today" as an example, the position of "今" can be expressed as before "天", the position of "天" can be expressed as after "今", before "天", and so on.
[0093] After obtaining the word embedding tensor and position tensor of each word in the text to be processed, the position tensor and word embedding tensor of each word can be combined to obtain the initial feature information of each word, thereby obtaining the initial feature information corresponding to the text to be processed.
[0094] For example, the multi-head attention mechanism can simply and effectively abstract contextual dependencies and capture syntactic and semantic features.
[0095] In the multi-head attention mechanism, input features are linearly mapped to different information subspaces through different fully-connected layers, and the same attention calculation is performed in each subspace to fully learn the underlying structure and semantics of the text. The output of each information subspace is combined through concatenation and then fused through a fully-connected layer.
[0096] like Figure 3A and Figure 3B As shown in the figure, normalization methods can be divided into two categories: one is the "online normalization method", represented by the layer normalization operator (Layer Normalization). This type of method still needs to dynamically calculate statistical values during inference and cannot be merged with adjacent linear operators. Therefore, it cannot accelerate the inference process; the other is the "offline normalization method", represented by the batch normalization operator (Batch Normalization). This type of method calculates statistical values in each corresponding channel dimension (Channel) and uses fixed statistical values for forward calculation during inference. Therefore, it can be merged with adjacent linear layers to accelerate model inference.
[0097] Considering that the layer normalization operator is inefficient in hardware deployment and consumes a lot of computational memory, in order to improve the deployment efficiency of the Transformer model and reduce computational memory consumption, the normalization operator in the Transformer model can be replaced with a batch normalization operator.
[0098] In order to reduce the performance degradation caused by the batch normalization operator, during the training of the Transformer model, a moving average strategy is used to correct the inference statistics to improve performance. At the same time, the "outlier filtering" strategy is applied to adaptively adjust the application of the moving average strategy for each layer.
[0099] When performing inference using the trained Transformer model, the statistics required for normalization can be fixed, and the Norm operator can be merged with adjacent linear operators to accelerate inference.
[0100] Exemplarily, considering that replacing the layer normalization operator in the Transformer with the batch normalization operator will lead to performance degradation, the following problems exist:
[0101] Problem 1: The distribution range of batch statistics is large (there are oscillations), resulting in inaccurate inference statistics, which in turn causes performance degradation;
[0102] Problem 2: There are large outliers in the distribution of batch statistics, which will bring instability to training.
[0103] For Problem 1, in the embodiments of the present application, the constraint of zero mean is removed, the squared mean is used as the statistics for scaling normalization, and it is proposed to use the moving average strategy both in the forward and backward normalization to suppress the oscillation of batch statistics, thereby calibrating the inference statistics and improving the model performance.
[0104] For Problem 2, in the embodiments of the present application, an adaptive judgment strategy is introduced to reduce the impact of outliers on the moving average strategy, thereby stabilizing the training process.
[0105] The specific implementation method is as follows:
[0106] a) Use the moving average strategy. During training, in the forward process, the geometric mean strategy is used to smooth the impact of the oscillation of activation value statistics on the inference statistics, and in the backward process, the arithmetic mean and exponential mean strategies are used to smooth the gradient statistics.
[0107] Exemplarily, assuming that the moving window size is M, during the forward propagation process, the first activation value statistics of the current batch is the first activation value statistics at time t Then the following formula can be used to smooth the first activation value statistics to obtain the second activation value statistics:
[0108]
[0109] Among them, is the first activation value statistics at time (t - i), that is, the first activation value statistics of the i-th batch between the current batches.
[0110] Assume that during the backpropagation process, the first gradient statistic value of the current batch is the first gradient statistic value at time t. Then, the following formula can be used to smooth the first gradient statistic value to obtain the second activation value statistic:
[0111]
[0112]
[0113] where is the second gradient statistic value at time t, is the second gradient statistic value at time t - 1, is the first gradient statistic value at time t - i.
[0114] b) Introduce an "outlier filtering" strategy to adaptively adjust the application of the moving average strategy and reduce the impact of outliers on training stability.
[0115] Example 1
[0116] Use the stored historical statistic values to determine the arithmetic mean AM and the geometric mean GM, and then use the upper bound of the inequality |AM - GM| to select an adaptive threshold at time t:
[0117]
[0118] where is the arithmetic mean of the first statistic values within the moving window M at time t, is the geometric mean of the first statistic values within the moving window M at time t, Var[σ t-1 is the variance after taking the square root of the first statistic values within the moving window M at time t - 1, and M is the size of the moving window.
[0119] Example 2
[0120] Use the stored historical statistic values n times the standard deviation (std) of the distribution as the adaptive threshold:
[0121]
[0122] where is the first statistic value at time t, is the standard deviation of the first statistic values from time (t - M + 1) to time t - 1, and n is an empirical value that can be set according to actual needs, such as 3 or 5 or 11, etc.
[0123] If the above inequality 1) or 2) does not hold, it is determined that there is an outlier in the first statistical value of the current batch (at time t). The sliding average strategies in the above forward propagation process / backward propagation process will be skipped. In this case, the first statistical value can be used (without further calculating the second statistical value, or determining that the second statistical value is equal to the first statistical value).
[0124] It should be noted that for the forward propagation of any batch, if the forward propagation of this batch (assumed to be at time t) meets the smoothing condition, the first activation value / gradient statistical value update window M can be used; if the forward propagation of this batch does not meet the smoothing condition, for example, the activation value statistical value of this batch is judged to contain an outlier, that is, it is not a representative statistical value, then in order to avoid the influence of the activation value statistical value of this batch on the sliding average strategy, the third activation value statistical value determined in the previous batch (that is, the third activation value statistical value at time t - 1) can be used as the first activation value statistical value at time t to update the window M; for the update of the window M of the first gradient statistical value, the second gradient statistical value at time t - 1 is used as the first gradient statistical value at time t for window M update.
[0125] It can be seen that by using the sliding average strategy to calibrate the inference statistical value, the performance of the model can be prevented from significantly degrading. However, using the sliding window strategy to approximately estimate the backward gradient will cause training instability due to the existence of outliers. Therefore, an adaptive judgment strategy is also needed to judge whether to use the sliding window strategy at the current moment to avoid training collapse.
[0126] The method provided in this application has been described above. Next, the device provided in this application will be described:
[0127] Please refer to Figure 4 , which is a schematic structural diagram of a task processing device provided in an embodiment of this application. As Figure 4 shown, the task processing device may include:
[0128] A training unit 410, configured to, during the training of the Transformer model, for any normalization layer in the Transformer model, determine the first statistical value of the current batch of this normalization layer, and based on the first statistical value and the statistical values of the historical batches of this normalization layer, use the sliding average strategy to smooth the first statistical value to obtain a second statistical value, and perform forward or backward propagation using the second statistical value of this normalization layer;
[0129] An inference unit 420, which is used to fix the statistical value of any normalization layer in the Transformer model to a third statistical value for inference calculation during the process of task processing using the trained Transformer model; wherein, the third statistical value is determined by using an exponential moving average strategy based on the second statistical value during the training process of the Transformer model.
[0130] In some embodiments, the training unit 410 determines the first statistical value of the current batch of the normalization layer, and based on the first statistical value and the statistical values of the historical batches of the normalization layer, uses a moving average strategy to smooth the first statistical value to obtain a second statistical value, including:
[0131] During the forward propagation process, perform a squared average operation on the input activation values of the current batch of the normalization layer to obtain the first activation value statistical value of the current batch of the normalization layer;
[0132] Based on the first activation value statistical value of the current batch of the normalization layer and the activation value statistical values of the historical batches of the normalization layer, use a geometric mean strategy to smooth the first activation value statistical value of the current batch of the normalization layer to obtain a second activation value statistical value.
[0133] In some embodiments, the training unit 410 determines the first statistical value of the current batch of the normalization layer, and based on the first statistical value and the statistical values of the historical batches of the normalization layer, uses a moving average strategy to smooth the first statistical value to obtain a second statistical value, including:
[0134] During the backward propagation process, determine the first gradient statistical value of the current batch of the normalization layer based on the second activation value statistical value of the current batch of the normalization layer;
[0135] Based on the first gradient statistical value of the current batch of the normalization layer and the first gradient statistical values of the historical batches of the normalization layer, use an arithmetic average and exponential moving average strategy to smooth the first gradient statistical value of the current batch of the normalization layer to obtain a second gradient statistical value.
[0136] In some embodiments, for any normalization layer in the Transformer model, after the training unit 410 determines the first statistical value of the current batch of the normalization layer, before smoothing the first statistical value to obtain a second statistical value based on the first statistical value and the statistical values of the historical batches of the normalization layer using a moving average strategy, it further includes:
[0137] Based on the first statistical value of the current batch of the normalization layer, determine whether the current batch of the normalization layer meets the smoothing condition;
[0138] When it is determined that the current batch of the normalization layer meets the smoothing condition, it is determined to perform the operation of smoothing the first statistical value according to the first statistical value and the statistical values of the historical batches of the normalization layer by using a moving average strategy to obtain a second statistical value.
[0139] In some embodiments, the current batch of the normalization layer meeting the smoothing condition includes:
[0140] The absolute value of the difference between the arithmetic mean of the first statistical value of the current batch of the normalization layer and the first statistical values of the historical batches and the geometric mean of the first statistical value of the current batch of the normalization layer and the first statistical values of the historical batches is less than or equal to a first threshold;
[0141] Wherein, the first threshold is determined according to the variance of the first statistical values of the historical batches of the normalization layer.
[0142] In some embodiments, the current batch of the normalization layer meeting the smoothing condition includes:
[0143] The absolute value of the first statistical value of the current batch of the normalization layer is less than or equal to a second threshold;
[0144] Wherein, the second threshold is determined according to the standard deviation of the distribution of the first statistical values of the historical batches of the normalization layer.
[0145] An embodiment of the present application provides an electronic device, including a processor and a memory. Among them, the memory stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the task processing method described above.
[0146] Please refer to Figure 5 , which is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. The electronic device may include a processor 501 and a memory 502 storing machine-executable instructions. The processor 501 and the memory 502 may communicate via a system bus 503. And by reading and executing the machine-executable instructions corresponding to the task processing logic in the memory 502, the processor 501 may execute the task processing method described above.
[0147] The memory 502 mentioned herein may be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium may be: RAM (Radom Access Memory, random access memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0148] In some embodiments, a machine-readable storage medium is also provided, such as the memory 502 in Figure 5 . The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the task processing method described above is implemented. For example, the storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0149] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0150] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.
Claims
1. A task processing method, characterized in that, Including: During the training process of the Transformer model, for any normalization layer in the Transformer model, determine the first statistical value of the current batch of this normalization layer. According to the first statistical value and the statistical values of historical batches of this normalization layer, use the moving average strategy to smooth the first statistical value to obtain the second statistical value, and use the second statistical value of this normalization layer for forward or backward propagation; the normalization operator of the Transformer model uses the batch normalization operator, the input of the Transformer model is sequence information, and the sequence information includes text, speech, or image blocks; the Transformer model is used for natural language processing tasks, speech tasks, or vision tasks; During the process of using the trained Transformer model for task processing, for any normalization layer in the Transformer model, fix the statistical value of this normalization layer as the third statistical value for inference calculation; where, during the training process of the Transformer model, the third statistical value is determined according to the second statistical value using the exponential average strategy; Wherein, the determining the first statistical value of the current batch of this normalization layer, according to the first statistical value and the statistical values of historical batches of this normalization layer, using the moving average strategy to smooth the first statistical value to obtain the second statistical value includes: During the forward propagation process, perform an arithmetic mean operation on the input activation values of the current batch of this normalization layer to obtain the first activation value statistical value of the current batch of this normalization layer; According to the first activation value statistical value of the current batch of this normalization layer and the activation value statistical values of historical batches of this normalization layer, use the geometric mean strategy to smooth the first activation value statistical value of the current batch of this normalization layer to obtain the second activation value statistical value; During the backward propagation process, determine the first gradient statistical value of the current batch of this normalization layer according to the second activation value statistical value of the current batch of this normalization layer; According to the first gradient statistical value of the current batch of this normalization layer and the first gradient statistical values of historical batches of this normalization layer, use the arithmetic mean and exponential average strategy to smooth the first gradient statistical value of the current batch of this normalization layer to obtain the second gradient statistical value.
2. The method according to claim 1, wherein For any normalization layer in the Transformer model, after determining the first statistical value of the current batch of this normalization layer and before smoothing the first statistical value according to the first statistical value and the statistical values of historical batches of this normalization layer using the moving average strategy to obtain the second statistical value, it further includes: Determine whether the current batch of this normalization layer meets the smoothing condition according to the first statistical value of the current batch of this normalization layer; In the case of determining that the current batch of this normalization layer meets the smoothing condition, determine to perform the operation of smoothing the first statistical value according to the first statistical value and the statistical values of historical batches of this normalization layer using the moving average strategy to obtain the second statistical value.
3. The method according to claim 2, characterized in that, The current batch of this normalization layer meets the smoothing condition, including: The absolute value of the difference between the arithmetic mean of the first statistical value of the current batch of this normalization layer and the first statistical value of the historical batches and the geometric mean of the first statistical value of the current batch of this normalization layer and the first statistical value of the historical batches is less than or equal to the first threshold; Among them, the first threshold is determined according to the variance of the first statistical value of the historical batches of this normalization layer.
4. The method according to claim 3, characterized in that, The current batch of this normalization layer meets the smoothing condition, including: The absolute value of the first statistical value of the current batch of this normalization layer is less than or equal to the second threshold; Among them, the second threshold is determined according to the standard deviation of the distribution of the first statistical value of the historical batches of this normalization layer.
5. A task processing device, characterized in that, Including: A training unit, which is used to determine the first statistical value of the current batch of any normalization layer in the Transformer model during the training of the Transformer model. According to the first statistical value and the statistical values of the historical batches of this normalization layer, using the moving average strategy, smooth the first statistical value to obtain a second statistical value, and use the second statistical value of this normalization layer for forward or backward propagation; the normalization operator of the Transformer model uses a batch normalization operator, the input of the Transformer model is sequence information, and the sequence information includes text, speech or image blocks; the Transformer model is used for natural language processing tasks, speech tasks or vision tasks; An inference unit, which is used to fix the statistical value of any normalization layer in the Transformer model to a third statistical value for inference calculation during the task processing using the trained Transformer model; among them, the third statistical value is determined according to the second statistical value using the exponential average strategy during the training of the Transformer model; Among them, the training unit determines the first statistical value of the current batch of this normalization layer. According to the first statistical value and the statistical values of the historical batches of this normalization layer, using the moving average strategy, smooth the first statistical value to obtain a second statistical value, including: During the forward propagation process, perform an arithmetic mean operation on the input activation values of the current batch of this normalization layer to obtain the first activation value statistical value of the current batch of this normalization layer; According to the first activation value statistical value of the current batch of this normalization layer and the activation value statistical values of the historical batches of this normalization layer, use the geometric mean strategy to smooth the first activation value statistical value of the current batch of this normalization layer to obtain a second activation value statistical value; During the backward propagation process, determine the first gradient statistical value of the current batch of this normalization layer according to the second activation value statistical value of the current batch of this normalization layer; According to the first gradient statistical value of the current batch of this normalization layer and the first gradient statistical values of the historical batches of this normalization layer, use the arithmetic mean and exponential average strategy to smooth the first gradient statistical value of the current batch of this normalization layer to obtain a second gradient statistical value.
6. The device according to claim 5, characterized in that, For any normalization layer in the Transformer model, after determining the first statistical value of the current batch of the normalization layer, before smoothing the first statistical value using a moving average strategy based on the first statistical value and the statistical values of historical batches of the normalization layer to obtain a second statistical value, the training unit further includes: Determining whether the current batch of the normalization layer meets the smoothing condition based on the first statistical value of the current batch of the normalization layer; When it is determined that the current batch of the normalization layer meets the smoothing condition, determining to perform the operation of smoothing the first statistical value using a moving average strategy based on the first statistical value and the statistical values of historical batches of the normalization layer to obtain a second statistical value; Among them, the current batch of the normalization layer meeting the smoothing condition includes: The absolute value of the difference between the arithmetic mean of the first statistical value of the current batch of the normalization layer and the first statistical values of historical batches and the geometric mean of the first statistical value of the current batch of the normalization layer and the first statistical values of historical batches is less than or equal to a first threshold; Among them, the first threshold is determined based on the variance of the first statistical values of historical batches of the normalization layer; Or, The current batch of the normalization layer meeting the smoothing condition includes: The absolute value of the first statistical value of the current batch of the normalization layer is less than or equal to a second threshold; Among them, the second threshold is determined based on the standard deviation of the distribution of the first statistical values of historical batches of the normalization layer.
7. An electronic device, characterized in that, It includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor is used to execute the machine-executable instructions to implement the method according to any one of claims 1-4.
8. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are executed by a processor, the method according to any one of claims 1-4 is implemented.
Citation Information
Patent Citations
Batch normalization layers
CN107278310A
Batch renormalization layers
CN110291540A