Batch layer normalization method and device applied to Transformer

By using batch layer normalization method in Transformer, normalizing and affine transformation of the feature map using global mean and standard deviation, the problem of excessive consumption of computing resources and quantization in the existing technology affecting the accuracy of the model is solved, and efficient computing resource usage and accuracy maintenance are achieved.

CN115952834BActive Publication Date: 2025-05-16ZHUHAI OUYEEL SEMICON CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211712332.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2025-05-16
Estimated Expiration
2042-12-29

AI Technical Summary

Technical Problem

When the data dimensions increase and image resolution increase, the number of data statistics increases significantly during the calculation of mean and standard deviation, resulting in excessive consumption of computing resources, and the use of quantitative parameters will affect the model accuracy.

Method used

A batch layer normalization method is proposed. By obtaining the global mean and global standard deviation of the Transformer neural network, the feature map is normalized and affine transformation is performed to obtain the normalized feature map. This method updates the global mean and global variance during the training process, reduces the number of data statistics of normalization operations, and quantifies the same mean and standard deviation of the feature point.

Benefits of technology

The number of statistics on data in normalization operations is reduced, the consumption of computing resources is reduced, and the accuracy of the model is maintained during the quantization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115952834B_ABST
    Figure CN115952834B_ABST
Patent Text Reader

Abstract

The present application discloses a batch layer normalization method and device applied to Transformer, the method comprising obtaining a feature map to be batch normalized, reading the global mean and global standard deviation of the Transformer neural network; normalizing the feature map based on the global mean and global standard deviation; and performing an affine transformation on the normalized feature map to obtain a normalized feature map. By obtaining the global mean and global standard deviation, the present application can reduce the number of statistical operations on data during the normalization operation and can reduce the computing resources required for normalization. At the same time, the present application has a unified mean and standard deviation for each feature point, so that when quantization is performed through the quantization parameters configured by the Transformer, the accuracy of the Transformer will not be affected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a batch layer normalization method and device applied to Transformer. Background Art

[0002] Transformer originated from the field of NLP. The most widely used normalization method for transformers is layer normalization (LayerNorm, hereinafter referred to as LN), which performs online normalization processing. However, with the increase in the data dimension used by Transformer (for example, from one-dimensional data to two-dimensional data) and the increase in image resolution, the mean and standard deviation of LN are both variable, resulting in a significant increase in the number of data statistics in the process of calculating the mean and standard deviation, which requires a large amount of computing resources.

[0003] Therefore the prior art still needs to be improved and enhanced. Summary of the invention

[0004] The technical problem to be solved by the present application is to provide a batch layer normalization method and device applied to Transformer in view of the deficiencies in the prior art.

[0005] In order to solve the above technical problems, the first aspect of the embodiment of the present application provides a batch layer normalization method applied to Transformer, wherein the method includes:

[0006] Obtain the feature map to be batch normalized, and read the global mean and global standard deviation of the Transformer neural network;

[0007] Normalizing the feature map based on the global mean and the global standard deviation;

[0008] Perform affine transformation on the normalized feature map to obtain a normalized feature map.

[0009] The batch layer normalization method applied to Transformer is characterized in that the global standard deviation of the Transformer neural network is determined based on the global variance of the Transformer neural network, and the global mean and the global variance are obtained based on the training sample set during the training process of the Transformer neural network.

[0010] The batch layer normalization method applied to Transformer, wherein the process of obtaining the global mean and global variance of the Transformer neural network specifically includes:

[0011] For each training batch in the training sample set, a batch mean corresponding to the training batch is calculated, wherein the batch mean is determined based on the feature values ​​of each dimension in the feature graph corresponding to the training batch;

[0012] Counting the batch variance corresponding to the training batch, wherein the batch variance is determined based on the feature variance of all pixels in the feature map corresponding to the training batch;

[0013] The global mean and global variance of the Transformer neural network are updated based on the batch mean and the batch variance to obtain the global mean and global variance of the Transformer neural network.

[0014] The batch layer normalization method applied to Transformer, wherein the updating of the global mean and the global standard deviation of the Transformer neural network based on the batch mean and the batch variance specifically includes:

[0015] Obtaining the global mean and global variance of the Transformer neural network;

[0016] Weighting the batch mean to the global mean to update the global mean;

[0017] The batch variance is weighted to the global variance to update the global variance.

[0018] The batch layer normalization method applied to Transformer, wherein the affine transformation is performed on the normalized feature map to obtain the normalized feature map is specifically:

[0019] An affine transformation is performed on the channel dimension of the normalized feature map to obtain a normalized feature map.

[0020] A second aspect of the embodiment of the present application provides a processing method using a Transformer neural network, and the processing method specifically includes:

[0021] Get the data to be processed;

[0022] The data to be processed is input into a trained Transformer neural network, and processed data corresponding to the data to be processed is determined by the Transformer neural network, wherein the Transformer neural network is configured with a batch layer normalization processing process, wherein the batch layer normalization layer processing process adopts the batch layer normalization method as described above.

[0023] The processing method using the Transformer neural network, wherein the Transformer neural network includes a plurality of stacked encoding modules, the encoding module includes a linear layer and a multi-head attention module, and the output items of the linear layer are input into the multi-head attention module after batch normalization processing.

[0024] The processing method using the Transformer neural network, wherein the batch layer normalization processing process is performed through a batch layer normalization layer set by the Transformer neural network, or the batch layer normalization processing process is embedded in a linear layer in the Transformer neural network and performed through the linear layer.

[0025] A third aspect of an embodiment of the present application provides a batch layer normalization processing device applied to a Transformer, the device comprising:

[0026] A reading unit, used to obtain the feature map to be batch normalized, and read the global mean and global standard deviation of the Transformer neural network;

[0027] A normalization unit, configured to normalize the feature map based on the global mean and the global standard deviation;

[0028] The affine unit is used to perform affine transformation on the normalized feature map to obtain a normalized feature map.

[0029] A fourth aspect of the embodiments of the present application provides a terminal device, comprising: a processor, a memory, and a communication bus; the memory stores a computer-readable program that can be executed by the processor;

[0030] The communication bus realizes the connection and communication between the processor and the memory;

[0031] When the processor executes the computer-readable program, the processor implements the steps in any of the above-described batch layer normalization methods applied to Transformer, and / or implements the steps in the above-described processing method applied to Transformer neural network.

[0032] Beneficial effects: Compared with the prior art, the present application provides a batch layer normalization method and device applied to Transformer, the method comprising obtaining a feature map to be batch normalized, reading the global mean and global standard deviation of the Transformer neural network; normalizing the feature map based on the global mean and global standard deviation; performing an affine transformation on the normalized feature map to obtain a normalized feature map. By obtaining the global mean and global standard deviation, the present application can reduce the number of statistics on data during the normalization operation and reduce the computing resources required for normalization. At the same time, the mean and standard deviation of each feature point in the present application are unified, so that the accuracy of the Transformer will not be affected when quantized through the quantization parameters configured by the Transformer. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without inventive work.

[0034] Figure 1 A flowchart of the batch layer normalization method applied to Transformer provided in this application.

[0035] Figure 2 A comparison diagram of the vector directions of the normalized feature vector and the original vector in the batch layer normalization method applied to Transformer provided in this application.

[0036] Figure 3 Flowchart of a processing method using a Transformer neural network.

[0037] Figure 4 This is a schematic diagram of the structure of the Transformer neural network using LN.

[0038] Figure 5 This is a schematic diagram of the structure of the Transformer neural network that uses LBN to replace LN.

[0039] Figure 6 The schematic diagram of the Transformer neural network with LBN replacing LN and adding LBN between the linear layer and the multi-head attention module.

[0040] Figure 7 for Figure 4 , Figure 5 and Figure 6Comparison chart of top1 accuracy curves.

[0041] Figure 8 for Figure 6 Schematic diagram of the structure after the LBN in is embedded in the adjacent linear layer.

[0042] Fig. 9 This is a schematic diagram of the structural principle of the batch layer normalization device applied to Transformer provided in this application.

[0043] Fig.10 This is a schematic diagram of the structure of the terminal device provided in this application. DETAILED DESCRIPTION

[0044] The present application provides a batch layer normalization method and device for Transformer. In order to make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described here are only used to explain the present application and are not used to limit the present application.

[0045] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0046] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.

[0047] It should be understood that the sequence numbers and sizes of the steps in this embodiment do not mean the order of execution. The execution order of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.

[0048] After research, it was found that Transformer originated from the field of NLP, and the normalization method widely used in Transformer is layer normalization (LayerNorm, hereinafter referred to as LN), which performs online normalization processing through LN. However, with the increase of data dimensions used by Transformer (for example, from one-dimensional data to two-dimensional data) and the increase of image resolution, since the mean and standard deviation of LN are both variable, the number of data statistics in the process of calculating the mean and standard deviation increases significantly, which requires a large amount of computing resources. In addition, when Transformer needs to be quantized after being deployed to the terminal, since LN uses a set of quantization parameters, and each feature point in the statistical normalization process of LN is independent of each other, the mean and standard deviation are very discrete, which will cause the accuracy of the model to be seriously affected after quantization using a set of quantization parameters.

[0049] In order to solve the above problems, in an embodiment of the present application, a feature map to be batch normalized is obtained, and the global mean and global standard deviation of the Transformer neural network are read; based on the global mean and global standard deviation, the feature map is normalized; and the normalized feature map is affine transformed to obtain a normalized feature map. By obtaining the global mean and global standard deviation, the present application can reduce the number of statistics on data during the normalization operation and reduce the computing resources required for normalization. At the same time, the mean and standard deviation of each feature point in the present application are unified, so that when quantization is performed through the quantization parameters configured by the Transformer, the accuracy of the Transformer will not be affected.

[0050] The application content is further explained below through the description of embodiments in conjunction with the accompanying drawings.

[0051] This embodiment provides a batch layer normalization method applied to Transformer, such as Figure 1 As shown, the method includes:

[0052] S10, obtaining feature maps to be batch normalized, and reading the global mean and global standard deviation of the Transformer neural network.

[0053] Specifically, the feature map can be the output item of a network layer in the Transformer neural network, and the feature map needs to be batch normalized. For example, the feature map is the output item of the multi-head attention module in the Transformer neural network. The global mean and the global standard deviation are both used to perform batch normalization on the feature map, and each feature point in the feature map is batch normalized using the global mean and the global standard deviation. Among them, the global mean and the global standard deviation can be obtained by statistically analyzing all feature points in the feature map, or can be determined during the training process of the Transformer neural network.

[0054] In one implementation, the global mean and global standard deviation are determined during the training process of the Transformer neural network, wherein the global standard deviation of the Transformer neural network is determined based on the global variance of the Transformer neural network. Based on this, the process of obtaining the global mean and global standard deviation of the Transformer neural network specifically includes:

[0055] For each training batch in the training sample set, a batch mean corresponding to the training batch is calculated, wherein the batch mean is determined based on the feature values ​​of each dimension in the feature graph corresponding to the training batch;

[0056] Counting the batch variance corresponding to the training batch, wherein the batch variance is determined based on the feature variance of all pixels in the feature map corresponding to the training batch;

[0057] The global mean and global variance of the Transformer neural network are updated based on the batch mean and the batch variance to obtain the global mean and global variance of the Transformer neural network.

[0058] Specifically, the feature map corresponding to the training batch includes four dimensions, namely, the number of batches, the number of feature map channels, the height of the feature map, and the width of the feature map. Therefore, the feature map corresponding to the training batch can be recorded as X, and the feature dimension of X is B×C×H×W, where B is the number of batches, C is the number of feature channels, H is the height of the feature map, and W is the width of the feature map.

[0059] The batch mean corresponding to the training batch is counted as the cumulative feature Figure X The features of each dimension are divided by the total number of features, where each dimension includes B dimension, C dimension, H dimension and W dimension. Therefore, the calculation formula of batch mean can be:

[0060]

[0061] Among them, N = B × C × H × W is the total number of features, μB is the batch mean, X ijkl As a feature.

[0062] Similarly, the batch variance corresponding to the training batch is counted as the cumulative feature Figure X The feature variance of each dimension is divided by the total number of features, where each dimension includes B dimension, C dimension, H dimension and W dimension. Therefore, the calculation formula for batch variance can be:

[0063]

[0064] Among them, μ represents the global mean of the Transformer neural network, σ 2 B Represents the batch variance corresponding to the training batch.

[0065] Furthermore, after obtaining the batch variance, in order to avoid the batch variance being zero and affecting the normalization of the feature graphs corresponding to the training batch based on the batch variance, when determining the candidate global standard deviation based on the batch variance, a preset adjustment parameter is accumulated on the batch variance, and then the candidate global standard deviation is calculated based on the batch variance with the preset adjustment parameter accumulated, and finally the feature graphs are normalized based on the candidate global standard deviation. Figure X Perform normalization.

[0066] Among them, based on the candidate global standard deviation feature Figure X The normalized processing can be expressed as:

[0067]

[0068] in, is the normalized feature, μ B is the batch mean, is the candidate global standard deviation, ∈ is the preset adjustment parameter, σ 2 B is the batch variance, X ijkl Features Figure X Features in .

[0069] Further, after obtaining the batch mean and the batch variance, the batch mean and the batch variance are used to update the global mean and the global variance of the Transformer neural network, wherein the updating process may include:

[0070] Obtaining the global mean and global variance of the Transformer neural network;

[0071] Weighting the batch mean to the global mean to update the global mean;

[0072] The batch variance is weighted to the global variance to update the global variance.

[0073] Specifically, the Transformer neural network is provided with a global mean and a global variance, wherein, when the training batch is the first training batch for training the Transformer neural network, the batch mean can be directly used as the global mean, and the batch variance can be used as the global variance, or the Transformer neural network is provided with an initial global mean and an initial global variance, and then the global mean is determined based on the batch mean and the initial global mean, and the global variance is determined based on the batch variance and the initial global variance, and the determined global mean and global variance are used as the global mean and global variance of the Transformer neural network. Conversely, when the training batch is not the first training batch for training the Transformer neural network, the global mean is directly updated based on the global mean and the batch mean in the Transformer neural network, and the global variance is updated based on the global variance and the batch variance in the Transformer neural network.

[0074] In one implementation, the update process of the global mean and the update process of the global variance are both obtained by weighting. In addition, since the global standard deviation is obtained by taking the square root of the global variance, the Transformer neural network can directly store the global variance and update the global variance through the batch variance. Therefore, the update formula of the global mean and the update formula of the global variance can be expressed as:

[0075] μ=α×μ+(1-α)×μ B

[0076] σ 2 =α×σ 2 +(1-α)×σ 2 B

[0077] Among them, μ is the global mean, α is the preset weight coefficient, μ B is the batch mean, σ 2 is the global variance, σ 2 B is the batch variance.

[0078] S20. Normalize the feature map based on the global mean and the global standard deviation.

[0079] Specifically, after obtaining the global mean and the global standard deviation, each feature in the feature map is normalized based on the global mean and the global standard deviation, where the normalization process can be expressed as:

[0080]

[0081] in, To determine the normalized features based on the global mean and global standard deviation, X ijkl is the feature in the feature map, μ is the global mean, and σ is the global standard deviation.

[0082] S30, performing affine transformation on the normalized feature map to obtain a normalized feature map.

[0083] Specifically, the affine transformation is used to perform an affine transformation on the feature map normalized based on the global mean and the global standard deviation, wherein the affine transformation is performed on the channel dimension of the normalized feature map, and the affine transformation is performed on the channel dimension of the normalized feature map to obtain a normalized feature map. In a specific implementation, the affine transformation can be expressed as:

[0084]

[0085] Among them, γ ijkl is the feature in the normalized feature map, γ j and β j are affine parameters.

[0086] Further, if Figure 2 As shown, the vector direction of the feature vector in the normalized feature map is the same as the vector direction of the original feature vector, that is, the feature vector C processed by the batch layer normalization method (LayerBatchNorm, LBN) provided in this embodiment LBN The vector direction of the original feature vector is the same as that of the original feature vector. At the same time, the feature vector C after layer normalization (LayerNorm, LN) LN The vector direction of is the same as the vector direction of the original feature vector, so the LBN provided in this embodiment can replace the LN in the Transformer neural network. The feature vector C after the existing offline batch normalization (BatchNorm, BN) processing BN The vector direction is the same as the original eigenvector C src The vector direction is different, so BN cannot be applied to Transformer neural network.

[0087] In summary, the present embodiment provides a batch layer normalization method and device applied to Transformer, the method comprising obtaining a feature map to be batch normalized, reading the global mean and global standard deviation of the Transformer neural network; normalizing the feature map based on the global mean and global standard deviation; and performing an affine transformation on the normalized feature map to obtain a normalized feature map. When the batch layer normalization method applied to Transformer provided in the present embodiment is normalized, all features share normalization parameters (i.e., global mean and global variance), which does not change the vector direction of the feature vector, but simply translates and scales. Therefore, the batch layer normalization method provided in the present embodiment can be applied to the Transformer neural network, replacing the LN layer in the Transformer neural network, reducing the consumption of batch normalization processing in the Transformer neural network, and thus reducing the consumption of the Transformer neural network. At the same time, since the batch layer normalization method provided in the present embodiment has a unified mean and standard deviation of each feature point, the accuracy of the Transformer will not be affected when quantized by the quantization parameters configured by the Transformer.

[0088] Based on the above batch layer normalization method applied to Transformer, this embodiment provides a processing method using Transformer neural network, such as Figure 3 As shown, the processing method specifically includes:

[0089] N10. Obtain data to be processed;

[0090] N20. Input the data to be processed into a trained Transformer neural network, and determine the processed data corresponding to the data to be processed through the Transformer neural network.

[0091] Specifically, the data to be processed is an input item of the Transformer neural network, and the data to be processed is processed by the Transformer neural network. For example, the data to be processed is a processed image, and the processed data is the target object area in the image to be processed. The Transformer neural network is configured with a batch layer normalization processing process, and the batch layer normalization layer processing process adopts the batch layer normalization method as described above. It can be understood that the Transformer neural network configuration performs normalization operations on the feature maps that need to be normalized in the Transformer neural network through the batch layer normalization method described above.

[0092] The Transformer neural network includes several stacked encoding modules, and the encoding module includes a linear layer and a multi-head attention module. The output items of the linear layer are input into the multi-head attention module after batch normalization processing, wherein the batch normalization processing can be performed through the batch normalization layer, or can be embedded in the linear layer, and the process performed in the linear layer is performed until the batch normalization processing is performed, so that the batch layer normalization processing does not increase the computing power consumption. In other words, the batch layer normalization processing process is performed through the batch layer normalization layer set by the Transformer neural network, or the batch layer normalization processing process is embedded in the linear layer in the Transformer neural network and performed through the linear layer.

[0093] like Figure 4 The Transformer neural network using LN shown in Figure 5 The Transformer neural network using LBN shown in Figure 6 The Transformer neural network using LBN and QKV implanted in LBN is trained on the image1 K data set for 170 epochs to obtain each Transformer neural network, as shown in Figure 7 As shown in the figure, when training for more than 15 epochs, the model accuracy of the Transformer neural network using LBN and implanting QKV into LBN is significantly higher than that of the other two Transformer neural networks.

[0094] In one implementation, if Figure 6 As shown, the encoder of the Transformer neural network includes a batch normalization layer, a linear layer, a batch normalization layer, a multi-head attention module, a batch normalization layer and a multi-layer perception module connected in sequence, wherein the batch normalization layers in the encoder of the Transformer neural network all adopt the batch layer normalization method applied to the Transformer neural network described in the above embodiment, thereby improving the model performance of the Transformer neural network.

[0095] In one implementation, since batch normalization is a linear operation, Figure 6 The batch normalization performed by the batch normalization layer in the encoder of the Transformer neural network shown is embedded in the linear layer adjacent to the batch normalization layer, resulting in Figure 8 The encoder of the Transformer neural network shown in . Figure 8 The encoder of the Transformer neural network shown in Figure 1 performs batch normalization when running the linear layer, so that batch layer normalization does not add additional computational overhead.

[0096] Based on the above batch layer normalization method applied to Transformer, this embodiment provides a batch layer normalization processing device applied to Transformer neural network, such as Fig. 9 As shown, the device comprises:

[0097] A reading unit 100 is used to obtain a feature map to be batch normalized and read a global mean and a global standard deviation of the Transformer neural network;

[0098] A normalization unit 200, configured to normalize the feature map based on the global mean and the global standard deviation;

[0099] The affine unit 300 is used to perform affine transformation on the normalized feature map to obtain a normalized feature map.

[0100] Based on the above-mentioned batch layer normalization method applied to Transformer, this embodiment provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps in the batch layer normalization method applied to Transformer as described in the above-mentioned embodiment.

[0101] Based on the above batch layer normalization method applied to Transformer, the present application also provides a terminal device, such as Fig.10 As shown, it includes at least one processor 20, a display screen 21, and a memory 22, and may also include a communication interface 23 and a bus 24. The processor 20, the display screen 21, the memory 22, and the communication interface 23 may communicate with each other through the bus 24. The display screen 21 is configured to display a preset user guide interface in the initial setting mode. The communication interface 23 may transmit information. The processor 20 may call the logic instructions in the memory 22 to execute the method in the above embodiment.

[0102] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0103] The memory 22 is a computer-readable storage medium that can be configured to store software programs, computer executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions or modules stored in the memory 22, that is, implementing the methods in the above embodiments.

[0104] The memory 22 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and at least one application required for a function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, a variety of media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, may also be a transient storage medium.

[0105] In addition, the specific process of loading and executing the multiple instruction processors in the above-mentioned storage medium and the terminal device has been described in detail in the above-mentioned method, and will not be described one by one here.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A processing method using a Transformer neural network, characterized in that: The processing method specifically includes: Acquire data to be processed, where the data to be processed is an image to be processed; Inputting the data to be processed into a trained Transformer neural network, and determining processed data corresponding to the data to be processed through the Transformer neural network, wherein the processed data is a target object area in the image to be processed; The Transformer neural network is configured with a batch layer normalization process, wherein the batch layer normalization process is embedded in the linear layer, and the batch normalization process is executed during the execution of the linear layer. The batch layer normalization process is: Obtain the feature map to be batch normalized, and read the global mean and global standard deviation of the Transformer neural network; Normalizing the feature map based on the global mean and the global standard deviation; An affine transformation is performed on the normalized feature map to obtain a normalized feature map, so that batch layer normalization processing does not increase the computing power consumption.

2. The processing method using a Transformer neural network according to claim 1, characterized in that: The global standard deviation of the Transformer neural network is determined based on the global variance of the Transformer neural network, and the global mean and the global variance are obtained based on the training sample set during the training process of the Transformer neural network.

3. The processing method using a Transformer neural network according to claim 2, characterized in that: The process of obtaining the global mean and global variance of the Transformer neural network specifically includes: For each training batch in the training sample set, a batch mean corresponding to the training batch is calculated, wherein the batch mean is determined based on the feature values ​​of each dimension in the feature graph corresponding to the training batch; Counting the batch variance corresponding to the training batch, wherein the batch variance is determined based on the feature variance of all pixels in the feature map corresponding to the training batch; The global mean and global variance of the Transformer neural network are updated based on the batch mean and the batch variance to obtain the global mean and global variance of the Transformer neural network.

4. The processing method using a Transformer neural network according to claim 3, characterized in that: The updating of the global mean and global standard deviation of the Transformer neural network based on the batch mean and batch variance specifically includes: Obtaining the global mean and global variance of the Transformer neural network; Weighting the batch mean to the global mean to update the global mean; The batch variance is weighted to the global variance to update the global variance.

5. According to the processing method using a Transformer neural network according to claim 1, the step of performing an affine transformation on the normalized feature map to obtain a normalized feature map is specifically: An affine transformation is performed on the channel dimension of the normalized feature map to obtain a normalized feature map.

6. The processing method using a Transformer neural network according to claim 1, characterized in that: The Transformer neural network includes several stacked encoding modules, each of which includes a linear layer and a multi-head attention module. The output items of the linear layer are input into the multi-head attention module after batch normalization processing.

7. A terminal device, characterized in that: include: Processor, memory and communication bus; The memory stores a computer-readable program executable by the processor; The communication bus realizes the connection and communication between the processor and the memory; When the processor executes the computer-readable program, the steps in the processing method using the Transformer neural network as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Normalization method and device for deep neural network, equipment and storage medium

    CN108921283A

  • Batch renormalization layers

    CN110291540A