A data processing method and related device
By inserting the target module into the transformer model for feature map fusion, the problem of degradation of accuracy after compression processing is solved, and the effect of improving data processing accuracy is achieved while reducing parameters and computing power overhead.
Patent Information
- Application Number
- CN202110611218.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-01
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-06-01
AI Technical Summary
After compression processing, the data processing accuracy of the transformer model has been greatly reduced, making it difficult to maintain high accuracy while reducing the model parameter amount and computing power overhead.
Insert the target module into the transformer model, and perform convolution-based nonlinear operation on the feature map output of the target network layer through the target module to generate more feature maps and fuse them with the original feature map to enhance the feature expression ability of the model.
While reducing the model parameter quantity and computing power overhead, the data processing accuracy of the transformer model is improved and the performance of the model in natural language processing tasks is enhanced.
Smart Images

Figure CN113505193B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a data processing method and related devices. Background Art
[0002] Artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0003] With the continuous development of artificial intelligence technology, natural language human-computer interaction systems that enable human-computer interaction through natural language have become increasingly important. For human-computer interaction to be possible through natural language, the system needs to be able to recognize the specific meaning of human natural language. Usually, the system identifies the specific meaning of a sentence by extracting key information from the natural language sentence.
[0004] The transformer structure has powerful semantic expression capabilities and can capture long-term dependencies in text. Since its proposal, it has significantly outperformed previous models in a series of natural language processing tasks represented by translation. Pre-trained language models based on the transformer structure have also achieved very good results in fields such as question-and-answer systems and voice assistants.
[0005] The transformer model has a large number of parameters and high requirements for computing and power consumption. Therefore, the transformer model can usually be compressed by pruning or other methods to obtain a lighter and more quantized transformer model. However, the compression process will significantly reduce the data processing accuracy of the transformer model. Summary of the Invention
[0006] In a first aspect, this application provides a data processing method, the method comprising:
[0007] Obtain a transformer model, the transformer model including a target network layer and a target module;
[0008] Among them, the terminal device or the cloud - side server can obtain a Transformer model for model inference. The Transformer model can be a trained Transformer model. For example, the Transformer model can be a pre - trained model or a model after model fine - tuning. The Transformer model can include a target network layer. Among them, the target network layer can be the attention layer or the feed - forward layer in the Transformer layer.
[0009] Among them, a target module can be inserted into the Transformer model to obtain the Transformer model in the embodiments of the present application.
[0010] Obtain the data to be processed, and process the data to be processed through the Transformer model to obtain a data processing result. Among them, the target module is used to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output to obtain the updated feature map output. The target operation is a non - linear operation based on convolution.
[0011] Among them, the role of the target module in the embodiments of the present application is similar to that of the Ghost module. Generally speaking, most linear operations can be used as the operations adopted in the Ghost module. However, in the Transformer model, simple linear operations do not help much in improving the performance of the model. Therefore, in the embodiments of the present application, non - linear operations are introduced on the basis of convolution operations.
[0012] Among them, the above - mentioned feature map output can be understood as the feature map output by the network layer (which can be the final output of the network layer or the intermediate layer output of the network layer). For example, the feature map output of the target network layer can be understood as the feature map output by the target network layer.
[0013] Among them, the data to be processed can be text data.
[0014] The data to be processed can be processed through the Transformer model. The data to be processed can be the input data in the model inference process, and the Transformer model is the model adopted in the model inference process.
[0015] In the embodiments of the present application, a target module is inserted into the transformer model. The target module generates more feature maps (i.e., the operation results obtained by the target module through convolution-based non-linear operations), and fuses the operation results with the input of the target module, increasing the information carried in the feature maps output by the target network layer in the transformer model. Moreover, since the number of parameters of the target module itself and the computing power overhead required during operation are very small, it is equivalent to improving the data processing accuracy of the model on the premise of reducing the number of model parameters and computing power overhead.
[0016] In a possible implementation, the weight parameters included in the convolution kernel used for the convolution are obtained through regularization processing. By performing regularization processing on the weight parameters included in the convolution kernel, the input and output values of the convolution operation can be made as close as possible. This can make the model more stable during the training process, make the model performance robust, and can reduce the waste of computing power resources caused by the redundant parameter adjustment process. Among them, the regularization processing method can include, but is not limited to, Softmax regularization, L1 regularization, etc.
[0017] In a possible implementation, the convolution kernel used for the convolution satisfies at least one of the following conditions: the difference between the sum of the weight parameters included in the convolution kernel and 1 is within a preset range; the weight parameters included in the convolution kernel are positive numbers. In order to ensure that the numerical sizes of the input and output of the target module are close, the weight parameters included in the convolution kernel used for the convolution operation can be regularized so that all the weight parameters included in the convolution kernel are positive, and the sum of the weight parameters is 1 or a value close to 1, for example, the difference from 1 can be within 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.1.
[0018] In a possible implementation, the length and width dimensions of the feature map output are the same as those of the updated feature map output. In order to enable the operation result (the updated feature map output) obtained through the target operation to be fused (such as added or concatenated) with the feature map output of the target network layer, it is necessary to make the length and width dimensions of the updated feature map output the same as those of the feature map output of the target network layer.
[0019] In a possible implementation, the non-linear operation is used to perform non-linear processing on the result obtained by the convolution. The non-linear operation can include, but is not limited to, ReLU.
[0020] In a possible implementation, the target network layer includes an attention layer.
[0021] In a possible implementation, the attention layer includes M attention heads, and the feature map output of the target network layer includes the M feature map outputs of the M attention heads;
[0022] Performing a target operation on the feature map output of the target network layer to obtain an operation result, and fusing the operation result with the feature map output, includes:
[0023] Performing N target operations on the M feature map outputs to obtain N first feature maps, and fusing the N first feature maps with the M feature map outputs of the M attention heads.
[0024] In a possible implementation, the fusing of the N first feature maps with the M feature map outputs of the M attention heads includes: performing an addition operation on the N first feature maps and the M feature map outputs of the M attention heads. Since the target module can generate more feature maps through inexpensive operations, after performing an addition operation on the N first feature maps and the M feature map outputs of the M attention heads, the information carried in the feature maps output by the attention layer can be enriched, thereby improving the data processing accuracy of the model on the premise of paying less parameter quantity and computing power cost.
[0025] In a possible implementation, the attention layer includes M attention heads, each of the M attention heads includes a first branch and a second branch, the output of the first branch is obtained by a dot product operation of the K vector and the Q vector, the output of the second branch is obtained according to the V vector, and the feature map output of the target network layer includes the outputs of the M first branches of the M attention heads; performing a target operation on the feature map output of the target network layer to obtain an operation result, and fusing the operation result with the feature map output, includes: performing N target operations on the outputs of the M first branches to obtain N second feature maps, and fusing the N second feature maps with the outputs of the M first branches.
[0026] In a possible implementation, the fusion of the N second feature maps and the outputs of the M first branches includes: performing a concatenation operation (concat) on the N second feature maps and the outputs of the M first branches. In the embodiments of the present application, the transformer model can be obtained by pruning the attention heads. For example, the K matrix and the Q matrix can be cropped to a size of A*M, while the V matrix is not cropped and retains a size of A*(M+N). And through the target module, a new matrix with a size of A*N can be generated and dot-multiplied with the N v vectors in the V matrix. Equivalent to looking at the outputs of the first branch and the second branch, the dimensions of the outputs are the same as those without cropping, and there is not much loss in the amount of data. And because the number of parameters of the target module itself and the computing power overhead required during operation are very small, it is equivalent to improving the data processing accuracy of the model on the premise of reducing the number of model parameters and the computing power overhead.
[0027] In a possible implementation, the attention layer includes M attention heads, and each of the M attention heads includes a third branch. The output of the third branch is obtained by performing a dot product operation on the K vector, the Q vector, and the V vector. The feature map output of the target network layer includes the outputs of the M third branches of the M attention heads; the target operation on the feature map output of the target network layer to obtain an operation result and fuse the operation result with the feature map output includes: performing N target operations on the outputs of the M third branches to obtain N third feature maps, and fusing the N third feature maps with the outputs of the M third branches. For example, the N third feature maps can be concatenated with the outputs of the M third branches.
[0028] In the embodiments of the present application, the transformer model can be obtained by pruning the attention heads. For example, the K matrix, the Q matrix, and the V matrix can be cropped from a size of dimension M+N to a size of dimension M. Furthermore, the dimension of the output of the third branch is also M. And through the target module, a new matrix with a dimension of N can be generated. After concatenating with the output of the third branch, a feature map with a dimension of M+N can be obtained. Equivalent to looking at the output of the third branch, the dimension is the same as that without cropping, and there is not much loss in the amount of data. And because the number of parameters of the target module itself and the computing power overhead required during operation are very small, it is equivalent to improving the data processing accuracy of the model on the premise of reducing the number of model parameters and the computing power overhead.
[0029] In a possible implementation, the target network layer includes a feed-forward layer FFN.
[0030] In a possible implementation, the FFN includes an intermediate layer, the intermediate layer includes X groups of neurons, and the feature map output of the target network layer includes X feature map outputs of the X groups of neurons; performing a target operation on the feature map output of the target network layer to obtain an operation result, and fusing the operation result with the feature map output includes: performing the target operation N times on the X feature map outputs to obtain N fourth feature maps, and fusing the N fourth feature maps with the feature map outputs of the X groups of neurons. For example, the N fourth feature maps can be concatenated with the X feature map outputs of the X groups of neurons.
[0031] In the embodiments of the present application, the transformer model can be obtained by pruning the intermediate layer of the FFN. For example, the output feature map of the neuron can be cropped from a size of dimension M + N to a size of dimension M, and a new matrix of dimension N can be generated by the target module. After concatenating with the X feature map outputs, a feature map of dimension M + N can be obtained. Equivalent to looking at the output of the intermediate layer, the dimension is the same as the output without cropping, and there is not much loss in the amount of data. And because the number of parameters of the target module itself and the computing power overhead required during operation are very small, it is equivalent to improving the data processing accuracy of the model on the premise of reducing the number of model parameters and computing power overhead.
[0032] In a possible implementation, the FFN includes an intermediate layer and an output layer, the intermediate layer includes X groups of neurons, the output layer is used to process the X feature map outputs of the X groups of neurons to obtain X output layer outputs, and the feature map output of the target network layer includes the X output layer outputs; performing a target operation on the feature map output of the target network layer to obtain an operation result, and updating the feature map output according to the operation result includes: performing the target operation N times on the X output layer outputs to obtain N fifth feature maps, and updating the X output layer outputs according to the N fifth feature maps. For example, the N fifth feature maps can be concatenated with the X output layer outputs.
[0033] In a possible implementation, after processing the data to be processed by the transformer model, the method further includes: training the transformer model according to the data processing result to obtain a trained transformer model.
[0034] In a possible implementation, before obtaining the transformer model, the method further includes:
[0035] Obtain performance requirements, where the performance requirements are used to indicate the data processing accuracy of the Transformer model and / or the model size of the Transformer model; according to the performance requirements, determine the number of target modules and the insertion positions in the Transformer model.
[0036] Specifically, the terminal device can send performance requirements to the cloud server. Among them, the performance requirements include, but are not limited to, at least one of accuracy requirements, latency requirements, or model compression ratio requirements. Furthermore, the cloud server can obtain the performance requirements.
[0037] In an embodiment of the present application, the cloud server may have an initial neural network model with a Transformer structure. After receiving the performance requirements sent by the terminal device, the cloud server can determine the pruning size of the Transformer model based on the received performance requirements. Specifically, when the accuracy requirement included in the performance requirements is relatively high, it can be determined that the pruning size of the Transformer model is relatively large; when the latency requirement included in the performance requirements is relatively high, it can be determined that the pruning size of the Transformer model is relatively small; when the model compression ratio requirement included in the performance requirements is relatively high, it can be determined that the pruning size of the Transformer model is relatively large. Specifically, the cloud server can determine the pruning size information of the Transformer model based on a preset functional relationship or based on a preset corresponding relationship (for example, by looking up a table).
[0038] In a possible implementation, the higher the data processing accuracy, the more the number of target modules; and / or, the higher the data processing accuracy, the closer the insertion position of the target module in the Transformer model is to the embedding layer in the Transformer model; and / or, the larger the model size, the more the number of target modules.
[0039] In an embodiment of the present application, the cloud server may have an initial neural network model with a Transformer structure. After receiving the performance requirements sent by the terminal device, the cloud server can determine the number of target modules and the insertion positions in the Transformer model based on the received performance requirements.
[0040] In a possible implementation, the higher the data processing accuracy, the more the number of target modules; and / or, the higher the data processing accuracy, the closer the insertion position of the target module in the Transformer model is to the embedding layer in the Transformer model.
[0041] For example, when the performance requirements include relatively high precision requirements, the number of target modules can be determined to be larger, or the insertion position of the target module in the Transformer model is closer to the embedding layer in the Transformer model. When the performance requirements include relatively high latency requirements, the number of target modules can be determined to be smaller.
[0042] In a possible implementation, the size of the Transformer model can be determined first, and then the number of target modules and their insertion positions in the Transformer model can be further determined according to performance parameters such as the remaining allocable number of parameters and FLOPs.
[0043] In a possible implementation, the Transformer model is a model after compression processing, where the compression processing can be pruning processing or quantization processing, etc.
[0044] It should be understood that the pruning operation on the Transformer model is optional, and the target module can also be directly applied to the Transformer model (for example, the Transformer model is a pre-trained model or a model obtained after fine-tuning) to obtain better model performance.
[0045] In a possible implementation, the processing of the data to be processed by the Transformer model includes: performing processing corresponding to a target task on the data to be processed through the Transformer model, where the target task includes: reading comprehension, text translation, paraphrase recognition, named entity recognition, text sentiment analysis, natural language inference, text automatic question answering, text intention recognition, text classification, text simplification, or text story generation.
[0046] In a second aspect, the present application provides a data processing device, where the device includes:
[0047] An acquisition module, configured to acquire a Transformer model, where the Transformer model includes a target network layer and a target module;
[0048] A data processing module, configured to acquire data to be processed, and process the data to be processed through the Transformer model to obtain a data processing result; where the target module is configured to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output to obtain an updated feature map output; the target operation is a non-linear operation based on convolution.
[0049] In a possible implementation, the weight parameters included in the convolution kernel used for the convolution are obtained through regularization processing.
[0050] In a possible implementation, the convolution kernel used for the convolution satisfies at least one of the following conditions:
[0051] The difference between the sum of the weight parameters included in the convolution kernel and 1 is within a preset range;
[0052] The weight parameters included in the convolution kernel are positive numbers.
[0053] In a possible implementation, the length and width dimensions of the output of the feature map are the same as those of the updated output of the feature map.
[0054] In a possible implementation, the non-linear operation is used to perform non-linear processing on the result obtained by the convolution.
[0055] In a possible implementation, the target network layer includes an attention layer.
[0056] In a possible implementation, the attention layer includes M attention heads, and the output of the feature map of the target network layer includes the M outputs of the feature maps of the M attention heads;
[0057] The data processing module is configured to perform N target operations on the M outputs of the feature maps to obtain N first feature maps, and fuse the N first feature maps with the M outputs of the feature maps of the M attention heads.
[0058] In a possible implementation, the data processing module is configured to perform an addition operation on the N first feature maps and the M outputs of the feature maps of the M attention heads.
[0059] In a possible implementation, the attention layer includes M attention heads, each of the M attention heads includes a first branch and a second branch, the output of the first branch is obtained according to the dot product operation of the K vector and the Q vector, the output of the second branch is obtained according to the V vector, and the output of the feature map of the target network layer includes the M outputs of the first branches of the M attention heads;
[0060] The data processing module is configured to perform N target operations on the M outputs of the first branches to obtain N second feature maps, and fuse the N second feature maps with the M outputs of the first branches.
[0061] In a possible implementation, the data processing module is configured to perform a concatenation operation (concat) on the N second feature maps and the M outputs of the first branches.
[0062] In a possible implementation, the attention layer includes M attention heads, each of the M attention heads includes a third branch, the output of the third branch is obtained according to the dot product operation of the K vector, the Q vector and the V vector, and the feature map output of the target network layer includes the outputs of the M third branches of the M attention heads;
[0063] The data processing module is configured to perform N target operations on the outputs of the M third branches to obtain N third feature maps, and fuse the N third feature maps with the outputs of the M third branches.
[0064] In a possible implementation, the data processing module is configured to perform a concatenation operation on the N third feature maps and the outputs of the M third branches.
[0065] In a possible implementation, the target network layer includes a feed-forward network (FFN).
[0066] In a possible implementation, the FFN includes an intermediate layer, the intermediate layer includes X groups of neurons, and the feature map output of the target network layer includes the X feature map outputs of the X groups of neurons;
[0067] The data processing module is configured to perform N target operations on the X feature map outputs to obtain N fourth feature maps, and fuse the N fourth feature maps with the feature map outputs of the X groups of neurons.
[0068] In a possible implementation, the data processing module is configured to perform a concatenation operation on the N fourth feature maps and the X feature map outputs of the X groups of neurons.
[0069] In a possible implementation, the FFN includes an intermediate layer and an output layer, the intermediate layer includes X groups of neurons, the output layer is configured to process the X feature map outputs of the X groups of neurons to obtain X output layer outputs, and the feature map output of the target network layer includes the X output layer outputs;
[0070] The data processing module is configured to perform N target operations on the X output layer outputs to obtain N fifth feature maps, and fuse the N fifth feature maps with the X output layer outputs.
[0071] In a possible implementation, the data processing module is configured to perform an addition operation on the N fifth feature maps and the X output layer outputs.
[0072] In a possible implementation, the apparatus further includes:
[0073] A model training module, configured to train the Transformer model according to the data processing result to obtain a trained Transformer model.
[0074] In a possible implementation, the obtaining module is configured to obtain a performance requirement, where the performance requirement is used to indicate the data processing accuracy of the Transformer model and / or the model size of the Transformer model;
[0075] Determine the number of the target modules and the insertion positions in the Transformer model according to the performance requirement.
[0076] In a possible implementation, the higher the data processing accuracy, the more the number of the target modules; and / or, the higher the data processing accuracy, the closer the insertion position of the target module in the Transformer model is to the embedding layer in the Transformer model; and / or, the larger the model size, the more the number of the target modules.
[0077] In a possible implementation, the Transformer model is a model after compression processing.
[0078] In a possible implementation, the processing of the data to be processed by the Transformer model includes:
[0079] Performing processing corresponding to a target task on the data to be processed by the Transformer model, where the target task includes: reading comprehension, text translation, paraphrase recognition, named entity recognition, text sentiment analysis, natural language inference, text automatic question answering, text intention recognition, text classification, text simplification, or text story generation.
[0080] In a third aspect, the present application provides a data processing method, where the method includes:
[0081] Receiving a performance requirement sent by a receiving end side, where the performance requirement is used to indicate the data processing accuracy of a Transformer model and / or the model size of the Transformer model;
[0082] Obtaining a target Transformer model that meets the performance requirement according to the performance requirement, where the target Transformer model includes a target network layer and target modules, and the target modules are configured to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output; the target operation is a non-linear operation based on convolution;
[0083] Send the target Transformer model to the terminal side.
[0084] In a possible implementation, the performance requirements include at least one of the following:
[0085] The accuracy requirement of the model, the latency requirement of the model, or the model compression ratio requirement of the model.
[0086] In a possible implementation, obtaining the target Transformer model according to the performance requirements includes:
[0087] Obtain a first Transformer model;
[0088] According to the performance requirements, determine the number M of the target modules and the insertion positions in the first Transformer model;
[0089] Obtain the target Transformer model according to the first Transformer model, the number M of the target modules, and the insertion positions.
[0090] In a possible implementation, the higher the data processing accuracy, the more the number of the target modules; and / or,
[0091] The higher the data processing accuracy, the closer the insertion position of the target module in the Transformer model is to the embedding layer in the Transformer model; and / or,
[0092] The larger the model size, the more the number of the target modules.
[0093] In a possible implementation, obtaining the target Transformer model according to the first Transformer model, the number of the target modules, and the insertion positions includes:
[0094] According to the number of the target modules and the insertion positions, insert the M target modules into the first Transformer model to obtain a second Transformer model;
[0095] Train the second Transformer model to obtain the target Transformer model.
[0096] In a possible implementation, obtaining the first Transformer model includes:
[0097] Receive the compression indication for the initial transformer model sent by the terminal side;
[0098] Obtain the initial transformer model, and perform compression processing on the initial transformer model to obtain the first transformer model.
[0099] In a fourth aspect, the present application provides a data processing device, and the device includes:
[0100] A receiving module, configured to receive the performance requirements sent by the terminal side, where the performance requirements are used to indicate the data processing accuracy of the transformer model and / or the model size of the transformer model;
[0101] An obtaining module, configured to obtain a target transformer model that meets the performance requirements according to the performance requirements, where the target transformer model includes a target network layer and a target module, and the target module is configured to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output; the target operation is a non-linear operation based on convolution;
[0102] A sending module, configured to send the target transformer model to the terminal side.
[0103] In a possible implementation, the performance requirements include at least one of the following:
[0104] The accuracy requirement of the model, the latency requirement of the model, or the model compression ratio requirement of the model.
[0105] In a possible implementation, the obtaining module is specifically configured to:
[0106] Obtain a first transformer model;
[0107] Determine the number M of the target modules and the insertion position in the first transformer model according to the performance requirements;
[0108] Obtain the target transformer model according to the first transformer model, the number M of the target modules, and the insertion position.
[0109] In a possible implementation, the higher the data processing accuracy, the more the number of the target modules; and / or,
[0110] The higher the data processing precision, the closer the insertion position of the target module in the Transformer model is to the embedding layer in the Transformer model; and / or,
[0111] The larger the model size, the more the number of target modules.
[0112] In a possible implementation, the obtaining the target Transformer model according to the first Transformer model, the number of target modules, and the insertion position;
[0113] According to the number of target modules and the insertion position, insert the M target modules into the first Transformer model to obtain a second Transformer model;
[0114] Perform model training on the second Transformer model to obtain the target Transformer model.
[0115] In a possible implementation, the obtaining module is specifically configured to:
[0116] Receive a compression instruction for the initial Transformer model sent by the terminal side;
[0117] Obtain the initial Transformer model, and perform compression processing on the initial Transformer model to obtain the first Transformer model.
[0118] In a fifth aspect, an embodiment of the present application provides an execution device, which may include a memory, a processor, and a bus system. Among them, the memory is used to store a program, and the processor is used to execute the program in the memory to execute the methods in the first aspect and any optional method thereof, and the methods in the third aspect and any optional method thereof as described above.
[0119] In a sixth aspect, an embodiment of the present application provides a training device, which may include a memory, a processor, and a bus system. Among them, the memory is used to store a program, and the processor is used to execute the program in the memory to execute the methods in the first aspect and any optional method thereof, and the methods in the third aspect and any optional method thereof as described above.
[0120] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, it causes the computer to execute the methods in the first aspect and any optional method thereof, and the methods in the third aspect and any optional method thereof as described above.
[0121] In an eighth aspect, an embodiment of the present application provides a computer program which, when running on a computer, causes the computer to execute the method according to the first aspect and any optional method thereof, and the method according to the third aspect and any optional method thereof.
[0122] In a ninth aspect, the present application provides a chip system. The chip system includes a processor for supporting an execution device or a training device to implement the functions involved in the above aspects, for example, sending or processing the data or information involved in the above methods. In a possible design, the chip system further includes a memory for storing necessary program instructions and data for the execution device or the training device. The chip system may be composed of chips or may include chips and other discrete devices.
[0123] An embodiment of the present application provides a data processing method. The method includes: obtaining a Transformer model, where the Transformer model includes a target network layer and a target module; obtaining data to be processed, and processing the data to be processed through the Transformer model to obtain a data processing result; where the target module is used to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output to obtain an updated feature map output; the target operation is a non-linear operation based on convolution. Through the above method, a target module is inserted into the Transformer model. The target module generates more feature maps (that is, the operation result obtained by the target module through the non-linear operation based on convolution), and fuses the operation result with the input of the target module, increasing the information carried in the feature maps output by the target network layer in the Transformer model. And since the number of parameters of the target module itself and the computing power overhead required during operation are very small, it improves the data processing accuracy of the model on the premise of reducing the number of model parameters and computing power overhead. Description of the Drawings
[0124] Figure 1 It is a schematic structural diagram of an artificial intelligence main framework;
[0125] Figure 2 It is a natural language processing system;
[0126] Figure 3 It is another natural language processing system;
[0127] Figure 4 It is a schematic diagram of related devices for natural language processing provided by an embodiment of the present application;
[0128] Figure 5a It is an architecture schematic of a Transformer layer;
[0129] Figure 5b It is a schematic diagram of an application architecture;
[0130] Figure 6a It is a schematic diagram of an embodiment of a data processing method provided by an embodiment of this application;
[0131] Figure 6b It is a schematic diagram of an application architecture provided by an embodiment of this application;
[0132] Figure 7 It is a schematic diagram of the structure of a neural network model in an embodiment of this application;
[0133] Figure 8 It is a schematic diagram of the structure of a transformer layer;
[0134] Figure 9 It is an operation schematic diagram of an attention head;
[0135] Figure 10 It is an operation schematic diagram of a target module provided by an embodiment of this application;
[0136] Figure 11 It is a schematic diagram of the structure of a neural network model provided by an embodiment of this application;
[0137] Figure 12 It is a schematic diagram of the structure of a neural network model provided by an embodiment of this application;
[0138] Figure 13 It is a schematic diagram of the structure of a neural network model provided by an embodiment of this application;
[0139] Figure 14 It is a schematic diagram of the structure of a neural network model provided by an embodiment of this application;
[0140] Figure 15 It is a schematic diagram of the structure of a neural network model provided by an embodiment of this application;
[0141] Figure 16 It is a schematic diagram of the structure of a neural network model provided by an embodiment of this application;
[0142] Figure 17 It is a schematic diagram of the structure of a neural network model provided by an embodiment of this application;
[0143] Figure 18 It is a schematic diagram of an application architecture provided by an embodiment of this application;
[0144] Figure 19 It is a schematic diagram of an application architecture provided by an embodiment of this application;
[0145] Figure 20a Schematic diagram of an embodiment of a data processing method provided by an embodiment of the present application;
[0146] Figure 20b Schematic diagram of an embodiment of a data processing method provided by an embodiment of the present application;
[0147] Figure 21 Schematic diagram of an embodiment of a data processing device provided by an embodiment of the present application;
[0148] Figure 22 Schematic diagram of a structure of an execution device provided by an embodiment of the present application;
[0149] Figure 23 Schematic diagram of a structure of a training device provided by an embodiment of the present application;
[0150] Figure 24 Schematic diagram of a structure of a chip provided by an embodiment of the present application. Detailed implementation manners
[0151] The embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention. The terms used in the embodiments of the present invention are only for explaining the specific embodiments of the present invention, rather than aiming to limit the present invention.
[0152] The embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0153] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0154] First, the overall working process of the artificial intelligence system will be described. Please refer to Figure 1 , Figure 1Shown is a schematic structural diagram of an artificial intelligence entity framework. The above artificial intelligence entity framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecological process of the system.
[0155] (1) Infrastructure
[0156] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform. It communicates with the external through sensors; the computing power is provided by intelligent chips (such as hardware acceleration chips like CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.
[0157] (2) Data
[0158] The data in the layer above the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the perception data such as force, displacement, liquid level, temperature, humidity, etc.
[0159] (3) Data Processing
[0160] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0161] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.
[0162] Reasoning refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formal information for machine thinking and problem-solving according to the reasoning control strategy. The typical function is search and matching.
[0163] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, prediction, etc.
[0164] (4) General Capabilities
[0165] After the data is processed by the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system, such as translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0166] (5) Intelligent Products and Industry Applications
[0167] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which are the encapsulation of the overall artificial intelligence solution, productize intelligent information decision-making, and realize the actual application. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.
[0168] This application can be applied to the field of natural language processing in the field of artificial intelligence. Below, multiple application scenarios implemented in products will be introduced.
[0169] To better understand the solution of the embodiments of this application, first, in combination with Figures 1 to 3 A simple introduction to the possible application scenarios of the embodiments of this application will be given.
[0170] Figure 2 A natural language processing system is shown. The natural language processing system includes a user device and a data processing device. Among them, the user device includes intelligent terminals such as mobile phones, personal computers, or information processing centers. The user device is the initiator of natural language data processing and is the initiator of requests such as language questions and queries. Usually, the user initiates requests through the user device.
[0171] The above-mentioned data processing device can be a device or server with data processing functions such as a cloud server, a network server, an application server, and a management server. The data processing device receives query statements / voices / texts, etc. (such as the data to be processed in the embodiments of this application) from the intelligent terminal through an interaction interface, and then performs language data processing in ways such as machine learning, deep learning, searching, reasoning, and decision-making through the memory for storing data and the processor for data processing, and feeds back the processing results (such as the data processing results in the embodiments of this application) to the user device. The memory in the data processing device can be a general term, including local storage and a database for storing historical data. The database can be on the data processing device or on other network servers.
[0172] In Figure 2In the natural language processing system shown, the user device can receive a user's instruction. For example, the user device can receive a piece of text input by the user, and then send a request to the data processing device, causing the data processing device to execute a natural language processing application (such as text classification, text reasoning, named entity recognition, translation, etc.) on the piece of text obtained by the user device, so as to obtain the processing result of the corresponding natural language processing application for the piece of text (such as classification result, reasoning result, named entity recognition result, translation result, etc.). Exemplarily, the user device can receive a piece of Chinese input by the user, and then send a request to the data processing device, causing the data processing device to perform entity classification on the piece of Chinese, so as to obtain the entity classification result for the piece of Chinese; Exemplarily, the user device can receive a piece of Chinese input by the user, and then send a request to the data processing device, causing the data processing device to translate the piece of Chinese into English, so as to obtain the English translation for the piece of Chinese.
[0173] In Figure 2 , the data processing device can execute the data processing method of the embodiment of the present application.
[0174] Figure 3 Another natural language processing system is shown. In Figure 3 , the user device directly serves as the data processing device. This user device can directly receive input from the user (such as the data to be processed in the embodiment of the present application) and be directly processed by the hardware of the user device itself. The specific process is similar to Figure 2 and can refer to the above description, which will not be elaborated here.
[0175] In Figure 3 the natural language processing system shown, the user device can receive a user's instruction. For example, the user device can receive a piece of text input by the user, and then the user device itself executes a natural language processing application (such as text classification, text reasoning, named entity recognition, translation, etc.) on the piece of text, so as to obtain the processing result of the corresponding natural language processing application for the piece of text (such as classification result, reasoning result, named entity recognition result, translation result, etc.). Exemplarily, the user device can receive a piece of Chinese input by the user and perform entity classification on the piece of Chinese, so as to obtain the entity classification result for the piece of Chinese; Exemplarily, the user device can receive a piece of Chinese input by the user and translate the piece of Chinese into English, so as to obtain the English translation for the piece of Chinese.
[0176] In an embodiment of the present application, a user device may store a Transformer model, and after the operating system (OS) or an application (APP) calls the model each time, an inference task is performed according to the Transformer model.
[0177] In Figure 3 , the user device itself can execute the data processing method of the embodiment of the present application.
[0178] Figure 4 is a schematic diagram of a related device 300 for natural language processing provided by an embodiment of the present application.
[0179] The above Figure 2 and Figure 3 The user device in Figure 4 may specifically be the local device 301 or the local device 302 in Figure 2 The data processing device in Figure 4 may specifically be the execution device 310 in
[0180] Figure 2 and Figure 3 The processor in
[0181] Since the embodiment of the present application involves a large number of neural network applications, for the convenience of understanding, the related terms and related concepts such as neural networks involved in the embodiment of the present application will be introduced below.
[0182] (1) Neural network
[0183] A neural network may be composed of neural units. A neural unit may refer to an operation unit that takes xs and an intercept 1 as inputs, and the output of the operation unit may be:
[0184] Among them, s = 1, 2, ……, n, where n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neuron. f is the activation function of the neuron, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer, and the activation function can be the sigmoid function. The neural network is a network formed by connecting multiple such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.
[0185] (2) Transformer layer
[0186] Refer to Figure 5a , Figure 5a which is a schematic diagram of the architecture of a Transformer layer. As Figure 5a shown, the neural network includes an embedding layer and at least one Transformer layer. The at least one Transformer layer can be N Transformer layers (N is an integer greater than 0). Among them, each Transformer layer includes an attention layer, an add & norm layer, a feed forward layer, and an add & norm layer that are adjacent to each other in sequence. In the embedding layer, the current input is embedded to obtain multiple feature vectors; in the attention layer, P input vectors are obtained from the previous layer of the first Transformer layer. Taking any first input vector among the P input vectors as the center, based on the correlation degree between each input vector within the preset attention window range and this first input vector, an intermediate vector corresponding to this first input vector is obtained, and thus P intermediate vectors corresponding to the P input vectors are determined; in the pooling layer, the P intermediate vectors are combined into Q output vectors, and the multiple output vectors obtained from the last Transformer layer in the Transformer layer are used as the feature representation of the current input.
[0187] Next, the above steps will be specifically introduced with specific examples.
[0188] First, in the embedding layer, the current input is embedded to obtain multiple feature vectors.
[0189] The embedding layer can be referred to as the input embedding layer. The current input can be a text input, for example, it can be a piece of text or a sentence. The text can be in Chinese, English, or other languages. After obtaining the current input, the embedding layer can perform embedding processing on each word in the current input to obtain the feature vectors of each word. In some embodiments, as Figure 1 shown, the embedding layer includes an input embedding layer and a positional encoding layer. In the input embedding layer, word embedding processing can be performed on each word in the current input to obtain the word embedding vectors of each word. In the positional encoding layer, the position of each word in the current input can be obtained, and then position vectors are generated for the positions of each word. In some examples, the position of each word can be the absolute position of each word in the current input. Taking the current input "When should I pay back Huabei" as an example, the position of "When" can be represented as the first position, the position of "should" can be represented as the second position, and so on. In some examples, the position of each word can be the relative position between each word. Still taking the current input "When should I pay back Huabei" as an example, the position of "When" can be represented as before "should", the position of "should" can be represented as after "When" and before "pay", and so on. When the word embedding vectors and position vectors of each word in the current input are obtained, the position vectors of each word can be combined with the corresponding word embedding vectors to obtain the feature vectors of each word, that is, multiple feature vectors corresponding to the current input are obtained. The multiple feature vectors can be represented as an embedding matrix with a preset dimension. It can be set that the number of feature vectors in the multiple feature vectors is M, and the preset dimension is H-dimensional, then the multiple feature vectors can be represented as an M×H embedding matrix.
[0190] Secondly, P input vectors can be obtained from the previous layer of the transformer layer. Taking any of the P input vectors as the center, based on the correlation between each input vector within the preset attention window range and this input vector, the intermediate vector corresponding to this input vector is obtained, and in this way, P intermediate vectors corresponding to the P input vectors are determined. The attention layer can also be referred to as the multi-head attention layer. In one example, the attention layer can be a fixed window multi-head attention layer.
[0191] (3) Attention mechanism
[0192] The attention mechanism mimics the internal process of biological observation behavior, that is, a mechanism that aligns internal experience and external sensations to increase the observation fineness of some areas, and can quickly screen out high-value information from a large amount of information using limited attention resources. The attention mechanism can quickly extract the important features of sparse data and is thus widely used in natural language processing tasks, especially machine translation. The self-attention mechanism is an improvement of the attention mechanism, which reduces the dependence on external information and is better at capturing the internal correlation of data or features. The essential idea of the attention mechanism can be rewritten as the following formula:
[0193] Among them, Lx = ||Source|| represents the length of Source. The meaning of the formula is to imagine the constituent elements in Source as being composed of a series of data pairs. At this time, given an element Query in the target Target, by calculating the similarity or correlation between Query and each Key, the weight coefficient of the Value corresponding to each Key is obtained, and then the Values are weighted and summed to obtain the final Attention value. So essentially, the Attention mechanism is to perform a weighted sum on the Value values of the elements in Source, and Query and Key are used to calculate the weight coefficients of the corresponding Values. Conceptually, Attention can be understood as selectively screening out a small amount of important information from a large amount of information and focusing on this important information, while ignoring most of the unimportant information. The focusing process is reflected in the calculation of the weight coefficients. The larger the weight, the more focused on the corresponding Value value, that is, the weight represents the importance of the information, and Value is the corresponding information. The self-attention mechanism can be understood as internal Attention (intra attention). The Attention mechanism occurs between the element Query in Target and all elements in Source. The self-attention mechanism refers to the Attention mechanism that occurs between the elements within Source or within Target, and can also be understood as the attention calculation mechanism in the special case where Target = Source. Its specific calculation process is the same, only the calculation object has changed.
[0194] (4) Natural language processing (NLP)
[0195] Natural language is the language of human beings. Natural language processing (NLP) is the processing of human language. Natural language processing is a process of systematically analyzing, understanding, and extracting information from text data in an intelligent and efficient manner. By using NLP and its components, we can manage very large amounts of text data, perform a large number of automated tasks, and solve various problems, such as automatic summarization, machine translation (MT), named entity recognition (NER), relation extraction (RE), information extraction (IE), sentiment analysis, speech recognition, question answering, and topic segmentation, etc.
[0196] Exemplarily, natural language processing tasks can be classified into the following categories.
[0197] Sequence labeling: Each word in a sentence requires the model to give a classification category based on the context. Such as Chinese word segmentation, part-of-speech tagging, named entity recognition, and semantic role labeling.
[0198] Classification task: The whole sentence outputs a classification value, such as text classification.
[0199] Sentence relation inference: Given two sentences, determine whether these two sentences have a certain nominal relationship. For example, entailment, QA, semantic rewriting, and natural language inference.
[0200] Generative task: Output a piece of text to generate another piece of text. Such as machine translation, text summarization, writing poems and sentences, and describing pictures in words.
[0201] The following are some exemplary natural language processing cases.
[0202] Word segmentation (or word breaker, WB): Segment continuous natural language text into a sequence of words with semantic rationality and integrity, which can solve the problem of cross ambiguity.
[0203] Named entity recognition (NER): Identify entities with specific meanings (people, places, organizations, times, works, etc.) in natural language text.
[0204] Part-of-speech tagging: Assign a part of speech (noun, verb, adjective, etc.) to each word in natural language text; Dependency parsing: Automatically analyze the syntactic components (subject, predicate, object, attributive, adverbial, complement, etc.) in a sentence, which can solve the problem of structural ambiguity.
[0205] Word embedding & semantic similarity: Represent words in a vectorized form and calculate the semantic similarity of words based on this, which can solve the problem of lexical language similarity.
[0206] Text semantic similarity: Rely on the vast amount of data across the network and deep neural network technology to achieve the ability to calculate the semantic similarity between texts, which can solve the problem of text semantic similarity.
[0207] (5) Ghost module
[0208] The existing Ghost module can generate more phantom feature maps with inexpensive linear operations, and the network performance can be improved through the fusion of phantom feature maps. Specifically, first, the network layers in the neural network need to be divided into two parts. Given the inherent feature maps of the first part, then use the Ghost module on the feature maps of the first part to generate more feature maps. Compared with the neural network without the Ghost module, without changing the size of the output feature maps, the total number of parameters and the computational complexity required in this Ghost module have both been reduced.
[0209] In the embodiments of this application, the target module is integrated into the Transformer model. The target module plays a role similar to that of the Ghost module, and the insertion position of the target module and the operations used in the target module are adaptively modified and adjusted.
[0210] The data processing method provided by the embodiments of the present application relates to the processing of natural language texts, and can be specifically applied to data processing methods such as data training, machine learning, and deep learning. It performs symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on training data, and finally obtains a trained transformer model. Moreover, the data processing method provided by the embodiments of the present application can use the above-mentioned trained transformer model to input input data (such as data to be processed) into the trained transformer model to obtain output data (such as data processing results). It should be noted that the model training method and data processing method related to the transformer model provided by the embodiments of the present application are inventions generated based on the same concept, and can also be understood as two parts of a system or two stages of an overall process: such as the model training stage and the model application stage.
[0211] Next, the architectures of the model training stage and the model application stage in the embodiments of the present application will be introduced.
[0212] The following Figure 5b will introduce the system architecture provided by the embodiments of the present application in detail. Figure 5b is a schematic diagram of the system architecture provided by the embodiments of the present application. As Figure 5b shown, the system architecture 500 includes an execution device 510, a training device 520, a database 530, a client device 540, a data storage system 550, and a data acquisition system 560.
[0213] The execution device 510 includes a computing module 511, an I / O interface 512, a preprocessing module 513, and a preprocessing module 514. The computing module 511 may include a target model / rule 501, and the preprocessing module 513 and the preprocessing module 514 are optional.
[0214] The data acquisition device 560 is used to acquire training samples. The training samples may be image data, text data, audio data, etc. In the embodiments of the present application, the training samples are the data used for training the transformer model (such as data to be processed). After acquiring the training samples, the data acquisition device 560 stores these training samples in the database 530.
[0215] It should be understood that the database 530 may also maintain a transformer model.
[0216] The training device 520 can train the transformer model based on the training samples maintained in the database 530 to obtain the target model / rule 501. In the embodiments of the present application, the target model / rule 501 may be a trained transformer model.
[0217] It should be noted that in practical applications, the training samples maintained in the database 530 may not all come from the collection of the data acquisition device 560, and it is also possible to be received from other devices. Additionally, it should be noted that the training device 520 may not necessarily train the target model / rule 501 entirely based on the training samples maintained in the database 530, and it is also possible to obtain training samples from the cloud or other places for model training. The above descriptions should not be regarded as limitations on the embodiments of the present application.
[0218] The target model / rule 501 trained by the training device 520 can be applied to different systems or devices, such as being applied to Figure 5b the execution device 510 as shown, and the execution device 510 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, a vehicle-mounted terminal, etc., or can also be a server or the cloud, etc.
[0219] Specifically, the training device 520 can transfer the transformer model to the execution device.
[0220] In Figure 5b it, the execution device 510 configures an input / output (I / O) interface 512 for data interaction with external devices, and a user can input data (such as the data to be processed in the embodiments of the present application) to the I / O interface 512 through the client device 540.
[0221] The preprocessing modules 513 and 514 are used to preprocess the input data received through the I / O interface 512. It should be understood that there may be no preprocessing modules 513 and 514 or only one preprocessing module. When the preprocessing modules 513 and 514 do not exist, the computing module 511 can directly process the input data.
[0222] During the preprocessing of the input data by the execution device 510, or during the relevant processing such as the computing module 511 of the execution device 510 performing calculations, the execution device 510 can call data, code, etc. in the data storage system 550 for corresponding processing, or can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 550.
[0223] Finally, the I / O interface 512 presents the processing result (such as the data processing result in the embodiments of the present application) to the client device 540, so as to provide it to the user.
[0224] In Figure 5bIn the case shown, the user can manually input data, and the "manual input of data" can be operated through the interface provided by the I / O interface 512. In another case, the client device 540 can automatically send input data to the I / O interface 512. If the client device 540 is required to automatically send input data and user authorization is needed, the user can set the corresponding permissions in the client device 540. The user can view the results output by the execution device 510 in the client device 540, and the specific presentation form can be specific ways such as display, sound, and action. The client device 540 can also be used as a data collection end to collect the input data input to the I / O interface 512 and the output results of the output I / O interface 512 as new sample data, and store them in the database 530. Of course, it is also possible not to collect data through the client device 540, but to directly store the input data input to the I / O interface 512 and the output results of the output I / O interface 512 as new sample data in the database 530.
[0225] It should be noted that Figure 5b This is only a schematic diagram of a system architecture provided by the embodiments of the present application. The positional relationships among the devices, components, modules, etc. shown in the figure do not constitute any limitations. For example, in Figure 5b the data storage system 550 is an external memory relative to the execution device 510. In other cases, the data storage system 550 can also be placed in the execution device 510. It should be understood that the above execution device 510 can be deployed in the client device 540.
[0226] From the inference side of the model:
[0227] In the embodiments of the present application, the computing module 511 of the above execution device 520 can obtain the code stored in the data storage system 550 to implement the data processing method in the embodiments of the present application.
[0228] In the embodiments of the present application, the computing module 511 of the execution device 520 may include hardware circuits (such as application specific integrated circuit (ASIC), field-programmable gate array (FPGA), general-purpose processor, digital signal processor (DSP), microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with the function of executing instructions, such as CPU, DSP, etc., or a hardware system without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above hardware systems without the function of executing instructions and hardware systems with the function of executing instructions.
[0229] Specifically, the computing module 511 of the execution device 520 may be a hardware system with the function of executing instructions. The data processing method provided in the embodiments of the present application may be software code stored in the memory. The computing module 511 of the execution device 520 may obtain the software code from the memory and execute the obtained software code to implement the data processing method provided in the embodiments of the present application.
[0230] It should be understood that the computing module 511 of the execution device 520 may be a combination of a hardware system without the function of executing instructions and a hardware system with the function of executing instructions. Some steps of the data processing method provided in the embodiments of the present application may also be implemented by the hardware system without the function of executing instructions in the computing module 511 of the execution device 520, which is not limited here.
[0231] From the perspective of model training:
[0232] In the embodiments of the present application, the above training device 520 may obtain the code stored in the memory ( Figure 5b not shown in the figure, which may be integrated with the training device 520 or deployed separately from the training device 520) to implement the data processing method in the embodiments of the present application.
[0233] In the embodiments of the present application, the training device 520 may include a hardware circuit (such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, etc.), or a combination of these hardware circuits. For example, the training device 520 may be a hardware system with the function of executing instructions, such as a CPU, a DSP, etc., or a hardware system without the function of executing instructions, such as an ASIC, an FPGA, etc., or a combination of the above hardware systems without the function of executing instructions and hardware systems with the function of executing instructions.
[0234] Specifically, the training device 520 may be a hardware system with the function of executing instructions. The data processing method provided by the embodiments of the present application may be software code stored in a memory. The training device 520 may obtain the software code from the memory and execute the obtained software code to implement the data processing method provided by the embodiments of the present application.
[0235] It should be understood that the training device 520 may be a combination of a hardware system without the function of executing instructions and a hardware system with the function of executing instructions. Some steps of the data processing method provided by the embodiments of the present application may also be implemented by the hardware system without the function of executing instructions in the training device 520, which is not limited here.
[0236] First, take the model inference stage as an example to illustrate the data processing method provided by the embodiments of the present application.
[0237] Refer to Figure 6a , Figure 6a is a schematic diagram of an embodiment of a data processing method provided by the embodiments of the present application. The data processing method provided by the embodiments of the present application may be applied in an execution device. The execution device may be a terminal device such as a mobile phone, a tablet computer, a laptop computer, a smart wearable device, etc. The execution device may also be a cloud-side server, as Figure 6a shown, a data processing method provided by the embodiments of the present application includes:
[0238] 601. Obtain a Transformer model; the Transformer model includes a target network layer and a target module.
[0239] In the embodiments of the present application, refer to Figure 6b , and a Transformer model may be obtained based on the service provided by the cloud side.
[0240] In the embodiments of the present application, a terminal device or a cloud-side server may obtain a Transformer model for model inference, where the Transformer model may be a trained Transformer model. For example, the Transformer model may be a pre-trained model or a model after model fine-tuning.
[0241] Next, the general structure of the Transformer model will be described:
[0242] Referring to Figure 7 , Figure 7 is a schematic structure of a Transformer model in the embodiments of the present application. The Transformer model may include an embedding layer and a plurality of Transformer layers connected in sequence. It should be understood that Figure 7 The structure is only an example, and the number of Transformer layers can be set as needed. For example, only one Transformer layer can be set, or more Transformer layers can be set.
[0243] The following describes the specific working processes of each layer in the Transformer model:
[0244] 1. Embedding layer
[0245] The embedding layer can perform embedding processing on the input to obtain a plurality of feature vectors. The core feature of the Transformer model lies in its unique attention mechanism. When processing natural language, such as a sentence, the Transformer model uses this attention mechanism to assign different attention coefficients to each word vector in the sentence, so as to more comprehensively consider the influence of the context in the sentence on each word. The embedding layer can obtain N embedding vectors X l . The attention layer is connected to the embedding layer, obtains N embedding vectors from the embedding layer as input vectors, synthesizes each input vector based on the correlation degree between the input vectors in the N input vectors, obtains N output vectors, and outputs them to the subsequent Transformer layers. The Transformer layer obtains the output of the previous layer as the input vector and performs an operation similar to that of the previous-level Transformer layer.
[0246] 2. Transformer layer
[0247] Referring to Figure 8 , Figure 8 is a schematic structure of a Transformer layer. As Figure 8As shown, the transformer layer may include a multi-head attention layer (or simply referred to as an attention layer), an add & norm layer, a feed forward net (FFN), and an add & norm layer that are adjacent to each other in sequence.
[0248] Among them, the multi-head attention layer obtains N input vectors X from its previous layer l , and the N input vectors X l can also be represented as a matrix X. The multi-head attention layer adopts a self-attention mechanism to transform each vector based on the correlation degree between vectors, obtaining N output vectors, and the N output vectors can also be represented as a matrix Y. It can be understood that when the multi-head attention layer is the layer directly connected to the embedding layer, such as Figure 7 the transformer layer directly connected to the embedding layer in [reference], the input vectors it obtains are the embedding vectors output by the embedding layer; when the multi-head attention layer is the multi-head attention layer included in the subsequent transformer layer, such as Figure 7 the multi-head attention layer included in the transformer layer directly connected to the previous-level transformer layer in [reference], the input vectors it obtains are the output vectors of the previous-level transformer layer. The multi-head attention layer may include multiple attention heads head (such as Figure 8 Head 1, Head 2,..., Head N shown in [reference]).
[0249] Figure 9 FIG. [reference] is a schematic diagram of the operation of an attention head head, which shows how the attention head head transforms the input matrix X into the output matrix Y. As shown in Figure 9 [reference], the attention head head can respectively use the first transformation matrix Q, the second transformation matrix K, and the third transformation matrix V to transform each input vector Xi in the N input vectors <X1, X2,..., XN>, obtaining the first intermediate vector (q vector), the second intermediate vector (k vector), and the third intermediate vector (v vector) corresponding to each input vector.
[0250] In terms of operation, the first transformation matrix Q, the second transformation matrix K, and the third transformation matrix V can be respectively used to perform a linear transformation on the input matrix X composed of N input vectors, obtaining the Q matrix, the K matrix, and the V matrix of the input matrix respectively, and then splitting the matrices respectively, so as to obtain the q vector, the k vector, and the v vector corresponding to each input vector.
[0251] It should be understood that the operation branch where the V matrix is obtained by performing a linear transformation on the input matrix X composed of N input vectors through the third transformation matrix V can also be referred to as the second branch in the embodiments of the present application.
[0252] Among them, for any i-th input vector Xi among the N input vectors, based on the dot product operation of the first intermediate vector (q vector, qi) corresponding to the i-th input vector and the second intermediate vectors (k vectors, kj) corresponding to each input vector Xj, the correlation degrees between the i-th input vector Xi and each input vector Xj are determined. Although the dot product result of qi and kj can also be directly determined as the correlation degree, more typically, the dot product result is first divided by a constant and then a softmax operation is performed, and the operation result is used as the correlation degree between the input vector Xi and Xj, that is:
[0253]
[0254] It should be understood that the operation branch where the dot product operation of the first intermediate vector (q vector, qi) and the second intermediate vector (k vector, kj) is located can also be referred to as the first branch in the embodiments of the present application.
[0255] Then, the correlation degrees αi,j between the i-th input vector Xi and each input vector Xj can be used as weight factors to perform a weighted combination of the third intermediate vectors (v vectors, vj) corresponding to each input vector Xj, and the i-th combined vector Ci corresponding to the i-th input vector Xi is obtained:
[0256]
[0257] Then, a vector sequence <C1, C2, …, CN> of N combined vectors corresponding to the N input vectors, or a matrix C, can be obtained. Based on this combined vector sequence, N output vectors can be obtained. Specifically, in one embodiment, the vector sequence of N combined vectors can be directly used as the N output vectors, that is, Yi = Ci. At this time, the output matrix Y is the combined vector matrix C, and can also be written as:
[0258]
[0259] The above is the description of the processing process of one attention head. In the MHA architecture, the MHA layer maintains m sets of transformation matrices, and each set of transformation matrices includes the aforementioned first transformation matrix Q, second transformation matrix K, and third transformation matrix V. Thus, the above operations can be performed in parallel to obtain m combined vector sequences (i.e., m matrices C), and each vector sequence includes N combined vectors obtained based on a set of transformation matrices. In such a case, the MHA layer performs a concatenation operation (concat) on the m combined vector sequences obtained, and obtains a concatenated matrix; then, the concatenated matrix is transformed by the fourth transformation matrix W to obtain the final output matrix Y. Splitting the output matrix Y corresponds to the N output vectors <Y1, Y2, …, YN>. Through the above operation process, the MHA layer performs transformation operations based on the correlation degrees between the N input vectors to obtain N output vectors.
[0260] In the embodiments of the present application, the transformer model may include a target network layer, where the target network layer may be an attention layer or a feed-forward layer in the transformer layer.
[0261] In a possible implementation, the target module may be located after the target network layer or embedded within the target network layer and located after the output of the intermediate layer of the target network layer, and the number of target modules may be one or more, which is not limited here.
[0262] In a possible implementation, the transformer model may be used to implement a target task, and the target task may include but is not limited to: reading comprehension, text translation, paraphrase recognition, named entity recognition, text sentiment analysis, natural language inference, text automatic question answering, text intent recognition, text classification, text simplification, or text story generation.
[0263] 602. Obtain the data to be processed, and process the data to be processed through the transformer model to obtain a data processing result; wherein, the target module is used to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result and the feature map output to obtain the updated feature map output; the target operation is a non-linear operation based on convolution.
[0264] In the embodiments of the present application, the data to be processed is obtained, where the data to be processed may be text data, and the data to be processed may be processed through the transformer model. The data to be processed may be the input data in the model inference process, and the transformer model is the model used in the model inference process.
[0265] In the embodiments of the present application, the target module may perform a target operation on the feature map output of the target network layer to obtain an operation result, where the feature map output may be the final output or the intermediate feature map output of the target network layer.
[0266] Next, the target operation is described:
[0267] Generally speaking, most linear operations can be used as the operations adopted in the target module. However, in the transformer model, simple linear operations do not help much in improving the performance of the model. Therefore, in the embodiments of the present application, non-linear operations are introduced on the basis of convolution operations.
[0268] Regarding the convolution operation:
[0269] The convolution operation in the embodiments of the present application may but is not limited to one-dimensional convolution operation, two-dimensional convolution operation, and depthwise separable convolution operation.
[0270] Among them, one-dimensional convolution (Conv1D) encodes local dependencies in the sequence direction and shows excellent performance for NLP tasks. For one-dimensional convolution, if the convolution operation is performed in the sequence direction (Conv1D_S), the input and output channels are d, and the dimension of the convolution kernel is W ∈ R d×d×k . After applying Conv1D_S, the output of the c-th channel of the i-th token can be expressed as:
[0271]
[0272] Similarly, if one-dimensional convolution (Conv1D_F) is performed in the feature direction, the input and output channels are n, and the dimension of the convolution kernel is W ∈ R n×n×k . After applying Conv1D_F, the output of the c-th channel of the i-th token can be expressed as:
[0273]
[0274] Among them, for two-dimensional convolution (Conv2D), both the input and output channels are 1, and the dimension of the convolution kernel is W ∈ R 1 ×1×k×k . After applying Conv2D, the output of the c-th channel of the i-th token can be expressed as:
[0275]
[0276] Although one-dimensional convolution (Conv1D_S) has strong expressive power, it requires introducing a lot of additional memory and computation. Compared with Conv1D, depthwise separable convolution (DWConv) performs convolution independently on each channel, and can reduce the number of parameters from d 2 k to dk. The weights representing the DWConv operation are W ∈ R d×k . After applying DWConv, the output of the c-th channel of the i-th token can be expressed as:
[0277]
[0278] Regarding the non-linear operation based on convolution:
[0279] In a possible implementation, to ensure that the numerical magnitudes of the input and output of the target module are close, the weight parameters included in the convolution kernel used in the convolution operation in the target module can be regularized so that all the weight parameters included in the convolution kernel are positive and the sum of the weight parameters is 1 or a value close to 1. For example, the difference from 1 can be within 0.01, 0.02, 0.03, 0.04, 0.05, 0.06, 0.07, 0.1.
[0280] In a possible implementation, the way of regularization processing can include, but is not limited to, Softmax regularization, L1 regularization, etc.
[0281] By regularizing the weight parameters included in the convolution kernel, the input and output numerical values of the convolution operation can be made as close as possible. This can make the model more stable during the training process, make the model performance robust, and reduce the waste of computing power resources caused by the redundant parameter tuning process.
[0282] In a possible implementation, the length and width dimensions of the feature map output are the same as those of the updated feature map output. To enable the operation result (the updated feature map output) obtained through the target operation to be fused (such as added or concatenated) with the feature map output of the target network layer, it is necessary to make the length and width dimensions of the updated feature map output the same as those of the feature map output of the target network layer.
[0283] Taking the depthwise separable convolution as the convolution operation and Softmax regularization as the regularization method as an example, refer to Figure 10 , Figure 10 which is an operation schematic diagram of a target module based on depthwise separable convolution. Softmax regularization is applied to the convolution kernel to ensure that the numerical magnitudes of the input and output of the target module are close.
[0284] In a possible implementation, the non-linear operation is used to perform non-linear processing on the result obtained by the convolution. The non-linear operation can include, but is not limited to, ReLU.
[0285] Next, the target network layer, that is, the insertion position of the target module in the transformer model, will be described:
[0286] In a possible implementation, the target network layer can include an attention layer.
[0287] Refer to Figure 11, the insertion position of the target module in the Transformer model can be after the output of the attention head. Specifically, the attention layer in the Transformer model can include M attention heads, where M is a positive integer. The feature map output of the target network layer can include the M feature map outputs of the M attention heads. The target module can perform N target operations on the M feature map outputs to obtain N first feature maps, and update the M feature map outputs of the M attention heads based on the N first feature maps.
[0288] Among them, each time the target module performs the above N target operations, a different convolution kernel is used for each target operation. That is to say, the target module can include N sub-modules, and each sub-module includes a convolution kernel. Furthermore, the input of the target module can be each of the M attention heads in the M attention heads, that is, the M feature map outputs of the M attention heads are input into the target module. The target module can perform N target operations on the M feature map outputs, and each target operation can obtain a first feature map.
[0289] Exemplarily, reference can be made to Figure 12 , Figure 12 in which the M feature map outputs (H1, H2, …, HM) of the M attention heads can undergo N target operations to obtain N first feature maps (G1, G2, …, GM). Specifically, if a Transformer layer in the Transformer model has N H attention heads, the outputs of multiple attention heads can be expressed as the sum of the outputs of these N H attention heads:
[0290]
[0291] Assume that the Transformer layer includes M attention heads. N ghost features (which can also be referred to as first feature maps in this embodiment) can be generated by these M attention heads through the target module. The calculation formula for the f-th ghost feature can be expressed as:
[0292]
[0293] Among them, Nonlinear is a non-linear operation, such as ReLU.
[0294] In a possible implementation, the M feature map outputs of the M attention heads can be updated based on the N first feature maps. For example, the N first feature maps can be subjected to an addition operation with the M feature map outputs of the M attention heads.
[0295] In the embodiments of the present application, since the target module can generate more feature maps through inexpensive operations, after adding and calculating the N first feature maps and the M feature map outputs of the M attention heads, the information carried in the feature maps output by the attention layer can be enriched, thereby improving the data processing accuracy of the model on the premise of paying less parameter quantity and computing power cost.
[0296] Referring to Figure 13 , the insertion position of the target module in the transformer model can be after the intermediate output of the attention head. The insertion position of the target module in the transformer model can be at the position after the dot product of the k vector and the q vector in the attention head and passing through the softmax process. Specifically, the attention layer includes M attention heads, and each of the M attention heads includes a first branch and a second branch. The output of the first branch is obtained according to the dot product operation of the K vector and the Q vector, and the output of the second branch is obtained according to the V vector. The feature map output of the target network layer includes the M outputs of the M first branches of the M attention heads. Furthermore, N target operations can be performed on the outputs of the M first branches to obtain N second feature maps, and the outputs of the M first branches can be updated according to the N second feature maps. For example, the N second feature maps can be concatenated with the outputs of the M first branches (concat).
[0297] In the embodiments of the present application, the transformer model can be obtained by pruning the attention heads. For example, the K matrix and the Q matrix can be cropped to the size of A*M, while the V matrix is not cropped and retains the size of A*(M+N). The target module can generate a new matrix with a size of A*N and perform a dot product with N v vectors in the V matrix. Equivalent to looking at the outputs of the first branch and the second branch, the output size is the same as that before cropping, and there is not much loss in the data volume. And because the parameter quantity of the target module itself and the computing power overhead required during operation are very small, it is equivalent to improving the data processing accuracy of the model on the premise of reducing the model parameter quantity and computing power overhead.
[0298] In a possible implementation, the attention layer includes M attention heads, and each of the M attention heads includes a third branch. The output of the third branch is obtained according to the dot product operation of the K vector, the Q vector, and the V vector. The feature map output of the target network layer includes the M outputs of the M third branches of the M attention heads;
[0299] Referring to Figure 14, the insertion position of the target module in the Transformer model can be after the intermediate output of the attention head. Specifically, the insertion position of the target module in the Transformer model can be at the position after the dot product of the k vector, q vector, and v vector in the attention head (at the output of the third branch). Specifically, the outputs of the M third branches can be subjected to N target operations to obtain N third feature maps, and the N third feature maps and the outputs of the M third branches can be fused. For example, the N third feature maps can be concatenated with the outputs of the M third branches.
[0300] In the embodiments of the present application, the Transformer model can be obtained by pruning the attention head. For example, the K matrix, Q matrix, and V matrix can be cropped from a size of M + N dimensions to a size of M dimensions. Furthermore, the output dimension of the third branch is also M, and the target module can generate a new matrix with a dimension of N. After concatenating with the output of the third branch, a feature map with a dimension of M + N can be obtained. Equivalent to looking at the output of the third branch, the dimension is the same as the output without cropping, and there is not much loss in the amount of data. Moreover, since the number of parameters of the target module itself and the computing power overhead required during operation are very small, it is equivalent to improving the data processing accuracy of the model on the premise of reducing the number of model parameters and computing power overhead.
[0301] In a possible implementation, the target network layer can include a feed-forward layer FFN.
[0302] Refer to Figure 15 , the insertion position of the target module in the Transformer model can be after the intermediate output of the FFN. Specifically, in a possible implementation, the FFN includes an intermediate layer, and the intermediate layer includes X groups of neurons. The feature map output of the target network layer includes the X feature map outputs of the X groups of neurons. That is to say, the insertion position of the target module in the Transformer model can be after the intermediate layer of the FFN and before the output layer of the FFN.
[0303] In a possible implementation, the X feature map outputs can be subjected to N target operations to obtain N fourth feature maps, and the N fourth feature maps can be fused with the feature map outputs of the X groups of neurons. For example, the N fourth feature maps can be concatenated with the X feature map outputs of the X groups of neurons.
[0304] In the embodiments of the present application, the transformer model can be obtained by pruning the intermediate layer of the FFN. For example, the output feature map of the neuron can be cropped from the size of dimension M+N to the size of dimension M, and the target module can generate a new matrix of dimension N. After concatenating with the X feature map outputs, a feature map of dimension M+N can be obtained. That is, from the perspective of the output of the intermediate layer, the dimension is the same as that of the output without cropping, and there is not much loss in the amount of data. Moreover, since the number of parameters of the target module itself and the computing power overhead required during operation are very small, it is equivalent to improving the data processing accuracy of the model on the premise of reducing the number of model parameters and computing power overhead.
[0305] Referring to Figure 16 , the insertion position of the target module in the transformer model can be after the final output of the FFN. Specifically, in one possible implementation, the FFN includes an intermediate layer and an output layer. The intermediate layer includes X groups of neurons, and the output layer is used to process the X feature map outputs of the X groups of neurons to obtain X output layer outputs. The feature map output of the target network layer includes the X output layer outputs. That is to say, the insertion position of the target module in the transformer model can be after the output layer of the FFN.
[0306] In one possible implementation, the X output layer outputs can be subjected to N target operations to obtain N fifth feature maps, and the N fifth feature maps can be fused with the X output layer outputs. For example, the N fifth feature maps can be added to the X output layer outputs.
[0307] In the embodiments of the present application, since the target module can generate more feature maps through inexpensive operations, after adding the N fifth feature maps to the X output layer outputs, the information carried in the feature maps output by the FFN can be enriched, thereby improving the data processing accuracy of the model on the premise of paying less parameter and computing power costs.
[0308] In one possible implementation, the target module can also be inserted at other positions in the transformer model. For example, it can be inserted at the position before the dot product of the k vector and the q vector after linear transformation by the first transformation matrix K. For example, it can be inserted at the position before the dot product of the k vector and the q vector after linear transformation by the second transformation matrix Q. For example, it can be inserted at the position before softmax after the dot product of the k vector and the q vector. For example, it can be inserted at the position before the dot product with the v vector after linear transformation by the third transformation matrix V (as Figure 17 shown).
[0309] In the embodiment of the present application, a target module is inserted into the transformer model. More feature maps are generated through the target module (that is, the operation result obtained by the target module through convolution-based non-linear operations), and the operation result is fused with the input of the target module, increasing the information carried in the feature maps output by the target network layer in the transformer model. Moreover, since the number of parameters of the target module itself and the computing power overhead required during operation are very small, it is equivalent to improving the data processing accuracy of the model on the premise of reducing the number of model parameters and computing power overhead.
[0310] As shown in Table 1, adding the target module to compressed or original transformer models such as BERT, RoBERTa, and ELECTRA significantly improves the data processing accuracy of the model with almost no additional parameters and calculations.
[0311] Table 1
[0312]
[0313] The embodiment of the present application provides a data processing method. The method includes: obtaining a transformer model, where the transformer model includes a target network layer and a target module; obtaining data to be processed, and processing the data to be processed through the transformer model to obtain a data processing result; where the target module is used to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output to obtain the updated feature map output; the target operation is a convolution-based non-linear operation. Through the above method, a target module is inserted into the transformer model. More feature maps are generated through the target module (that is, the operation result obtained by the target module through convolution-based non-linear operations), and the operation result is fused with the input of the target module, increasing the information carried in the feature maps output by the target network layer in the transformer model. Moreover, since the number of parameters of the target module itself and the computing power overhead required during operation are very small, it is equivalent to improving the data processing accuracy of the model on the premise of reducing the number of model parameters and computing power overhead.
[0314] First, take the model training stage as an example to illustrate the data processing method provided by the embodiment of the present application.
[0315] Refer to Figure 18 , Figure 18 which is a schematic diagram of an application architecture of an embodiment of the present application. Among them, the target module can be used to help users improve the effect of a given basic model and give a new model that meets the hardware constraints. AsFigure 18 As shown, the user can input the performance requirements of the required model, such as the computational load constraint. The cloud-side server can calculate the number of target modules that can be added and the insertion positions based on the performance requirements, and output a new model that meets the user's needs.
[0316] Referring to Figure 19 , Figure 19 , which is a schematic diagram of an application architecture of an embodiment of the present application. Among them, the target module proposed by the present invention can cooperate with other model compression methods (pruning, quantization, etc.) to provide cloud services for model compression. As Figure 19 shown, according to the performance requirements on the terminal side, the cloud-side server can compress the basic model (pruning, quantization, etc.), and select an appropriate number of target modules to insert into the appropriate positions of the compressed model, and return it to the terminal side.
[0317] Referring to Figure 20a , Figure 20a , which is a schematic diagram of an embodiment of a data processing method provided by an embodiment of the present application. The data processing method provided by an embodiment of the present application can be applied to a cloud-side server. As Figure 20a shown, the data processing method provided by an embodiment of the present application includes:
[0318] 2001. Obtain a Transformer model, where the Transformer model includes a target network layer and a target module;
[0319] In an embodiment of the present application, the cloud-side server can obtain a Transformer model for model training. Among them, the Transformer model can be a pre-trained model or a model after model fine-tuning. The Transformer model can also be a model after pruning processing. For example, the Transformer model can be a model after pruning the attention heads of the attention layer, and the Transformer model can also be a model after pruning the neurons in the intermediate layer of the FFN.
[0320] In a possible implementation, performance requirements can be obtained, where the performance requirements are used to indicate the data processing accuracy and / or model size of the Transformer model, and based on the performance requirements, the number of target modules and the insertion positions in the Transformer model are determined.
[0321] Next, how to obtain performance requirements is described.
[0322] In an embodiment of the present application, the terminal device can send the performance requirements of the terminal device to the cloud-side server.
[0323] Specifically, the terminal device can send performance requirements to the cloud server, where the performance requirements include, but are not limited to, at least one of accuracy requirements, latency requirements, or model compression ratio requirements. Furthermore, the cloud server can obtain the performance requirements.
[0324] In the embodiments of this application, after receiving the performance requirements sent by the terminal device, the cloud server can compress the initial transformer model based on the received performance requirements, for example, performing pruning or quantization.
[0325] Taking pruning as an example:
[0326] In the embodiments of this application, the cloud server can obtain an initial neural network model with
[0327] a transformer structure. After receiving the performance requirements sent by the terminal device, the cloud server can determine the pruning size of the transformer model based on the received performance requirements. Specifically, when the accuracy requirement included in the performance requirements is relatively high, a larger pruning size of the transformer model can be determined; when the latency requirement included in the performance requirements is relatively high, a smaller pruning size of the transformer model can be determined; when the model compression ratio included in the performance requirements is relatively high, a larger pruning size of the transformer model can be determined. Specifically, the cloud server can determine the pruning size information of the transformer model based on a preset functional relationship or based on a preset corresponding relationship (for example, by looking up a table).
[0328] In a possible implementation, the size information can include the width size and depth size of the transformer model. Specifically, the width size information can include the number of attention heads included in each transformer layer in the transformer model and the number of neurons included in the intermediate layer in the feed-forward layer; the depth size information can include the number of transformer layers included in the transformer model.
[0329] In the embodiments of the present application, the calculation in the multi-head attention mechanism can be split into calculations for each attention head and then added together. Therefore, the pruning operation of the MHA layer can be performed on the number of attention heads. By changing the number of neurons included in the intermediate layer of the fully connected network (feed-forward layer), the intermediate layer of the fully connected network (feed-forward layer) is also scalable. For a transformer layer, the width can be pruned for the attention heads of the MHA and the neurons in the intermediate layer of the feed-forward layer. Exemplarily, the base model of BERT has 12 attention heads. Then, there can be 12 choices for the corresponding width size scaling, that is, the width can be any one of 1, 2, …, 12. Similarly, any number of neurons can be retained in the intermediate layer of the feed-forward layer.
[0330] In the embodiments of the present application, the cloud-side server may have an initial neural network model with a transformer structure. After receiving the performance requirements sent by the terminal device, the cloud-side server may determine the number of target modules and their insertion positions in the transformer model based on the received performance requirements.
[0331] In a possible implementation, the higher the data processing accuracy, the more the number of target modules; and / or, the higher the data processing accuracy, the closer the insertion position of the target module in the transformer model is to the embedding layer in the transformer model; and / or, the larger the model size, the more the number of target modules.
[0332] For example, when the performance requirements include a high accuracy requirement, it can be determined that the number of target modules is more, or the insertion position of the target module in the transformer model is closer to the embedding layer in the transformer model. When the performance requirements include a high latency requirement, it can be determined that the number of target modules is less.
[0333] It should be understood that when the target module is located at a position far from the embedding layer, the magnitude of the distance has a relatively small impact on the improvement of the model performance.
[0334] In a possible implementation, the size of the transformer model can be determined first, and then the number of target modules and their insertion positions in the transformer model can be further determined according to performance parameters such as the remaining available number of parameters and FLOPs that can be allocated.
[0335] It should be understood that the above pruning operation for the Transformer model is optional, and the target module can also be directly applied to the Transformer model (for example, the Transformer model is a pre-trained model or a model obtained after fine-tuning) to obtain better model performance.
[0336] 2002. Obtain the data to be processed, and process the data to be processed through the Transformer model to obtain a data processing result; wherein, the target module is used to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output to obtain the updated feature map output; the target operation is a non-linear operation based on convolution.
[0337] Step 2002 is the feed-forward process in model training, and the specific description can refer to the above description of step 602, which will not be elaborated here.
[0338] 2003. According to the data processing result, perform model training on the Transformer model to obtain a trained Transformer model.
[0339] After obtaining the data processing result, a loss can be constructed based on the data processing result, and the Transformer model can be trained based on the loss to obtain a trained Transformer model.
[0340] In a possible implementation, the above model training process can be knowledge distillation. Specifically, a model with better model accuracy can be used as the teacher model, and the above Transformer model with the target module added can be used as the student model. The knowledge distillation method is used to transfer the knowledge learned by the original large model (teacher model) to the model with the target module added (student model) after pruning. The objective function can include various distillation objectives, such as making the logits, word vector layer, and hidden layer state of the student model approximate those of the teacher model, etc.
[0341] In a possible implementation, after obtaining the trained Transformer model, the trained Transformer model can be fine-tuned. Specifically, the trained Transformer model can be fine-tuned with the true labels of the downstream task.
[0342] After obtaining the trained Transformer model, the cloud-side server can send the trained Transformer model back to the user device. Subsequently, the user device can use the model returned by the cloud side (the trained Transformer model) for inference. The inference process can refer to the descriptions of steps 601 and 602 in the above embodiments and will not be elaborated here.
[0343] Refer to Figure 20b , Figure 20b is a structural schematic diagram of a data processing method provided by an embodiment of the present application. As Figure 20b shown, the method includes:
[0344] 2004. Receive the performance requirements sent by the receiving end side, where the performance requirements are used to indicate the data processing accuracy of the Transformer model and / or the model size of the Transformer model;
[0345] In an embodiment of the present application, as a service on the cloud side, the cloud-side server can receive the performance requirements sent by the receiving end side. The performance requirements can be used to indicate the data processing accuracy of the Transformer model and / or the model size; among them, the performance requirements include but are not limited to at least one of accuracy requirements, latency requirements, or model compression ratio requirements.
[0346] In a possible implementation, the cloud-side server can obtain a first Transformer model and insert a target module into the first Transformer model to obtain the target Transformer model.
[0347] Among them, the first Transformer model can be a model to be trained specified by the receiving end side, or a model to be trained selected by the cloud-side server itself.
[0348] In a possible implementation, the cloud-side server can receive the compression indication for the initial Transformer model sent by the receiving end side, obtain the initial Transformer model, and perform compression processing on the initial Transformer model to obtain the first Transformer model. Among them, the initial Transformer model can be a model to be trained specified by the receiving end side, or a model to be trained selected by the cloud-side server itself.
[0349] That is to say, the above first Transformer model can be a compressed model, or an uncompressed model (for example, the first Transformer model is a pre-trained model or a model obtained after fine-tuning).
[0350] In 2005, according to the performance requirements, a target Transformer model that meets the performance requirements is obtained, where the target Transformer model includes a target network layer and a target module. The target module is used to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output; the target operation is a non-linear operation based on convolution.
[0351] In the embodiments of the present application, the cloud-side server can obtain a target Transformer model that meets the performance requirements according to the performance requirements. In the embodiments of the present application, after receiving the performance requirements sent by the terminal device, the cloud-side server can determine the number of target modules and the insertion position in the Transformer model (such as the first Transformer model in the above embodiments) based on the received performance requirements.
[0352] In a possible implementation, the higher the data processing accuracy, the more the number of target modules; and / or, the higher the data processing accuracy, the closer the insertion position of the target module in the Transformer model is to the embedding layer in the Transformer model; and / or, the larger the model size, the more the number of target modules.
[0353] For example, when the performance requirements include a relatively high accuracy requirement, it can be determined that the number of target modules is more, or the insertion position of the target module in the Transformer model is closer to the embedding layer in the Transformer model. When the performance requirements include a relatively high latency requirement, it can be determined that the number of target modules is less.
[0354] It should be understood that when the target module is located at a position far from the embedding layer, the magnitude of the distance has a relatively small impact on the improvement of the model performance.
[0355] In a possible implementation, the size of the Transformer model can be determined first, and then the number of target modules and the insertion position in the Transformer model can be further determined according to performance parameters such as the remaining available number of parameters and FLOPs that can be allocated.
[0356] Furthermore, the cloud server can obtain the target Transformer model according to the first Transformer model, the number M of the target modules, and the insertion position. Specifically, according to the number of the target modules and the insertion position, the M target modules can be inserted into the first Transformer model to obtain a second Transformer model, and then the second Transformer model can be trained to obtain the target Transformer model. Among them, the above model training can be knowledge distillation. Specifically, a model with better model accuracy can be used as the teacher model, and the above Transformer model with the target modules added can be used as the student model. The method of knowledge distillation can be used to transfer the knowledge learned by the original large model (teacher model) to the model with the target modules added (student model after pruning). The objective function can include various distillation objectives, such as making the logits, word vector layer, and hidden layer state of the student model approximate those of the teacher model, etc.
[0357] For the description of the target module, reference can be made to the description related to the target module in step 602 of the above embodiment, which will not be elaborated here.
[0358] 2006. Send the target Transformer model to the edge side.
[0359] After obtaining the target Transformer model, the cloud server can send the target Transformer model back to the user device. Furthermore, the user device can use the model returned by the cloud (target Transformer model) for inference. The inference process can refer to the descriptions of step 601 and step 602 in the above embodiment, which will not be elaborated here.
[0360] Refer to Figure 21 , Figure 21 FIG. shows a structural schematic of a data processing device provided by an embodiment of the present application. The device 2100 includes:
[0361] An acquisition module 2101, configured to acquire a Transformer model, where the Transformer model includes a target network layer and target modules;
[0362] For the specific description of the acquisition module 2101, reference can be made to the description of step 601 or step 2001, which will not be elaborated here.
[0363] A data processing module 2102, configured to obtain data to be processed, and process the data to be processed through the transformer model to obtain a data processing result; wherein, the target module is configured to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output to obtain the updated feature map output; the target operation is a non-linear operation based on convolution.
[0364] For the specific description of the data processing module 2102, reference may be made to the description in step 602 or step 2002, which will not be elaborated here.
[0365] In a possible implementation, the weight parameters included in the convolution kernel used for the convolution are obtained through regularization processing.
[0366] In a possible implementation, the convolution kernel used for the convolution satisfies at least one of the following conditions:
[0367] The difference between the sum of the weight parameters included in the convolution kernel and 1 is within a preset range;
[0368] The weight parameters included in the convolution kernel are positive numbers.
[0369] In a possible implementation, the length and width dimensions of the feature map output and the updated feature map output are the same.
[0370] In a possible implementation, the non-linear operation is used to perform non-linear processing on the result obtained by the convolution.
[0371] In a possible implementation, the target network layer includes an attention layer.
[0372] In a possible implementation, the attention layer includes M attention heads, and the feature map output of the target network layer includes the M feature map outputs of the M attention heads;
[0373] The data processing module is configured to perform the target operation N times on the M feature map outputs to obtain N first feature maps, and fuse the N first feature maps with the M feature map outputs of the M attention heads.
[0374] In a possible implementation, the data processing module is configured to perform an addition operation on the N first feature maps and the M feature map outputs of the M attention heads.
[0375] In a possible implementation, the attention layer includes M attention heads, each of the M attention heads includes a first branch and a second branch, the output of the first branch is obtained according to the dot product operation of the K vector and the Q vector, the output of the second branch is obtained according to the V vector, and the feature map output of the target network layer includes the outputs of the M first branches of the M attention heads;
[0376] The data processing module is configured to perform N target operations on the outputs of the M first branches to obtain N second feature maps, and fuse the N second feature maps with the outputs of the M first branches.
[0377] In a possible implementation, the data processing module is configured to perform a concatenation operation (concat) on the N second feature maps and the outputs of the M first branches.
[0378] In a possible implementation, the attention layer includes M attention heads, each of the M attention heads includes a third branch, the output of the third branch is obtained according to the dot product operation of the K vector, the Q vector, and the V vector, and the feature map output of the target network layer includes the outputs of the M third branches of the M attention heads;
[0379] The data processing module is configured to perform N target operations on the outputs of the M third branches to obtain N third feature maps, and fuse the N third feature maps with the outputs of the M third branches.
[0380] In a possible implementation, the data processing module is configured to perform a concatenation operation on the N third feature maps and the outputs of the M third branches.
[0381] In a possible implementation, the target network layer includes a feed-forward layer FFN.
[0382] In a possible implementation, the FFN includes an intermediate layer, the intermediate layer includes X groups of neurons, and the feature map output of the target network layer includes the X feature map outputs of the X groups of neurons;
[0383] The data processing module is configured to perform N target operations on the X feature map outputs to obtain N fourth feature maps, and fuse the N fourth feature maps with the feature map outputs of the X groups of neurons.
[0384] In a possible implementation, the data processing module is configured to perform a concatenation operation on the N fourth feature maps and the X feature map outputs of the X groups of neurons.
[0385] In a possible implementation, the FFN includes an intermediate layer and an output layer. The intermediate layer includes X groups of neurons. The output layer is configured to process the X feature map outputs of the X groups of neurons to obtain X output layer outputs. The feature map output of the target network layer includes the X output layer outputs;
[0386] The data processing module is configured to perform N target operations on the X output layer outputs to obtain N fifth feature maps, and fuse the N fifth feature maps with the X output layer outputs.
[0387] In a possible implementation, the data processing module is configured to perform an addition operation on the N fifth feature maps and the X output layer outputs.
[0388] In a possible implementation, the apparatus further includes:
[0389] A model training module 2103, configured to perform model training on the transformer model according to the data processing result to obtain a trained transformer model.
[0390] For the specific description of the model training module 2103, reference may be made to the description of step 2003, which will not be elaborated here.
[0391] In a possible implementation, the acquisition module is configured to acquire a performance requirement, where the performance requirement is used to indicate the data processing accuracy of the transformer model;
[0392] According to the performance requirement, determine the number of the target modules and the insertion positions in the transformer model.
[0393] In a possible implementation, the higher the data processing accuracy, the more the number of the target modules; and / or, the higher the data processing accuracy, the closer the insertion position of the target module in the transformer model is to the embedding layer in the transformer model.
[0394] In a possible implementation, the transformer model is a compressed model.
[0395] In a possible implementation, processing the data to be processed by the transformer model includes:
[0396] The to-be-processed data is processed by the transformer model for a target task, where the target task includes: reading comprehension, text translation, paraphrase recognition, named entity recognition, text sentiment analysis, natural language inference, text automatic question answering, text intention recognition, text classification, text simplification, or text story generation.
[0397] An embodiment of the present application further provides a data processing device, where the device includes:
[0398] A receiving module, configured to receive a performance requirement sent by the terminal side, where the performance requirement is used to indicate the data processing accuracy of the transformer model and / or the model size of the transformer model;
[0399] An obtaining module, configured to obtain a target transformer model that meets the performance requirement according to the performance requirement, where the target transformer model includes a target network layer and a target module, and the target module is configured to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output; the target operation is a non-linear operation based on convolution;
[0400] A sending module, configured to send the target transformer model to the terminal side.
[0401] In a possible implementation, the performance requirement includes at least one of the following:
[0402] The accuracy requirement of the model, the latency requirement of the model, or the model compression ratio requirement of the model.
[0403] In a possible implementation, the obtaining module is specifically configured to:
[0404] Obtain a first transformer model;
[0405] Determine the number M of the target modules and the insertion position in the first transformer model according to the performance requirement;
[0406] Obtain the target transformer model according to the first transformer model, the number M of the target modules, and the insertion position.
[0407] In a possible implementation, the higher the data processing accuracy, the more the number of the target modules; and / or,
[0408] The higher the data processing precision, the closer the insertion position of the target module in the Transformer model is to the embedding layer in the Transformer model; and / or,
[0409] The larger the model size, the more the number of target modules.
[0410] In a possible implementation, obtaining the target Transformer model according to the first Transformer model, the number of target modules, and the insertion position;
[0411] Inserting the M target modules into the first Transformer model according to the number of target modules and the insertion position to obtain a second Transformer model;
[0412] Training the second Transformer model to obtain the target Transformer model.
[0413] In a possible implementation, the obtaining module is specifically configured to:
[0414] Receive a compression instruction for the initial Transformer model sent by the terminal side;
[0415] Obtain the initial Transformer model, and perform compression processing on the initial Transformer model to obtain the first Transformer model.
[0416] Next, an execution device provided in an embodiment of the present application is introduced. Please refer to Figure 22 , Figure 22 FIG. is a schematic structural diagram of an execution device provided in an embodiment of the present application. The execution device 2200 may specifically be a virtual reality (VR) device, a mobile phone, a tablet computer, a laptop computer, a smart wearable device, a monitoring data processing device, or a server, etc., which is not limited herein. Specifically, the execution device 2200 includes: a receiver 2201, a transmitter 2202, a processor 2203, and a memory 2204 (where the number of processors 2203 in the execution device 2200 may be one or more, Figure 22 taking one processor as an example), where the processor 2203 may include an application processor 22031 and a communication processor 22032. In some embodiments of the present application, the receiver 2201, the transmitter 2202, the processor 2203, and the memory 2204 may be connected through a bus or other means.
[0417] The memory 2204 may include a read-only memory and a random access memory, and provide instructions and data to the processor 2203. A part of the memory 2204 may also include a non-volatile random access memory (NVRAM). The memory 2204 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.
[0418] The processor 2203 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, where the bus system may include, in addition to a data bus, a power bus, a control bus, a status signal bus, etc. However, for the sake of clear illustration, all kinds of buses are referred to as a bus system in the figure.
[0419] The method disclosed in the embodiments of the present application above can be applied to the processor 2203 or implemented by the processor 2203. The processor 2203 may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above method can be completed by the integrated logic circuit in the hardware of the processor 2203 or instructions in software form. The above-mentioned processor 2203 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and may further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 2203 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 2204, and the processor 2203 reads the information in the memory 2204 and combines its hardware to complete the steps of the above method.
[0420] The receiver 2201 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the execution device. The transmitter 2202 can be used to output digital or character information through the first interface; the transmitter 2202 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 2202 can also include a display device such as a display screen.
[0421] In one embodiment of the present application, the processor 2203 is configured to execute Figure 6a The data processing method executed by the device in the corresponding embodiment.
[0422] The present application also provides a training device. Figure 23 , Figure 23 is a structural diagram of a training device provided in an embodiment of the present application. The training device 2300 may be deployed with Figures 17 to 20a The data processing device described in the corresponding embodiment, specifically, the training device 2300 is implemented by one or more servers, and the training device 2300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 2323 (for example, one or more processors) and memory 2332, one or more storage media 2330 (for example, one or more mass storage devices) storing application programs 2342 or data 2344. Among them, the memory 2332 and the storage medium 2330 can be short-term storage or permanent storage. The program stored in the storage medium 2330 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the training device. Furthermore, the central processing unit 2323 can be configured to communicate with the storage medium 2330 to execute a series of instruction operations in the storage medium 2330 on the training device 2300.
[0423] The training device 2300 may also include one or more power supplies 2326, one or more wired or wireless network interfaces 2350, one or more input and output interfaces 2358; or, one or more operating systems 2341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0424] In the embodiment of the present application, the central processor 2323 is used to execute Figure 20a as well as Figure 20b The data processing method in the corresponding embodiment.
[0425] An embodiment of the present application also provides a computer program product, which, when running on a computer, causes the computer to execute the steps performed by the aforementioned execution device, or causes the computer to execute the steps performed by the aforementioned training device.
[0426] An embodiment of the present application also provides a computer-readable storage medium storing a program for signal processing, which, when running on a computer, causes the computer to execute the steps performed by the aforementioned execution device, or causes the computer to execute the steps performed by the aforementioned training device.
[0427] The execution device, training device, or terminal device provided in the embodiment of the present application may specifically be a chip, which includes a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit to cause the chip in the execution device to execute the data processing method described in the above embodiment, or to cause the chip in the training device to execute the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit inside the chip, such as a register, a cache, etc. The storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0428] Specifically, please refer to Figure 24 , Figure 24 which is a schematic structural diagram of the chip provided in the embodiment of the present application. The chip may be embodied as a neural network processor NPU 2400, and the NPU 2400 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 2403, which extracts matrix data from the memory and performs multiplication operations under the control of the controller 2404.
[0429] In some implementations, the arithmetic circuit 2403 includes multiple processing units (Process Engine, PE) inside. In some implementations, the arithmetic circuit 2403 is a two-dimensional systolic array. The arithmetic circuit 2403 may also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 2403 is a general matrix processor.
[0430] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 2402 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 2401 and performs matrix operations with matrix B. The partial results or final results of the obtained matrix are stored in the accumulator 2408.
[0431] The unified memory 2406 is used to store input data and output data. The weight data is directly transferred through the Direct Memory Access Controller (DMAC) 2405 and is transported to the weight memory 2402. The input data is also transported to the unified memory 2406 through the DMAC.
[0432] The BIU is the Bus Interface Unit, i.e., the bus interface unit 2410, which is used for the interaction between the AXI bus and the DMAC and the Instruction Fetch Buffer (IFB) 2409.
[0433] The bus interface unit 2410 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch memory 2409 to obtain instructions from the external memory, and is also used for the storage unit access controller 2405 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0434] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 2406, or transfer the weight data to the weight memory 2402, or transfer the input data to the input memory 2401.
[0435] The vector calculation unit 2407 includes multiple arithmetic processing units. When needed, it further processes the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.
[0436] In some implementations, the vector computing unit 2407 can store the processed output vectors into the unified memory 2406. For example, the vector computing unit 2407 can apply a linear function; or, a non-linear function to the output of the arithmetic circuit 2403, such as performing linear interpolation on the feature planes extracted by the convolutional layer, or, for another example, vectors of accumulated values, to generate activation values. In some implementations, the vector computing unit 2407 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vectors can be used as activation inputs to the arithmetic circuit 2403, such as for use in subsequent layers in a neural network.
[0437] The instruction fetch buffer 2409 connected to the controller 2404 is used to store the instructions used by the controller 2404;
[0438] The unified memory 2406, the input memory 2401, the weight memory 2402, and the instruction fetch buffer 2409 are all On-Chip memories. The external memory is private to this NPU hardware architecture.
[0439] Wherein, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above programs.
[0440] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0441] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, software program implementation is a better embodiment in more cases. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.
[0442] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0443] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are generated in whole or in part. The computer can be a general computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
Claims
1. A data processing method, characterized in that, the method includes: obtaining a Transformer model, the Transformer model including a target network layer and a target module; obtaining data to be processed, and processing the data to be processed through the Transformer model to obtain a data processing result; wherein, the target module is obtained based on a Ghost module, and a neural network with the Ghost module added generates more feature maps than a neural network without the Ghost module added, and the target module is used to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output to obtain the updated feature map output; the target operation is a non-linear operation based on convolution, and the data to be processed is at least one of the following: text, speech; Before obtaining the Transformer model, the method further includes: obtaining the performance requirements of the terminal device, the performance requirements being used to indicate the data processing accuracy of the Transformer model and / or the model size of the Transformer model; determining the number of the target modules and the insertion positions in the Transformer model according to the performance requirements; The processing the data to be processed through the Transformer model includes: performing processing corresponding to a target task on the data to be processed through the Transformer model, the target task including: reading comprehension, text translation, repetition recognition, named entity recognition, text sentiment analysis, natural language inference, text automatic question answering, text intention recognition, text classification, text simplification or text story generation.
2. The method according to claim 1, characterized in that, the weight parameters included in the convolution kernel adopted by the convolution are obtained through regularization processing.
3. The method according to claim 2, characterized in that, the convolution kernel adopted by the convolution satisfies at least one of the following conditions: the difference between the sum of the weight parameters included in the convolution kernel and 1 is within a preset range; the weight parameters included in the convolution kernel are positive numbers.
4. The method according to any one of claims 1 to 3, characterized in that, the length and width dimensions of the feature map output and the updated feature map output are the same.
5. The method according to any one of claims 2 to 3, characterized in that, the non-linear operation is used to perform non-linear processing on the result obtained by the convolution.
6. The method according to any one of claims 1 to 3, characterized in that, the target network layer includes an attention layer.
7. The method according to claim 6, characterized in that, the attention layer includes M attention heads, and the feature map output of the target network layer includes the M feature map outputs of the M attention heads; The performing a target operation on the feature map output of the target network layer to obtain an operation result, and fusing the operation result with the feature map output includes: Perform N target operations on the M feature map outputs to obtain N first feature maps, and fuse the N first feature maps with the M feature map outputs of the M attention heads.
8. The method according to claim 7, wherein, the fusing of the N first feature maps with the M feature map outputs of the M attention heads includes: performing an addition operation on the N first feature maps and the M feature map outputs of the M attention heads.
9. The method according to claim 6, wherein, the attention layer includes M attention heads, each attention head in the M attention heads includes a first branch and a second branch, the output of the first branch is obtained by a dot product operation of a K vector and a Q vector, the output of the second branch is obtained according to a V vector, and the feature map output of the target network layer includes the outputs of the M first branches of the M attention heads; the performing of a target operation on the feature map output of the target network layer to obtain an operation result and fusing the operation result with the feature map output includes: performing N target operations on the outputs of the M first branches to obtain N second feature maps, and fusing the N second feature maps with the outputs of the M first branches.
10. The method according to claim 9, wherein, the fusing of the N second feature maps with the outputs of the M first branches includes: performing a concatenation operation (concat) on the N second feature maps and the outputs of the M first branches.
11. The method according to claim 6, wherein, the attention layer includes M attention heads, each attention head in the M attention heads includes a third branch, the output of the third branch is obtained by a dot product operation of a K vector, a Q vector, and a V vector, and the feature map output of the target network layer includes the outputs of the M third branches of the M attention heads; the performing of a target operation on the feature map output of the target network layer to obtain an operation result and fusing the operation result with the feature map output includes: performing N target operations on the outputs of the M third branches to obtain N third feature maps, and fusing the N third feature maps with the outputs of the M third branches.
12. The method according to claim 11, wherein, the fusing of the N third feature maps with the outputs of the M third branches includes: performing a concatenation operation on the N third feature maps and the outputs of the M third branches.
13. The method according to any one of claims 1 to 3, wherein, the target network layer includes a feed-forward layer FFN.
14. The method according to claim 13, wherein, the FFN includes an intermediate layer, the intermediate layer includes X groups of neurons, and the feature map output of the target network layer includes the X feature map outputs of the X groups of neurons; Performing a target operation on the feature map output of the target network layer to obtain an operation result, and fusing the operation result with the feature map output, includes: Performing N target operations on the X feature map outputs to obtain N fourth feature maps, and fusing the N fourth feature maps with the feature map outputs of the X groups of neurons.
15. The method according to claim 14, wherein, The fusing of the N fourth feature maps with the feature map outputs of the X groups of neurons includes: Performing a splicing operation on the N fourth feature maps and the X feature map outputs of the X groups of neurons.
16. The method according to claim 13, wherein, The FFN includes an intermediate layer and an output layer. The intermediate layer includes X groups of neurons. The output layer is configured to process the X feature map outputs of the X groups of neurons to obtain X output layer outputs. The feature map output of the target network layer includes the X output layer outputs; Performing a target operation on the feature map output of the target network layer to obtain an operation result, and fusing the operation result with the feature map output, includes: Performing N target operations on the X output layer outputs to obtain N fifth feature maps, and fusing the N fifth feature maps with the X output layer outputs.
17. The method according to claim 16, wherein, The fusing of the N fifth feature maps with the X output layer outputs includes: Performing an addition operation on the N fifth feature maps and the X output layer outputs.
18. The method according to any one of claims 1 to 3, wherein, After processing the data to be processed by the transformer model, the method further includes: Performing model training on the transformer model according to the data processing result to obtain a trained transformer model.
19. The method according to claim 1, wherein, The higher the data processing accuracy, the more the number of target modules; and / or, the higher the data processing accuracy, the closer the insertion position of the target module in the transformer model is to the embedding layer in the transformer model; and / or, the larger the model size, the more the number of target modules.
20. The method according to claim 1, wherein, The transformer model is a compressed model.
21. A data processing method, wherein, The method includes: Receiving the performance requirements of the terminal device sent by the receiving end, where the performance requirements are used to indicate the data processing accuracy of the transformer model and / or the model size of the transformer model; According to the performance requirements, obtain a target Transformer model that meets the performance requirements. Among them, the target Transformer model includes a target network layer and a target module. The target module is obtained based on the Ghost module. A neural network with the Ghost module added generates more feature maps than a neural network without the Ghost module added. The target module is used to perform a target operation on the feature map output of the target network layer to obtain an operation result, and fuse the operation result with the feature map output; the target operation is a non-linear operation based on convolution. Send the target Transformer model to the edge side. The obtaining of the target Transformer model that meets the performance requirements according to the performance requirements includes: Obtain a first Transformer model. According to the performance requirements, determine the number M of the target modules and the insertion positions in the first Transformer model. According to the first Transformer model, the number M of the target modules, and the insertion positions, obtain the target Transformer model. The target Transformer model is used to perform processing corresponding to a target task on the data to be processed. The target tasks include: reading comprehension, text translation, paraphrase recognition, named entity recognition, text sentiment analysis, natural language inference, text automatic question answering, text intention recognition, text classification, text simplification, or text story generation; the data to be processed is at least one of the following: text, speech.
22. The method according to claim 21. It is characterized in that The performance requirements further include at least one of the following: The accuracy requirement of the model, the latency requirement of the model, or the model compression ratio requirement of the model.
23. The method according to claim 21. It is characterized in that The higher the data processing accuracy, the more the number of the target modules; and / or The higher the data processing accuracy, the closer the insertion position of the target module in the Transformer model is to the embedding layer in the Transformer model; and / or The larger the model size, the more the number of the target modules.
24. The method according to claim 21 or 23. It is characterized in that The obtaining of the target Transformer model according to the first Transformer model, the number of the target modules, and the insertion positions includes: According to the number of the target modules and the insertion positions, insert the M target modules into the first Transformer model to obtain a second Transformer model. Perform model training on the second Transformer model to obtain the target Transformer model.
25. The method according to any one of claims 21 to 23. It is characterized in that The obtaining of the first Transformer model includes: Receive the compression indication for the initial Transformer model sent by the terminal side; Obtain the initial Transformer model, and perform compression processing on the initial Transformer model to obtain the first Transformer model.
26. A data processing device, Characterized in that, The device includes a memory and a processor; the memory stores code, and the processor is configured to obtain the code and execute the method according to any one of claims 1 to 25.
27. A computer-readable storage medium, Characterized in that, It includes computer-readable instructions, and when the computer-readable instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 25.
28. A computer program product, Characterized in that, It includes computer-readable instructions, and when the computer-readable instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 25.
Citation Information
Patent Citations
Data processing method and related equipment
CN111368993A
Model structure, model training method, image enhancement method and equipment
CN112529150A