A kind of attention model, feature extraction method and related device
By introducing parallel neural network layers into the self-attention network to transform and add features, the problem of feature collapse is solved, the diversity and expression ability of features are improved, and the performance of attention model is improved.
Patent Information
- Application Number
- CN202110731775.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-29
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-06-29
AI Technical Summary
In self-attention networks, as the network deepens, features tend to become indistinguishable, resulting in features collapse, and thus degrading model performance. Shortcuts added in the prior art cannot effectively enhance the expressive ability of features.
One or more parallel neural network layers are introduced, and connected in parallel with the self-attention module and/or multi-layer perceptron. The input features are transformed through these parallel neural network layers, and the transformed features are added with the output features of the self-attention module and/or multi-layer perceptron to increase the diversity and expression capabilities of the features.
By increasing the diversity and expression ability of features, the performance of the attention model is improved, feature collapse is avoided, and the representation ability of the model is enhanced.
Smart Images

Figure CN113627163B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an attention model, a feature extraction method and related devices. Background Art
[0002] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0003] In recent years, self-attention networks have been well applied in many natural language processing (NLP) tasks, such as machine translation, sentiment analysis, and question answering. With the widespread application of self-attention networks, self-attention networks originated from the field of natural language processing have also achieved high performance in tasks such as image classification, object detection, and image processing.
[0004] In a self-attention network, due to the processing of features by the self-attention network layer, the features of the input data tend to become indistinguishable as the network deepens, and these indistinguishable features have weak representation capabilities. This phenomenon of features becoming indistinguishable as the network deepens is usually called feature collapse.
[0005] Currently, adding shortcuts to the self-attention network can alleviate the phenomenon of feature collapse and avoid the situation where features cannot be distinguished. However, the shortcut added to the self-attention network simply copies the input features of the self-attention network layer to the output of the self-attention network layer, which cannot enhance the expressiveness of the features, resulting in poor performance of the self-attention network. Summary of the invention
[0006] The present application provides an attention model and a feature extraction method, which can increase the diversity of features extracted by the attention model, enhance the expressiveness of the features, and thus improve the performance of the attention model.
[0007] In a first aspect, the present application provides an attention model, comprising: one or more serially connected self-attention networks, wherein the self-attention network comprises a self-attention module, a multi-layer perceptron, and a first neural network layer.
[0008] The self-attention module includes a plurality of parallel feature extraction layers and a fusion layer, wherein the fusion layer is respectively connected to the plurality of parallel feature extraction layers. The self-attention module is a network using a self-attention mechanism, which can associate different positions of an input sequence to calculate a representation of the same sequence.
[0009] The multilayer perceptron is serially connected to the self-attention module, and the multilayer perceptron includes a plurality of serial first fully connected layers. Specifically, the multilayer perceptron can also be called a fully connected neural network (FCN), and the multilayer perceptron includes an input layer, a hidden layer, and an output layer, and the number of hidden layers can be one or more layers. Among them, the network layers in the multilayer perceptron are all fully connected layers. That is, the input layer and the hidden layer of the multilayer perceptron are fully connected, and the hidden layer and the output layer of the multilayer perceptron are also fully connected.
[0010] The first neural network layer is connected in parallel to the self-attention module and one or more of the multilayer perceptrons, wherein the first neural network layer is used to perform feature transformation.
[0011] In this solution, another parallel neural network layer is introduced on the basis of the self-attention module and the multi-layer perceptron, and the parallel neural network layer performs feature transformation operation on the input features to obtain the transformed features. In addition, the transformed features are added to the output features of the self-attention module and / or the multi-layer perceptron to increase the diversity of the features output by the intermediate layer of the self-attention network, enhance the expressiveness of the features, and thus improve the performance of the attention model.
[0012] In one possible implementation, the self-attention network also includes: a second neural network layer, which is used to perform feature transformation; the first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron, or the second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.
[0013] In this scheme, by introducing the first neural network layer and the second neural network layer which are parallel to the self-attention module and the multi-layer perceptron respectively, the two parallel neural network layers perform feature transformation operations on the input features to obtain transformed features. In addition, the transformed features are added to the output features of the self-attention module and the output features of the multi-layer perceptron respectively to increase the diversity of the features output by the intermediate layer of the self-attention network, enhance the expressiveness of the features, and thus improve the performance of the attention model.
[0014] In a possible implementation, the first neural network layer includes a weight matrix and an activation function, the weight matrix is used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix. The activation function can be, for example, a nonlinear function such as a Sigmoid function, a Tanh function, or a ReLU function.
[0015] In a possible implementation, the weight matrix includes a plurality of sub-matrices, each of which is a circulant matrix. In simple terms, the weight matrix may be composed of a plurality of sub-matrices, and each of which is a circulant matrix.
[0016] In a possible implementation manner, the multiple parallel feature extraction layers respectively include different weight matrices.
[0017] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer, the second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.
[0018] In a possible implementation, the self-attention module and / or the multilayer perceptron also has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.
[0019] Compared with the first neural network layer, shortcuts do not change the input features, that is, they do not perform feature transformation on the input features. Therefore, shortcuts can also be regarded as a special feature processing method. On the basis of the first neural network layer, the introduction of parallel shortcuts can further enhance the diversity of features obtained by the self-attention network, thereby increasing the expressiveness of features.
[0020] In one possible implementation, the model includes a computer vision model or a natural language processing model.
[0021] The second aspect of the present application provides a feature extraction method, including: obtaining data to be processed; inputting the data to be processed into one or more serially connected self-attention networks to obtain features of the data to be processed; wherein the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes multiple parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the multiple parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, the multi-layer perceptron includes multiple serial first fully connected layers, the first neural network layer is parallelly connected to the self-attention module and one or more of the multi-layer perceptrons, and the first neural network layer is used to perform feature transformation.
[0022] In one possible implementation, the self-attention network also includes: a second neural network layer, which is used to perform feature transformation; the first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron, or the second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.
[0023] In one possible implementation, the first neural network layer includes a weight matrix and an activation function, the weight matrix is used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.
[0024] In a possible implementation manner, the weight matrix includes multiple sub-matrices, and each sub-matrix is a circulant matrix.
[0025] In a possible implementation manner, the multiple parallel feature extraction layers respectively include different weight matrices.
[0026] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer, the second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.
[0027] In a possible implementation, the self-attention module and / or the multilayer perceptron also has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.
[0028] In a possible implementation, the method is applied to a computer vision task or a natural language processing task.
[0029] The third aspect of the present application provides an image processing method, including: obtaining an image to be processed; inputting the image to be processed into an image processing model to extract image features through an attention model in the image processing model, wherein the attention model is the attention model described in the first aspect or any implementation method of the first aspect; processing the image to be processed according to the image features.
[0030] In a possible implementation, the processing the image to be processed according to the image features includes: performing one or more of the following tasks on the image to be processed according to the image features: image recognition, target detection, semantic segmentation, and image generation.
[0031] The fourth aspect of the present application provides a natural language processing method, including: obtaining a text to be processed; inputting the text to be processed into a natural language processing model to extract text features through an attention model in the natural language processing model, wherein the attention model is the attention model described in the first aspect or any implementation method of the first aspect; processing the text to be processed according to the text features.
[0032] In a possible implementation, the processing of the text to be processed according to the text features includes: performing one or more of the following tasks on the text to be processed according to the text features: machine translation, public opinion monitoring, automatic summary generation, opinion extraction, text classification, question answering, and text semantic comparison.
[0033] In a fifth aspect, the present application provides a feature extraction device, comprising: an acquisition unit and a processing unit; the acquisition unit is used to acquire data to be processed; the processing unit is used to input the data to be processed into one or more serially connected self-attention networks to obtain the features of the data to be processed; wherein the self-attention network comprises a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module comprises a plurality of parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the plurality of parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, the multi-layer perceptron comprises a plurality of serial first fully connected layers, the first neural network layer is parallelly connected to the self-attention module and one or more of the multi-layer perceptrons, and the first neural network layer is used to perform feature transformation.
[0034] In a possible implementation, the self-attention network further includes: a second neural network layer, the second neural network layer being used to perform feature transformation;
[0035] The first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron,
[0036] or,
[0037] The second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.
[0038] In one possible implementation, the first neural network layer includes a weight matrix and an activation function, the weight matrix is used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.
[0039] In a possible implementation manner, the weight matrix includes multiple sub-matrices, and each sub-matrix is a circulant matrix.
[0040] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer, the second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.
[0041] In a possible implementation, the self-attention module and / or the multilayer perceptron also has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.
[0042] The sixth aspect of the present application provides an image processing device, including: an acquisition unit and a processing unit; the acquisition unit is used to acquire an image to be processed; the processing unit is used to input the image to be processed into an image processing model to extract image features through an attention model in the image processing model, and the attention model is the attention model described in the first aspect or any implementation method of the first aspect; the processing unit is also used to process the image to be processed according to the image features.
[0043] In a possible implementation, the processing unit is further configured to perform one or more of the following tasks on the image to be processed according to the image features: image recognition, target detection, semantic segmentation, and image generation.
[0044] The seventh aspect of the present application provides a natural language processing device, including: an acquisition unit and a processing unit; the acquisition unit is used to acquire a text to be processed; the processing unit is used to input the text to be processed into a natural language processing model to extract text features through an attention model in the natural language processing model, and the attention model is the attention model described in the first aspect or any implementation method of the first aspect; the processing unit is also used to process the text to be processed according to the text features.
[0045] In one possible implementation, the processing unit is also used to perform one or more of the following tasks on the text to be processed according to the text features: machine translation, public opinion monitoring, automatic summary generation, opinion extraction, text classification, question answering, and text semantic comparison.
[0046] In an eighth aspect, the present application provides an electronic device, which may include a processor, the processor and a memory are coupled, the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method described in the first aspect or the second aspect is implemented. For the steps in each possible implementation method of the processor executing the second aspect, the details can be referred to the second aspect, and no further description is given here.
[0047] In a ninth aspect of the present application, a server is provided, which may include a processor, the processor is coupled to a memory, the memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method described in the second aspect is implemented. For the steps in each possible implementation method of the processor executing the second aspect, the details can be referred to the second aspect, and will not be repeated here.
[0048] In a tenth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method described in the second aspect.
[0049] In an eleventh aspect, the present application provides a circuit system, which includes a processing circuit, and the processing circuit is configured to execute the method described in the second aspect.
[0050] A twelfth aspect of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method described in the second aspect.
[0051] The thirteenth aspect of the present application provides a chip system, which includes a processor for supporting a server or a threshold value acquisition device to implement the functions involved in the first aspect, for example, sending or processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the server or communication device. The chip system can be composed of chips, or it can include chips and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 A structural diagram of the main framework of artificial intelligence;
[0053] Figure 2 A schematic diagram of a convolutional neural network provided in an embodiment of the present application;
[0054] Figure 3 A schematic diagram of a convolutional neural network provided in an embodiment of the present application;
[0055] Figure 4 A schematic diagram of a system architecture provided for an embodiment of the present application;
[0056] Figure 5 A schematic diagram of the structure of an attention model provided in an embodiment of the present application;
[0057] Figure 6a A schematic diagram of the structure of the attention model provided in the embodiment of the present application;
[0058] Figure 6b Another structural schematic diagram of the attention model provided in the embodiment of the present application;
[0059] Figure 6c Another structural schematic diagram of the attention model provided in the embodiment of the present application;
[0060] Figure 6d Another structural schematic diagram of the attention model provided in the embodiment of the present application;
[0061] Figure 6e Another structural diagram of a self-attention network provided in an embodiment of the present application;
[0062] Figure 7 A schematic diagram of a feature processing provided in an embodiment of the present application;
[0063] Figure 8 A schematic diagram showing the performance comparison of different models provided in the embodiments of the present application on the Imagenet dataset;
[0064] Fig. 9 A schematic diagram of the structure of a feature extraction device provided in an embodiment of the present application;
[0065] Fig.10 A schematic diagram of the structure of an execution device provided in an embodiment of the present application;
[0066] Fig.11 A schematic diagram of the structure of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0068] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0069] First, the overall workflow of the artificial intelligence system is described. Figure 1 , Figure 1 The figure shows a structural diagram of the main framework of artificial intelligence. The following is an explanation of the above artificial intelligence theme framework from the two dimensions of "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, the data has undergone a condensation process of "data-information-knowledge-wisdom". The "IT value chain" reflects the value that artificial intelligence brings to the information technology industry from the underlying infrastructure of human intelligence, information (providing and processing technology implementation) to the industrial ecology process of the system.
[0070] (1) Infrastructure.
[0071] The infrastructure provides computing power support for the AI system, enables communication with the outside world, and supports it through the basic platform. It communicates with the outside world through sensors; computing power is provided by smart chips (CPU, NPU, GPU, ASIC, FPGA and other hardware acceleration chips); the basic platform includes distributed computing frameworks and networks and other related platform guarantees and support, which can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to obtain data, and these data are provided to the smart chips in the distributed computing system provided by the basic platform for calculation.
[0072] (2)Data.
[0073] The data on the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data of traditional devices, including business data of existing systems and perception data such as force, displacement, liquid level, temperature, and humidity.
[0074] (3) Data processing
[0075] Data processing usually includes data training, machine learning, deep learning, search, reasoning, decision-making and other methods.
[0076] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.
[0077] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.
[0078] Decision-making refers to the process of making decisions after intelligent information is reasoned, usually providing functions such as classification, sorting, and prediction.
[0079] (4) General ability.
[0080] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as an algorithm or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0081] (5) Smart products and industry applications.
[0082] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical applications. Its application areas mainly include: smart electronic devices, smart transportation, smart medical care, autonomous driving, smart cities, etc.
[0083] The following describes the method provided by this application from the perspectives of model training and model application:
[0084] The model training method provided in the embodiment of the present application can be specifically applied to data training, machine learning, deep learning and other data processing methods to perform symbolic and formalized intelligent information modeling, extraction, preprocessing, training, etc. on the training data, and finally obtain a trained neural network model (such as the target neural network model in the embodiment of the present application); and the target neural network model can be used for model reasoning, specifically, the input data can be input into the target neural network model to obtain output data.
[0085] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application are first introduced below.
[0086] (1) Neural network.
[0087] A neural network may be composed of neural units, and a neural unit may refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input, and the output of the operation unit may be:
[0088] Where s=1, 2, ...n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolution layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the characteristics of the local receptive field. The local receptive field can be an area composed of several neural units.
[0089] (2) Convolutional neural network (CNN) is a deep neural network with a convolutional structure. The convolutional neural network includes a feature extractor consisting of a convolution layer and a subsampling layer. The feature extractor can be regarded as a filter, and the convolution process can be regarded as using a trainable filter to convolve with an input image or convolution feature plane (feature map). The convolution layer refers to the neuron layer in the convolutional neural network that performs convolution processing on the input signal (for example, the first convolution layer and the second convolution layer in this embodiment). In the convolution layer of the convolutional neural network, a neuron can only be connected to some neurons in the adjacent layers. A convolutional layer usually contains several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights here are the convolution kernels. Shared weights can be understood as the way of extracting image information is independent of position. The implicit principle is that the statistical information of a part of the image is the same as that of other parts. This means that the image information learned in one part can also be used in another part. So for all positions on the image, we can use the same learned image information. In the same convolution layer, multiple convolution kernels can be used to extract different image information. Generally speaking, the more convolution kernels there are, the richer the image information reflected by the convolution operation.
[0090] The convolution kernel can be initialized in the form of a matrix of random size, and the convolution kernel can obtain reasonable weights through learning during the training process of the convolutional neural network. In addition, the direct benefit of shared weights is to reduce the connections between the layers of the convolutional neural network, while reducing the risk of overfitting.
[0091] Specifically, Figure 2 As shown, the convolutional neural network (CNN) 100 may include an input layer 110 , a convolutional layer / pooling layer 120 , wherein the pooling layer is optional, and a neural network layer 130 .
[0092] Among them, the structure composed of the convolution layer / pooling layer 120 and the neural network layer 130 can be the first convolution layer and the second convolution layer described in the present application, the input layer 110 is connected to the convolution layer / pooling layer 120, the convolution layer / pooling layer 120 is connected to the neural network layer 130, the output of the neural network layer 130 can be input to the activation layer, and the activation layer can perform nonlinear processing on the output of the neural network layer 130.
[0093] Convolutional layer / pooling layer 120. Convolutional layer: Figure 2 The convolution layer / pooling layer 120 shown may include layers 121-126 as shown in the examples. In one implementation, layer 121 is a convolution layer, layer 122 is a pooling layer, layer 123 is a convolution layer, layer 124 is a pooling layer, layer 125 is a convolution layer, and layer 126 is a pooling layer. In another implementation, layers 121 and 122 are convolution layers, layer 123 is a pooling layer, layers 124 and 125 are convolution layers, and layer 126 is a pooling layer. That is, the output of a convolution layer can be used as the input of a subsequent pooling layer, or as the input of another convolution layer to continue the convolution operation.
[0094] Taking the convolution layer 121 as an example, the convolution layer 121 may include a plurality of convolution operators, which are also called kernels. The convolution operator is equivalent to a filter that extracts specific information from the input image matrix in image processing. The convolution operator can essentially be a weight matrix, which is usually predefined. In the process of performing convolution operations on the image, the weight matrix is usually processed one pixel after another (or two pixels after two pixels... depending on the value of the step length stride) in the horizontal direction on the input image, thereby completing the work of extracting specific features from the image. The size of the weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix is the same as the depth dimension of the input image. In the process of performing convolution operations, the weight matrix will extend to the entire depth of the input image. Therefore, convolution with a single weight matrix will produce a convolution output with a single depth dimension, but in most cases, a single weight matrix is not used, but multiple weight matrices of the same dimension are applied. The output of each weight matrix is stacked to form the depth dimension of the convolved image. Different weight matrices can be used to extract different features in the image. For example, one weight matrix is used to extract image edge information, another weight matrix is used to extract specific colors of the image, and another weight matrix is used to blur unnecessary noise in the image... The multiple weight matrices have the same dimensions, and the feature maps extracted by the multiple weight matrices with the same dimensions are also of the same dimension. The extracted feature maps with the same dimensions are then merged to form the output of the convolution operation.
[0095] The weight values in these weight matrices need to be obtained through a lot of training in practical applications. The weight matrices formed by the weight values obtained through training can extract information from the input image, thereby helping the convolutional neural network 100 to make correct predictions.
[0096] When the convolutional neural network 100 has multiple convolutional layers, the initial convolutional layer (for example, 121) often extracts more general features, which can also be called low-level features. As the depth of the convolutional neural network 100 increases, the features extracted by the later convolutional layers (for example, 126) become more and more complex, such as high-level semantic features. Features with higher semantics are more suitable for the problem to be solved.
[0097] Pooling layer: Since it is often necessary to reduce the number of training parameters, it is often necessary to periodically introduce a pooling layer after the convolution layer, such as Figure 2 The layers 121-126 shown in 120 may be a convolution layer followed by a pooling layer, or multiple convolution layers may be followed by one or more pooling layers.
[0098] Neural network layer 130: After being processed by the convolution layer / pooling layer 120, the convolution neural network 100 is not sufficient to output the required output information. As mentioned above, the convolution layer / pooling layer 120 only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other related information), the convolution neural network 100 needs to use the neural network layer 130 to generate one or a group of outputs of the required number of classes. Therefore, the neural network layer 130 may include multiple hidden layers (such as Figure 2 131, 132 to 13n) and the output layer 140 shown, the parameters contained in the multi-layer hidden layers can be pre-trained according to relevant training data of a specific task type. For example, the task type may include image recognition, image classification, image super-resolution reconstruction, etc.
[0099] After the multiple hidden layers in the neural network layer 130, that is, the last layer of the entire convolutional neural network 100 is the output layer 140, which has a loss function similar to the classification cross entropy, specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network 100 (such as Figure 2 The propagation from 110 to 140 is forward propagation), and the reverse propagation (such as Figure 2 The propagation from 140 to 110 is called back propagation) and then the weight values and biases of the aforementioned layers will begin to be updated to reduce the loss of the convolutional neural network 100 and the error between the result output by the convolutional neural network 100 through the output layer and the ideal result.
[0100] It should be noted that if Figure 2 The convolutional neural network 100 shown is only an example of a convolutional neural network. In specific applications, the convolutional neural network can also exist in the form of other network models, for example, Figure 3 The multiple convolutional layers / pooling layers shown are operated in parallel, and the features extracted respectively are input to the full neural network layer 130 for processing.
[0101] (3) Deep neural networks.
[0102] Deep Neural Network (DNN), also known as multi-layer neural network, can be understood as a neural network with many hidden layers. There is no special metric for "many" here. From the position of different layers of DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since DNN has many layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscripts correspond to the third layer index 2 of the output and the second layer index 4 of the input.
[0103] In summary, the coefficients from the kth neuron in the L-1th layer to the jth neuron in the Lth layer are defined as
[0104] It should be noted that the input layer does not have a W parameter. In a deep neural network, more hidden layers allow the network to better describe complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrix of all layers of the trained deep neural network (the weight matrix formed by many layers of vector W).
[0105] (4) Loss function
[0106] In the process of training a deep neural network, because we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the target value we really want, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the deep neural network). For example, if the predicted value of the network is high, adjust the weight vector to make it predict a lower value, and keep adjusting until the deep neural network can predict the target value we really want or a value very close to the target value we really want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function, which are important equations used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0107] (5) Back propagation algorithm.
[0108] Convolutional neural networks can use the error back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during the training process, so that the reconstruction error loss of the super-resolution model becomes smaller and smaller. Specifically, the forward transmission of the input signal to the output will generate error loss, and the error loss information is back-propagated to update the parameters in the initial super-resolution model, so that the error loss converges. The back propagation algorithm is a back propagation movement dominated by error loss, aiming to obtain the optimal parameters of the super-resolution model, such as the weight matrix.
[0109] (6) Linear operation.
[0110] Linearity refers to the proportional, linear relationship between quantities. Mathematically, it can be understood as a function whose first-order derivative is a constant. Linear operations can be, but are not limited to, addition operations, null operations, identity operations, convolution operations, batch normalization (BN) operations, and pooling operations. Linear operations can also be called linear mappings. Linear mappings need to meet two conditions: homogeneity and additivity. If either condition is not met, it is nonlinear.
[0111] Among them, homogeneity means f(ax)=af(x); additivity means f(x+y)=f(x)+f(y); for example, f(x)=ax is linear. It should be noted that x, a, and f(x) here are not necessarily scalars, but can be vectors or matrices to form a linear space of any dimension. If x and f(x) are n-dimensional vectors, when a is a constant, they are equivalent to satisfying homogeneity, and when a is a matrix, they are equivalent to satisfying additivity. Relatively speaking, a function graph that is a straight line does not necessarily conform to a linear mapping. For example, f(x)=ax+b neither satisfies homogeneity nor additivity, so it belongs to a nonlinear mapping.
[0112] In the embodiment of the present application, the combination of multiple linear operations may be referred to as a linear operation, and each linear operation included in a linear operation may also be referred to as a sub-linear operation.
[0113] (7) Attention model.
[0114] An attention model is a neural network that uses an attention mechanism. In deep learning, the attention mechanism can be broadly defined as a weight vector that describes importance: this weight vector is used to predict or infer an element. For example, for a pixel in an image or a word in a sentence, the attention vector can be used to quantitatively estimate the correlation between the target element and other elements, and the weighted sum of the attention vectors is used as an approximation of the target.
[0115] The attention mechanism in deep learning simulates the attention mechanism of the human brain. For example, when humans look at a painting, although the human eye can see the whole picture of the whole painting, when humans observe it in depth and carefully, the eyes actually focus on only a part of the pattern in the whole painting. At this time, the human brain mainly focuses on this small pattern. In other words, when humans observe an image carefully, the human brain does not pay equal attention to the whole image, but has a certain weight distinction. This is the core idea of the attention mechanism.
[0116] Simply put, the human visual processing system tends to selectively focus on certain parts of the image and ignore other irrelevant information, which helps the human brain perceive. Similarly, in the attention mechanism of deep learning, in some problems involving language, speech or vision, some parts of the input may be more relevant than other parts. Therefore, through the attention mechanism in the attention model, the attention model can dynamically focus on only the part of the input that helps to effectively perform the task at hand.
[0117] (8) Self-attention network.
[0118] The self-attention network is a neural network that uses the self-attention mechanism. The self-attention mechanism is an extension of the attention mechanism. The self-attention mechanism is actually an attention mechanism that associates different positions of a single sequence to calculate the representation of the same sequence. The self-attention mechanism can play a key role in machine reading, abstract summarization, or image description generation.
[0119] Taking the application of self-attention network in natural language processing as an example, the self-attention network processes input data of arbitrary length and generates a new feature expression of the input data, and then converts the feature expression into the target word. The self-attention network layer in the self-attention network uses the attention mechanism to obtain the relationship between all other words, thereby generating a new feature expression for each word. The advantage of the self-attention network is that the attention mechanism can directly capture the relationship between all words in a sentence without considering the position of the words.
[0120] Figure 4 is a schematic diagram of a system architecture provided by an embodiment of the present application. Figure 4 In the embodiment, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with an external device. A user can input data to the I / O interface 112 through a client device 140 .
[0121] When the execution device 120 preprocesses the input data, or when the computing module 111 of the execution device 120 performs calculation and other related processing (such as implementing the function of the neural network in the present application), the execution device 120 can call the data, code, etc. in the data storage system 150 for the corresponding processing, and can also store the data, instructions, etc. obtained by the corresponding processing in the data storage system 150.
[0122] Finally, the I / O interface 112 returns the processing result to the client device 140 so as to provide it to the user.
[0123] Optionally, the client device 140, for example, may be a control unit in an autonomous driving system, or a functional algorithm module in a mobile electronic device. For example, the functional algorithm module may be used to implement related tasks.
[0124] It is worth noting that the training device 120 can generate corresponding target models / rules (such as the target neural network model in this embodiment) based on different training data for different goals or different tasks. The corresponding target models / rules can be used to achieve the above goals or complete the above tasks, thereby providing users with the desired results.
[0125] exist Figure 4In the case shown in the figure, the user can manually give input data, and the manual giving can be operated through the interface provided by the I / O interface 112. In another case, the client device 140 can automatically send input data to the I / O interface 112. If the client device 140 is required to automatically send input data and needs to obtain the user's authorization, the user can set the corresponding authority in the client device 140. The user can view the results output by the execution device 110 on the client device 140, and the specific form can be display, sound, action and other specific methods. The client device 140 can also be used as a data acquisition terminal to collect the input data of the input I / O interface 112 and the output results of the output I / O interface 112 as new sample data, and store them in the database 130. Of course, it is also possible not to collect through the client device 140, but the I / O interface 112 directly stores the input data of the input I / O interface 112 and the output results of the output I / O interface 112 as new sample data in the database 130.
[0126] It is worth noting that Figure 4 This is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 4 In the embodiment, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 may also be placed in the execution device 110.
[0127] The attention model and feature extraction method provided in the embodiments of the present application can be applied to electronic devices, especially electronic devices that need to perform data processing tasks based on self-attention networks. Exemplarily, the electronic device can be, for example, a server, a smart phone, a personal computer (PC), a laptop, a tablet computer, a smart TV, a mobile Internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless electronic device in industrial control, a wireless electronic device in self-driving, a wireless electronic device in remote medical surgery, a wireless electronic device in smart grid, a wireless electronic device in transportation safety, a wireless electronic device in a smart city, a wireless electronic device in a smart home, etc.
[0128] The above introduces the devices to which the attention model and feature extraction method provided in the embodiments of the present application are applied. The following will introduce the scenarios to which the attention model and feature extraction method provided in the embodiments of the present application are applied.
[0129] The attention model and feature extraction method provided in the embodiment of the present application can be applied to computer vision or natural language processing. That is, the electronic device can perform computer vision tasks or natural language processing tasks through the above-mentioned attention model and feature extraction method.
[0130] Among them, natural language processing is an important direction in the fields of computer science and artificial intelligence. Natural language processing studies various theories and methods that can achieve effective communication between humans and computers using natural language. Generally speaking, natural language processing tasks mainly include machine translation, public opinion monitoring, automatic summary generation, opinion extraction, text classification, question answering, text semantic comparison, and speech recognition.
[0131] Computer vision is a science that studies how to make machines learn to see. To put it more specifically, computer vision refers to machine vision that uses cameras and computers to replace human eyes to identify, track, and measure targets, and further performs image processing to make the processed images more suitable for human eye observation or transmission to instruments for detection. Generally speaking, computer vision tasks include image classification, object detection, semantic segmentation, and image generation.
[0132] Image recognition is a common classification problem, also commonly referred to as image classification. Specifically, in the image recognition task, the input of the neural network is image data, and the output value is the probability that the current image data belongs to each category. Usually, the category with the largest probability value is selected as the predicted category of the image data. Image recognition is one of the earliest tasks that successfully applied deep learning. Classic network models include the VGG series, Inception series, and ResNet series.
[0133] Object detection refers to the automatic detection of the approximate location of common objects in an image through algorithms. A bounding box is usually used to represent the approximate location of the object and to classify the category information of the object in the bounding box.
[0134] Semantic segmentation refers to the automatic segmentation and recognition of the content in an image through an algorithm. Semantic segmentation can be understood as the classification problem of each pixel, that is, analyzing the category of the object to which each pixel belongs.
[0135] Image generation refers to obtaining a highly realistic generated image by learning the distribution of real images and sampling from the learned distribution. For example, a clear image is generated from a blurred image; a defogged image is generated from a foggy image.
[0136] The above introduces the scenarios in which the attention model and feature extraction method provided in the embodiments of the present application are applied. The following will introduce the specific structure of the model provided in the embodiments of the present application.
[0137] The attention model provided in the embodiment of the present application includes a self-attention network, or multiple self-attention networks connected in series. When the attention model includes multiple self-attention networks connected in series, the structure of each self-attention network in the attention model is the same, but the weight parameters in different self-attention networks may be different. Figure 5 As shown, Figure 5 A schematic diagram of the structure of an attention model provided in an embodiment of the present application. Figure 5 In the example, the attention model includes N self-attention networks connected in series, namely, self-attention network 1, self-attention network 2, ... self-attention network N. The structures of the N networks from self-attention network 1 to self-attention network N may be the same, but the weight parameters in self-attention network 1 to self-attention network N may be different. Among them, the input of self-attention network 1 is the input data of the attention model, the input of self-attention network 2 is the output of self-attention network 1, and the input of self-attention network N is the output of self-attention network N-1.
[0138] Specifically, in the attention model, each self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer. The self-attention module includes multiple parallel feature extraction layers and a fusion layer, the fusion layer is respectively connected to the multiple parallel feature extraction layers, and the fusion layer is used to fuse the features output by the multiple parallel feature extraction layers. Among them, the self-attention module is a network that adopts the self-attention mechanism, which can associate different positions of the input sequence to calculate the representation of the same sequence.
[0139] The multilayer perceptron (MLP) is serially connected to the self-attention module, and the multilayer perceptron includes a plurality of serial first fully connected layers. Specifically, the multilayer perceptron can also be called a fully connected neural network, and the multilayer perceptron includes an input layer, a hidden layer and an output layer, and the number of hidden layers can be one or more layers. Among them, the network layers in the multilayer perceptron are all fully connected layers. That is, the input layer and the hidden layer of the multilayer perceptron are fully connected, and the hidden layer and the output layer of the multilayer perceptron are also fully connected. Among them, the fully connected layer means that each neuron in the fully connected layer is connected to all neurons in the previous layer, so as to integrate the features extracted from the previous layer.
[0140] The first neural network layer is connected in parallel with the self-attention module and one or more of the multilayer perceptrons, and the first neural network layer is used to perform feature transformation.
[0141] In this embodiment, the input of the attention model is data in the form of a sequence, that is, the input data of the attention model is sequence data. For example, the input data of the attention model can be a sentence sequence consisting of multiple consecutive words; for another example, the input data of the attention model can be an image block sequence consisting of multiple consecutive image blocks, and the multiple consecutive image blocks are obtained by segmenting a complete image.
[0142] For ease of understanding, various implementation methods of the above-mentioned self-attention network will be described in detail below with reference to the accompanying drawings.
[0143] Implementation method 1, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, and the first neural network layer is connected in parallel to the self-attention module.
[0144] See also Figure 6a , Figure 6a A schematic diagram of the structure of the attention model provided in the embodiment of the present application. Figure 6a As shown, in the self-attention network, the self-attention module is connected to the multi-layer perceptron in series, and the first neural network layer is connected to the self-attention module in parallel. During the working process of the self-attention network, the self-attention module and the first neural network layer process the data to be processed input to the self-attention network in parallel, and the output of the self-attention module and the output of the first neural network layer are added to obtain the input of the multi-layer perceptron. Finally, after the output of the self-attention module and the output of the first neural network layer are added, they continue to be processed by the multi-layer perceptron to obtain the output data of the self-attention network.
[0145] It can be understood that the attention model includes multiple serial self-attention networks, and Figure 6aWhen the self-attention network shown is the first self-attention network in the attention model, Figure 6a The data to be processed that is input into the self-attention network is the original sequence data, such as the text data to be processed or the image data to be processed. Figure 6a When the self-attention network shown is not the first self-attention network in the attention model, Figure 6a The data to be processed that is input into the self-attention network is the feature data output by the previous self-attention network.
[0146] In this solution, a parallel first neural network layer is introduced on the basis of the self-attention module, and the parallel first neural network layer performs feature transformation operation on the input features to obtain transformed features. In addition, the transformed features are added to the output features of the self-attention module to increase the diversity of features output by the intermediate layer of the self-attention network, enhance the expressiveness of the features, and thus improve the performance of the attention model.
[0147] Implementation method 2, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, and the first neural network layer is connected in parallel to the multi-layer perceptron.
[0148] See also Figure 6b , Figure 6b Another structural diagram of the attention model provided in the embodiment of the present application. Figure 6b As shown, in the self-attention network, the self-attention module is connected to the multi-layer perceptron in series, and the first neural network layer is connected to the multi-layer perceptron in parallel. During the working process of the self-attention network, the self-attention module processes the data to be processed input into the self-attention network, and the output of the self-attention module is simultaneously used as the input of the multi-layer perceptron and the first neural network layer, and the multi-layer perceptron and the first neural network layer process the output of the self-attention module in parallel. Finally, after performing an addition operation on the output of the multi-layer perceptron and the output of the first neural network layer, the output data of the self-attention network is obtained.
[0149] In this scheme, a parallel first neural network layer is introduced on the basis of a multi-layer perceptron, and the parallel first neural network layer performs a feature transformation operation on the input features of the multi-layer perceptron to obtain the transformed features. In addition, the transformed features are added to the output features of the multi-layer perceptron to increase the diversity of the features output by the middle layer of the self-attention network, enhance the expressiveness of the features, and thus improve the performance of the attention model.
[0150] Implementation method 3, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, and the first neural network layer is parallel to the self-attention module and the multi-layer perceptron.
[0151] See also Figure 6c , Figure 6c Another structural diagram of the attention model provided in the embodiment of the present application. Figure 6c As shown, in the self-attention network, the self-attention module is connected to the multi-layer perceptron in series, and the first neural network layer is connected to the multi-layer perceptron in parallel. During the working process of the self-attention network, the self-attention module processes the data to be processed input into the self-attention network, and the output of the self-attention module is simultaneously used as the input of the multi-layer perceptron and the first neural network layer, and the multi-layer perceptron and the first neural network layer process the output of the self-attention module in parallel. Finally, after performing an addition operation on the output of the multi-layer perceptron and the output of the first neural network layer, the output data of the self-attention network is obtained.
[0152] In this solution, a parallel first neural network layer is introduced on the basis of the self-attention module and the multi-layer perceptron, and the parallel first neural network layer performs feature transformation operation on the input features of the self-attention network to obtain the transformed features. In addition, the transformed features are added to the output features of the multi-layer perceptron to increase the diversity of the features output by the intermediate layer of the self-attention network, enhance the expressiveness of the features, and thus improve the performance of the attention model.
[0153] Implementation method 4, the self-attention network includes a self-attention module, a multi-layer perceptron, a first neural network layer and a second neural network layer. The first neural network layer is parallel to the self-attention module, and the second neural network layer is parallel to the multi-layer perceptron; or, the second neural network layer is parallel to the self-attention module, and the first neural network layer is parallel to the multi-layer perceptron. Both the first neural network layer and the second neural network layer are used to perform feature transformation. In addition, the network structures of the first neural network layer and the second neural network layer may be the same, but the weight parameters in the first neural network layer and the second neural network layer may be different.
[0154] See also Figure 6d , Figure 6d Another structural diagram of the attention model provided in the embodiment of the present application. Figure 6dAs shown, in the self-attention network, the self-attention module is connected to the multi-layer perceptron in series, and the first neural network layer is connected to the self-attention module in parallel, and the second neural network layer is connected to the multi-layer perceptron in parallel. During the working process of the self-attention network, the self-attention module and the first neural network layer process the data to be processed input to the self-attention network, and the result obtained by adding the output of the self-attention module and the output of the first neural network layer is used as the input of the multi-layer perceptron and the second neural network layer. That is, the result obtained by adding the output of the self-attention module and the output of the first neural network layer is simultaneously used as the input of the multi-layer perceptron and the second neural network layer, and is processed in parallel by the multi-layer perceptron and the second neural network layer. Finally, after performing the addition operation on the output of the multi-layer perceptron and the output of the second neural network layer, the output data of the self-attention network is obtained.
[0155] In this scheme, by introducing the first neural network layer and the second neural network layer which are parallel to the self-attention module and the multi-layer perceptron respectively, the two parallel neural network layers perform feature transformation operations on the input features to obtain transformed features. In addition, the transformed features are added to the output features of the self-attention module and the output features of the multi-layer perceptron respectively to increase the diversity of the features output by the intermediate layer of the self-attention network, enhance the expressiveness of the features, and thus improve the performance of the attention model.
[0156] Optionally, based on the aforementioned four implementation methods, in addition to introducing a neural network layer parallel to the self-attention module and / or the multi-layer perceptron, the self-attention network can also include multiple neural network layers, which are parallel to the aforementioned first neural network layer and / or second neural network layer.
[0157] For example, in the aforementioned implementation 1, in the self-attention network, in addition to the first neural network layer connected in parallel with the self-attention module, multiple neural network layers may also be included. The multiple neural network layers are all connected in parallel with the first neural network layer, that is, the self-attention module is simultaneously connected in parallel with the first neural network layer and the multiple neural network layers. Among them, the multiple neural network layers are all used to perform feature transformation. In addition, the network structure of the multiple neural network layers may be the same as the network structure of the first neural network layer, but the weight parameters in the first neural network layer and the multiple neural network layers may be different.
[0158] Taking the above implementation method 4 as an example, please refer to Figure 6e , Figure 6e Another structural diagram of the self-attention network provided in the embodiment of the present application. Figure 6eAs shown, the self-attention network includes a self-attention module and a multi-layer perceptron connected in series, as well as n neural network layers connected in parallel to the self-attention module and n neural network layers connected in parallel to the multi-layer perceptron. Specifically, in the self-attention network, the self-attention module is connected in series to the multi-layer perceptron. The self-attention module has n neural network layers in parallel, namely neural network layer A1, neural network layer A2...neural network layer An. The multi-layer perceptron also has n neural network layers in parallel, namely neural network layer B1, neural network layer B2...neural network layer Bn.
[0159] In this scheme, by introducing multiple parallel neural network layers, the diversity of features output by the self-attention network can be further enriched.
[0160] In a possible embodiment, based on the above four implementations, the self-attention module in the self-attention network and / or the multi-layer perceptron also has a shortcut in parallel, and the input feature and output feature of the shortcut are the same. That is, the self-attention module and / or the multi-layer perceptron simultaneously have a shortcut of the first neural network layer and the identity mapping in parallel.
[0161] For example, taking the above implementation method 1 as an example, please refer to Figure 7 , Figure 7 A schematic diagram of a feature processing provided in an embodiment of the present application. Figure 7 As shown in the figure, the self-attention module in the self-attention network has a shortcut and a first neural network layer in parallel. The self-attention module, shortcut and first neural network layer process the same input features respectively. Then, the output features of the self-attention module, the output features of the shortcut and the output features of the first neural network layer are added to obtain the feature addition result, which is used as the input of the subsequent multi-layer perceptron. And, by Figure 7 It can be seen that the input features and output features of the shortcut are the same. After the feature transformation of the first neural network layer, the input features and output features of the first neural network layer are different.
[0162] Compared with the first neural network layer, shortcuts do not change the input features, that is, they do not perform feature transformation on the input features. Therefore, shortcuts can also be regarded as a special feature processing method. On the basis of the first neural network layer, the introduction of parallel shortcuts can further enhance the diversity of features obtained by the self-attention network, thereby increasing the expressiveness of features.
[0163] The above introduces various implementation methods of self-attention network. The following will introduce the specific modules in the self-attention network in detail.
[0164] In a possible embodiment, the first neural network layer includes a weight matrix and an activation function, the weight matrix is used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix. That is to say, in the process of processing the input features through the first neural network layer, the input features are first multiplied by the weight matrix in the first neural network layer to obtain the multiplication result; then, the multiplication result is processed by the activation function in the first neural network layer to obtain the output features of the first neural network layer.
[0165] Exemplarily, the process of processing the input features through the first neural network layer to obtain output features can be expressed by Formula 1.
[0166]
[0167] in, Represents the output features of the first neural network layer; Z l Represents the input features of the first neural network layer; Θ li represents the weight matrix; σ represents the activation function; i represents the number of blocks of the weight matrix, and the weight matrix can be divided into multiple sub-matrices to perform multiplication operations with the input features; i∈{1, 2, …, T].
[0168] Among them, the activation function refers to the function running on the neurons of the neural network model, which is responsible for mapping the input of the neuron to the output. The activation function plays a very important role in the neural network model to learn and understand very complex and nonlinear functions. The activation function can introduce nonlinear characteristics into the neural network model. For example, in the neuron, after the input features are weighted and summed through the weight matrix, a function is also applied, which is the activation function. The activation function is introduced to increase the nonlinearity of the neural network model. In the neural network model, each network layer without an activation function is equivalent to matrix multiplication. Therefore, in the absence of an activation function, the output obtained after superimposing several network layers is actually the result of a matrix multiplication.
[0169] In simple terms, if the activation function is not used in the neural network, the output of each neural network layer is a linear function of the upper layer input. No matter how many layers the neural network has, the output is a linear combination of the input. If the activation function is used in the neural network, the activation function introduces nonlinear factors to the neurons, allowing the neural network to arbitrarily approximate any nonlinear function, so that the neural network can be applied to many nonlinear models.
[0170] Exemplarily, the activation function may be a nonlinear function such as a Sigmoid function, a Tanh function, or a ReLU function. This embodiment does not specifically limit the specific type of the activation function.
[0171] It can be understood that the second neural network layer has the same structure as the first neural network layer, that is, the second neural network layer may also include a weight matrix and an activation function.
[0172] Specifically, when the self-attention module in the self-attention network has the first neural network layer and the shortcut in parallel, it is necessary to add the output features of the self-attention module, the output features of the first neural network layer and the output features of the shortcut to obtain the added output features. The added output features can be used as the input of the subsequent multi-layer perceptron.
[0173] Exemplarily, the process of adding the output features of the self-attention module, the output features of the first neural network layer, and the output features of the shortcut to obtain the added output features can be expressed by Formula 2.
[0174]
[0175] Among them, AugMSA(Z l ) represents the output feature after addition; Z l Represents input features; MSA(Z l ) represents the features obtained based on the self-attention module; It represents the i-th neural network layer parallel to the self-attention module. There are T neural network layers parallel to the self-attention module in total. The first neural network layer is one of the neural network layers parallel to the self-attention module.
[0176] It is understandable that after the introduction of the parallel first neural network layer, the process of processing the features through the parallel first neural network layer will involve matrix multiplication operations. If the matrix multiplication operation is performed directly on the input features and the weight matrix, it will bring huge computational overhead. Therefore, in this embodiment, it is proposed to use a block circulant matrix to implement the calculation of the first neural network layer, thereby reducing the computational overhead caused by the introduction of the parallel first neural network layer.
[0177] Specifically, in the first neural network layer in parallel with the self-attention module and / or the multi-layer perceptron, the weight matrix of the first neural network layer includes multiple sub-matrices, and each sub-matrix is a circulant matrix. In simple terms, the weight matrix can be composed of multiple sub-matrices, and each sub-matrix constituting the weight matrix is a circulant matrix.
[0178] Among them, the circulant matrix is a special form of matrix. Each element of the row vector of the circulant matrix is the result of shifting each element in the previous row vector right by one position. The circulant matrix can be efficiently calculated through Fourier transform, thereby saving the computational overhead of performing feature transformation on features through parallel neural network layers.
[0179] Since each submatrix is a circulant matrix, in practical applications, each submatrix can be represented by a set of vectors. For any set of vectors, the circulant matrix corresponding to the set of vectors can be obtained by right-shifting the elements in the set of vectors in sequence to obtain each row in the circulant matrix. In this way, in the process of training the self-attention network, when adjusting the weight parameters of the weight matrix in the self-attention network based on the loss function, it is actually adjusting the vector corresponding to the submatrix in the weight matrix. After the vector corresponding to the submatrix is adjusted, the adjusted weight matrix can be obtained based on the adjusted vector.
[0180] Exemplarily, the weight matrix may be represented by the following formula 3.
[0181]
[0182] Among them, Θ represents the weight matrix; the weight matrix Θ is divided into b*b sub-matrices in total, and each sub-matrix is a circulant matrix.
[0183] For any submatrix C in the weight matrix ij , sub-matrix C ij It can be obtained by looping over the elements in a vector. Specifically, the submatrix C ij It can be expressed based on the following formula 4.
[0184]
[0185] From formula 4, we can know that the submatrix C ij It can be achieved by looping through the vector [C1 ij , C2 ij , …, C d-1 ij , C d ij ] are obtained.
[0186] Exemplarily, when the weight matrix of the first neural network layer is divided into multiple sub-matrices and each sub-matrix is a circulant matrix, multiplying the input features of the first neural network layer by the weight matrix in the first neural network layer may specifically include the following steps.
[0187] First, the electronic device performs Fourier transform on multiple sub-matrices of the first neural network layer and input features of the first neural network layer, respectively, to obtain multiple transformed sub-matrices and transformed features.
[0188] Then, the electronic device performs element-by-element multiplication operations on the transformed features and the multiple transformed sub-matrices to obtain multiple sub-features, wherein the element-by-element multiplication operation refers to multiplying the elements at the same position in the two matrices one by one.
[0189] Finally, the electronic device performs inverse Fourier transform on the multiple sub-features, and accumulates the multiple sub-features after the inverse Fourier transform to obtain the processed first feature.
[0190] In this scheme, based on the fact that the weight matrix can be divided into multiple circulant matrices, Fourier transform is used to achieve efficient calculation between input features and weight features, thereby saving the computational overhead of performing feature transformation on features through parallel network layers.
[0191] Furthermore, in order to facilitate parallel processing of the input features of the first neural network layer on the same hardware, the electronic device may also divide the input features into multiple feature matrices, and perform the above processing on each feature matrix based on multiple sub-matrices. Finally, the obtained multiple output features are concatenated to obtain features processed based on the weight matrix in the first neural network layer.
[0192] Exemplarily, the process of processing the input features based on the weight matrix of the first neural network layer can be shown as Formula 5 and Formula 6.
[0193]
[0194]
[0195] in, Represents the features processed based on the weight matrix in the first neural network layer; represents some features after processing based on the weight matrix in the first neural network layer; Z represents the input features, Z is a matrix of size N*d, N represents the number of image blocks, and d represents the dimension of each image block; Z j represents the j-th feature matrix of the input matrix Z, with a size of N*(d / b), where b is the number of feature matrices divided by the input features; c ij represents a vector of length d / b; FFT represents fast Fourier transform; IFFT represents inverse fast Fourier transform; ○ represents element-by-element multiplication.
[0196] The above introduces the process of processing features by the first neural network layer provided in the embodiment of the present application. The following will describe in detail the process of processing features by the self-attention module provided in the embodiment of the present application.
[0197] In a possible embodiment, the self-attention module includes multiple parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the multiple parallel feature extraction layers, and the multiple parallel feature extraction layers respectively include different weight matrices. Specifically, for the input features of the self-attention module, the input features are processed by the multiple parallel feature extraction layers to obtain multiple extracted features; then, the fusion layer in the self-attention module fuses the multiple extracted features to obtain the output features of the self-attention module.
[0198] Exemplarily, the self-attention module may include three parallel feature extraction layers, each of which includes a weight matrix W Q , the weight matrix W K and the weight matrix W V For the input feature of the self-attention module, the input feature includes multiple sub-features. Each sub-feature in the input feature can be firstly compared with the weight matrix W Q , the weight matrix W K and the weight matrix W V Multiply them to get feature Q, feature K and feature V. Among them, feature Q, feature K and feature V are the inputs of the fusion layer of the self-attention module, and are further processed by the fusion layer.
[0199] Specifically, the fusion layer performs a dot product operation on feature Q and feature K to obtain the feature value score. Then, the fusion layer processes the feature value score through the softmax activation function to obtain the feature value softmax. The fusion layer then performs a dot product operation on the feature value softmax to obtain the score value v corresponding to each sub-feature. Finally, the score value v corresponding to each sub-feature is added to obtain the output feature of the self-attention module.
[0200] Optionally, the self-attention module can be a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer, the second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.
[0201] For example, assuming that the multi-head self-attention module includes 8 parallel self-attention units, the input features of the multi-head self-attention module can be input into the 8 parallel self-attention units respectively to obtain the feature matrix Z output by the 8 parallel self-attention units respectively. i, i∈{1, 2, ..., 8}. Then, the 8 feature matrices Z i The columns are spliced into a total feature matrix, and the total feature matrix is processed by the second fully connected layer to obtain the output features of the multi-head self-attention module.
[0202] In order to facilitate verification of the advantages of the attention model provided in the embodiment of the present application in data processing, the beneficial effects of the attention model provided in the embodiment of the present application will be explained based on specific experiments below.
[0203] In the embodiment of the present application, based on the existing network model and the attention model provided in the embodiment of the present application, extensive experiments are conducted on a large-scale dataset to conduct empirical research on the proposed method. The large-scale dataset may be the Imagenet (ILSVRC-2012) dataset, which contains 1.28 million training images and verification images from 1,000 categories.
[0204] See also Figure 8 , Figure 8 This is a schematic diagram showing the performance comparison of different models provided in the embodiments of this application on the Imagenet dataset. Figure 8 In the figure, Aug-ViT, Aug-PVT and Aug-T2T represent image processing models using the attention model provided in the embodiment of the present application. Resolution represents image resolution; Top-1 Accuracy represents accuracy; Params represents parameter quantity; FLOPs represents computational effort (number of multiplications and additions).
[0205] Depend on Figure 8 It can be seen that compared with the image processing model in the prior art, the image processing model using the attention model provided in the embodiment of the present application can achieve higher accuracy while keeping the amount of calculation and the amount of parameters basically unchanged.
[0206] An embodiment of the present application also provides a feature extraction method, including: an electronic device obtains data to be processed; the electronic device inputs the data to be processed into one or more serially connected self-attention networks to obtain features of the data to be processed; wherein the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes multiple parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the multiple parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, the multi-layer perceptron includes multiple serial first fully connected layers, the first neural network layer is parallelly connected to the self-attention module and one or more of the multi-layer perceptrons, and the first neural network layer is used to perform feature transformation.
[0207] In one possible implementation, the self-attention network also includes: a second neural network layer, which is used to perform feature transformation; the first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron, or the second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.
[0208] In one possible implementation, the first neural network layer includes a weight matrix and an activation function, the weight matrix is used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.
[0209] In a possible implementation manner, the weight matrix includes multiple sub-matrices, and each sub-matrix is a circulant matrix.
[0210] In a possible implementation manner, the multiple parallel feature extraction layers respectively include different weight matrices.
[0211] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer, the second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.
[0212] In a possible implementation, the self-attention module and / or the multilayer perceptron also has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.
[0213] In a possible implementation, the method is applied to a computer vision task or a natural language processing task.
[0214] An embodiment of the present application also provides an image processing method, including: obtaining an image to be processed; inputting the image to be processed into an image processing model to extract image features through an attention model in the image processing model, wherein the attention model is the attention model described in the aforementioned embodiment; and processing the image to be processed according to the image features.
[0215] In a possible implementation, the processing the image to be processed according to the image features includes: performing one or more of the following tasks on the image to be processed according to the image features: image recognition, target detection, semantic segmentation, and image generation.
[0216] An embodiment of the present application also provides a natural language processing method, including: obtaining a text to be processed; inputting the text to be processed into a natural language processing model to extract text features through an attention model in the natural language processing model, wherein the attention model is the attention model described in the aforementioned embodiment; and processing the text to be processed according to the text features.
[0217] In a possible implementation, the processing of the text to be processed according to the text features includes: performing one or more of the following tasks on the text to be processed according to the text features: machine translation, public opinion monitoring, automatic summary generation, opinion extraction, text classification, question answering, and text semantic comparison.
[0218] In a possible embodiment, the present application provides a feature extraction device. Fig. 9 , Fig. 9 A structural schematic diagram of a feature extraction device provided in an embodiment of the present application. The feature extraction device 900 includes: an acquisition unit 901 and a processing unit 902; the acquisition unit 901 is used to acquire data to be processed; the processing unit 902 is used to input the data to be processed into one or more self-attention networks connected in series to obtain the features of the data to be processed; wherein the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes a plurality of parallel feature extraction layers and a fusion layer, the fusion layer is respectively connected to the plurality of parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, the multi-layer perceptron includes a plurality of serial first fully connected layers, the first neural network layer is parallelly connected to the self-attention module and one or more of the multi-layer perceptrons, and the first neural network layer is used to perform feature transformation.
[0219] In a possible implementation, the self-attention network further includes: a second neural network layer, the second neural network layer being used to perform feature transformation;
[0220] The first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron, or,
[0221] The second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.
[0222] In one possible implementation, the first neural network layer includes a weight matrix and an activation function, the weight matrix is used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.
[0223] In a possible implementation manner, the weight matrix includes multiple sub-matrices, and each sub-matrix is a circulant matrix.
[0224] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer, the second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.
[0225] In a possible implementation, the self-attention module and / or the multilayer perceptron also has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.
[0226] In a possible embodiment, the embodiment of the present application also provides an image processing device, including: an acquisition unit and a processing unit; the acquisition unit is used to acquire the image to be processed; the processing unit is used to input the image to be processed into an image processing model to extract image features through an attention model in the image processing model, and the attention model is the attention model described in the first aspect or any implementation method of the first aspect; the processing unit is also used to process the image to be processed according to the image features.
[0227] In a possible implementation, the processing unit is further configured to perform one or more of the following tasks on the image to be processed according to the image features: image recognition, target detection, semantic segmentation, and image generation.
[0228] In a possible embodiment, the embodiment of the present application also provides a natural language processing device, including: an acquisition unit and a processing unit; the acquisition unit is used to acquire a text to be processed; the processing unit is used to input the text to be processed into a natural language processing model to extract text features through an attention model in the natural language processing model, and the attention model is the attention model described in the first aspect or any implementation method of the first aspect; the processing unit is also used to process the text to be processed according to the text features.
[0229] In one possible implementation, the processing unit is also used to perform one or more of the following tasks on the text to be processed according to the text features: machine translation, public opinion monitoring, automatic summary generation, opinion extraction, text classification, question answering, and text semantic comparison.
[0230] Next, an execution device provided by an embodiment of the present application is introduced. Fig.10 , Fig.10This is a schematic diagram of the structure of the execution device provided in the embodiment of the present application. The execution device 1000 can be specifically manifested as a mobile phone, a tablet, a laptop, a smart wearable device, a server, etc., which is not limited here. Among them, the execution device 1000 can be deployed with Fig.10 The data processing device described in the corresponding embodiment is used to implement Fig.10 The data processing function in the corresponding embodiment. Specifically, the execution device 1000 includes: a receiver 1001, a transmitter 1002, a processor 1003 and a memory 1004 (wherein the number of the processor 1003 in the execution device 1000 can be one or more, Fig.10 In the example of FIG. 1003 , the processor 1003 may include an application processor 10031 and a communication processor 10032. In some embodiments of the present application, the receiver 1001, the transmitter 1002, the processor 1003 and the memory 1004 may be connected via a bus or other means.
[0231] The memory 1004 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1003. A portion of the memory 1004 may also include a non-volatile random access memory (NVRAM). The memory 1004 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0232] The processor 1003 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are referred to as bus systems in the figure.
[0233] The method disclosed in the above embodiment of the present application can be applied to the processor 1003, or implemented by the processor 1003. The processor 1003 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit or software instructions in the processor 1003. The above processor 1003 can be a general processor, a digital signal processor (digital signal processing, DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (application specific integrated circuit, ASIC), a field programmable gate array (field-programmable gate array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The processor 1003 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor are combined to execute. The software module may be located in a storage medium mature in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 1004, and the processor 1003 reads the information in the memory 1004 and completes the steps of the above method in combination with its hardware.
[0234] The receiver 1001 can be used to receive input digital or character information and generate signal input related to the relevant settings and function control of the execution device. The transmitter 1002 can be used to output digital or character information through the first interface; the transmitter 1002 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1002 can also include a display device such as a display screen.
[0235] Also provided in an embodiment of the present application is a computer program product which, when executed on a computer, enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0236] A computer-readable storage medium is also provided in an embodiment of the present application, which stores a program for signal processing. When the computer-readable storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0237] The execution device, training device or electronic device provided in the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit, wherein the processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, a pin or a circuit, etc. The processing unit may execute the computer execution instructions stored in the storage unit so that the chip in the execution device executes the image processing method described in the above embodiment, or so that the chip in the training device executes the image processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit may also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0238] For details, please refer to Fig.11 , Fig.11 A schematic diagram of the structure of a chip provided in an embodiment of the present application, the chip can be expressed as a neural network processor NPU 1100, NPU 1100 is mounted on the host CPU (Host CPU) as a coprocessor, and the host CPU assigns tasks. The core part of the NPU is the operation circuit 1103, which is controlled by the controller 1104 to extract matrix data from the memory and perform multiplication operations.
[0239] In some implementations, the operation circuit 1103 includes multiple processing units (Process Engine, PE) inside. In some implementations, the operation circuit 1103 is a two-dimensional systolic array. The operation circuit 1103 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the operation circuit 1103 is a general-purpose matrix processor.
[0240] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The operation circuit takes the corresponding data of matrix B from the weight memory 1102 and caches it on each PE in the operation circuit. The operation circuit takes the matrix A data from the input memory 1101 and performs matrix operations with matrix B, and the partial results or final results of the matrix are stored in the accumulator 1108.
[0241] The unified memory 1106 is used to store input data and output data. The weight data is directly transferred to the weight memory 1102 through the direct memory access controller (DMAC) 1105. The input data is also transferred to the unified memory 1106 through the DMAC.
[0242] BIU stands for Bus Interface Unit, i.e., bus interface unit 1111 , which is used for interaction between AXI bus, DMAC and instruction fetch buffer (IFB) 1109 .
[0243] The bus interface unit 1111 (Bus Interface Unit, BIU for short) is used for the instruction fetch memory 1109 to obtain instructions from the external memory, and is also used for the storage unit access controller 1105 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0244] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 1106 or to transfer weight data to the weight memory 1102 or to transfer input data to the input memory 1101.
[0245] The vector calculation unit 1107 includes multiple operation processing units, and further processes the output of the operation circuit 1103 when necessary, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of feature planes, etc.
[0246] In some implementations, the vector calculation unit 1107 can store the processed output vector to the unified memory 1106. For example, the vector calculation unit 1107 can apply a linear function; or a nonlinear function to the output of the operation circuit 1103, such as linear interpolation of the feature plane extracted by the convolution layer, and then, for example, a vector of accumulated values to generate an activation value. In some implementations, the vector calculation unit 1107 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1103, for example, for use in a subsequent layer in a neural network.
[0247] An instruction fetch buffer 1109 connected to the controller 1104 is used to store instructions used by the controller 1104;
[0248] Unified memory 1106, input memory 1101, weight memory 1102 and instruction fetch memory 1109 are all on-chip memories. External memories are private to the NPU hardware architecture.
[0249] The processor mentioned in any of the above places may be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.
[0250] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.
[0251] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0252] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0253] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.
Claims
1. A feature extraction method, characterized in that: include: Acquire data to be processed, where the data to be processed is an image or text; Inputting the data to be processed into one or more serially connected self-attention networks to obtain features of the data to be processed; Among them, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes multiple parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the multiple parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, and the multi-layer perceptron is located after the self-attention module, the multi-layer perceptron includes multiple serial first fully connected layers, the first neural network layer is connected to the self-attention module in parallel, or the first neural network layer is connected to the multi-layer perceptron in parallel, or the first neural network layer is connected to the self-attention module and the multi-layer perceptron in parallel, and the first neural network layer is used to perform feature transformation.
2. The method according to claim 1, characterized in that The self-attention network further includes: a second neural network layer, the second neural network layer being used to perform feature transformation; The first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron, or, The second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.
3. The method according to claim 1 or 2, characterized in that: The first neural network layer includes a weight matrix and an activation function, the weight matrix is used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.
4. The method according to claim 3, characterized in that The weight matrix includes a plurality of sub-matrices, and each sub-matrix is a circulant matrix.
5. The method according to claim 1 or 2, characterized in that: The self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer. The second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.
6. The method according to claim 1 or 2, characterized in that: The self-attention module and / or the multi-layer perceptron also has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.
7. An image processing method, characterized in that: include: Get the image to be processed; Input the image to be processed into an image processing model to extract image features through an attention model in the image processing model, wherein the attention model includes one or more self-attention networks connected in series, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes a plurality of parallel feature extraction layers and a fusion layer, the fusion layer is respectively connected to the plurality of parallel feature extraction layers, the multi-layer perceptron is connected in series to the self-attention module, and the multi-layer perceptron is located after the self-attention module, the multi-layer perceptron includes a plurality of serial first fully connected layers, the first neural network layer is connected in parallel to the self-attention module, or the first neural network layer is connected in parallel to the multi-layer perceptron, or the first neural network layer is connected in parallel to the self-attention module and the multi-layer perceptron, and the first neural network layer is used to perform feature transformation; The image to be processed is processed according to the image features.
8. A natural language processing method, characterized in that: include: Get the text to be processed; Input the text to be processed into a natural language processing model to extract text features through an attention model in the natural language processing model, wherein the attention model includes one or more self-attention networks connected in series, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes a plurality of parallel feature extraction layers and a fusion layer, the fusion layer is respectively connected to the plurality of parallel feature extraction layers, the multi-layer perceptron is connected in series to the self-attention module, and the multi-layer perceptron is located after the self-attention module, the multi-layer perceptron includes a plurality of serial first fully connected layers, the first neural network layer is connected in parallel to the self-attention module, or the first neural network layer is connected in parallel to the multi-layer perceptron, or the first neural network layer is connected in parallel to the self-attention module and the multi-layer perceptron, and the first neural network layer is used to perform feature transformation; The text to be processed is processed according to the text features.
9. A feature extraction device, comprising: Acquisition unit and processing unit; The acquisition unit is used to acquire the data to be processed, wherein the data to be processed is an image or text; The processing unit is used to input the data to be processed into one or more self-attention networks connected in series to obtain features of the data to be processed; Among them, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes multiple parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the multiple parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, and the multi-layer perceptron is located after the self-attention module, the multi-layer perceptron includes multiple serial first fully connected layers, the first neural network layer is connected to the self-attention module in parallel, or the first neural network layer is connected to the multi-layer perceptron in parallel, or the first neural network layer is connected to the self-attention module and the multi-layer perceptron in parallel, and the first neural network layer is used to perform feature transformation.
10. The device according to claim 9, characterized in that The self-attention network further includes: a second neural network layer, the second neural network layer being used to perform feature transformation; The first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron, or, The second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.
11. The device according to claim 9 or 10, characterized in that The first neural network layer includes a weight matrix and an activation function, the weight matrix is used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.
12. The device according to claim 11, characterized in that The weight matrix includes a plurality of sub-matrices, and each sub-matrix is a circulant matrix.
13. The device according to claim 9 or 10, characterized in that The self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer. The second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.
14. The device according to claim 9 or 10, characterized in that The self-attention module and / or the multi-layer perceptron also has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.
15. An image processing device, characterized in that: include: Acquisition unit and processing unit; The acquisition unit is used to acquire the image to be processed; The processing unit is used to input the image to be processed into the image processing model to extract image features through the attention model in the image processing model, the attention model includes one or more self-attention networks connected in series, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes a plurality of parallel feature extraction layers and a fusion layer, the fusion layer is respectively connected to the plurality of parallel feature extraction layers, the multi-layer perceptron is connected in series with the self-attention module, and the multi-layer perceptron is located after the self-attention module, the multi-layer perceptron includes a plurality of serial first fully connected layers, the first neural network layer is connected in parallel with the self-attention module, or the first neural network layer is connected in parallel with the multi-layer perceptron, or the first neural network layer is connected in parallel with the self-attention module and the multi-layer perceptron, and the first neural network layer is used to perform feature transformation; The processing unit is further used to process the image to be processed according to the image features.
16. A natural language processing device, characterized in that: include: Acquisition unit and processing unit; The acquisition unit is used to acquire the text to be processed; The processing unit is used to input the text to be processed into a natural language processing model to extract text features through an attention model in the natural language processing model, wherein the attention model includes one or more self-attention networks connected in series, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes a plurality of parallel feature extraction layers and a fusion layer, the fusion layer is respectively connected to the plurality of parallel feature extraction layers, the multi-layer perceptron is connected in series to the self-attention module, and the multi-layer perceptron is located after the self-attention module, the multi-layer perceptron includes a plurality of serial first fully connected layers, the first neural network layer is connected in parallel to the self-attention module, or the first neural network layer is connected in parallel to the multi-layer perceptron, or the first neural network layer is connected in parallel to the self-attention module and the multi-layer perceptron, and the first neural network layer is used to perform feature transformation; The processing unit is further used to process the text to be processed according to the text features.
17. An electronic device, characterized in that: The electronic device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the electronic device executes the method according to any one of claims 1 to 8.
18. The electronic device according to claim 17, characterized in that: The electronic device includes a smart car, a smart phone, a smart TV, a virtual reality device, an augmented reality device, a wearable device or a server.
19. A computer storage medium, characterized in that: The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method of any one of claims 1 to 8.
20. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Intention recognition method and device based on ordering dialogue text and electronic device
CN109857844A
Non-uniform motion blurred image adaptive restoration method based on attention model
CN111275637A