Attention model, feature extraction method and related device

By introducing parallel neural network layers and connecting them with multi-layer perceptrons in the self-attention network, the feature collapse problem is solved, the feature expression ability is enhanced, and the performance of the attention model is improved.

CN120633642APending Publication Date: 2025-09-12HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510542048.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2021-06-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In self-attention networks, as the network depth increases, features tend to collapse, making them indistinguishable. Existing shortcut methods are unable to enhance feature expression capabilities, affecting model performance.

Method used

Parallel neural network layers are introduced to connect with self-attention modules and multi-layer perceptrons, and feature diversity is increased and expression capabilities are enhanced through feature transformation operations.

Benefits of technology

The feature expression ability and performance of the attention model are improved, the feature collapse problem is solved, and the processing effect of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633642A_ABST
    Figure CN120633642A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an attention model and a feature extraction method, and is applied to the technical field of artificial intelligence. The attention model comprises one or more self-attention networks which are connected in series, wherein each self-attention network comprises a self-attention module, a multi-layer perceptron and a first neural network layer; the self-attention module comprises a plurality of parallel feature extraction layers and a fusion layer, and the fusion layer is respectively connected with the plurality of parallel feature extraction layers; the multi-layer perceptron is in serial connection with the self-attention module, and the multi-layer perceptron comprises a plurality of serial first full connection layers; the first neural network layer is connected with one or more of the self-attention module and the multi-layer perceptron in parallel, and the first neural network layer is used for executing feature transformation. Based on the scheme, the diversity of the features extracted by the attention model can be increased, the expression ability of the features is enhanced, and thus the performance of the attention model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202110731775.6, and the original application date is June 29, 2021. The entire content of the original application is incorporated into this application by reference. Technical Field

[0002] The present application relates to the field of artificial intelligence technology, and in particular to an attention model, a feature extraction method, and related devices. Background Art

[0003] Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that seeks to understand the essence of intelligence and develop new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.

[0004] In recent years, self-attention networks have been widely used in many natural language processing (NLP) tasks, such as machine translation, sentiment analysis, and question answering. With the widespread application of self-attention networks, self-attention networks originating from the field of natural language processing have also achieved high performance in tasks such as image classification, object detection, and image processing.

[0005] In self-attention networks, due to the processing of features by the self-attention network layers, the features of the input data can easily become indistinguishable as the network deepens. These indistinguishable features have weak representational power. This phenomenon of features becoming indistinguishable as the network deepens is often called feature collapse.

[0006] Currently, adding shortcuts to self-attention networks can alleviate the phenomenon of feature collapse and prevent features from becoming indistinguishable. However, these shortcuts simply copy the input features of the self-attention network layer to the output layer, failing to enhance the expressive power of the features and resulting in poor performance of the self-attention network. Summary of the Invention

[0007] This application provides an attention model and a feature extraction method, which can increase the diversity of features extracted by the attention model, enhance the expressiveness of the features, and thus improve the performance of the attention model.

[0008] In a first aspect, the present application provides an attention model, comprising: one or more serially connected self-attention networks, wherein the self-attention network comprises a self-attention module, a multi-layer perceptron, and a first neural network layer.

[0009] The self-attention module includes multiple parallel feature extraction layers and a fusion layer, wherein the fusion layer is connected to the multiple parallel feature extraction layers. The self-attention module is a network that uses a self-attention mechanism to associate different positions of the input sequence to calculate a representation of the same sequence.

[0010] The multilayer perceptron is serially connected to the self-attention module, and the multilayer perceptron includes a plurality of serial first fully connected layers. Specifically, the multilayer perceptron can also be called a fully connected neural network (FCN), and the multilayer perceptron includes an input layer, a hidden layer, and an output layer, and the number of hidden layers can be one or more layers. Among them, the network layers in the multilayer perceptron are all fully connected layers. That is, the input layer and the hidden layer of the multilayer perceptron are fully connected, and the hidden layer and the output layer of the multilayer perceptron are also fully connected.

[0011] The first neural network layer is connected in parallel to the self-attention module and one or more of the multilayer perceptrons, wherein the first neural network layer is used to perform feature transformation.

[0012] This approach introduces a parallel neural network layer based on the self-attention module and multi-layer perceptron. This parallel neural network layer performs feature transformation on the input features to generate transformed features. These transformed features are then added to the output features of the self-attention module and / or multi-layer perceptron to increase the diversity of features output by the intermediate layers of the self-attention network, enhancing the expressiveness of the features and thus improving the performance of the attention model.

[0013] In one possible implementation, the self-attention network further includes: a second neural network layer, which is used to perform feature transformation; the first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multi-layer perceptron, or the second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multi-layer perceptron.

[0014] This approach introduces a first neural network layer and a second neural network layer, running in parallel with the self-attention module and the multi-layer perceptron, respectively. These two parallel neural network layers perform feature transformation operations on the input features to generate transformed features. These transformed features are then added to the output features of the self-attention module and the multi-layer perceptron, respectively, to increase the diversity of features output by the intermediate layers of the self-attention network, enhance the expressiveness of the features, and thus improve the performance of the attention model.

[0015] In one possible implementation, the first neural network layer includes a weight matrix and an activation function. The weight matrix is ​​used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix. The activation function can be, for example, a nonlinear function such as a Sigmoid function, a Tanh function, or a ReLU function.

[0016] In one possible implementation, the weight matrix includes a plurality of sub-matrices, each of which is a circulant matrix. In simple terms, the weight matrix may be composed of a plurality of sub-matrices, and each sub-matrix constituting the weight matrix is ​​a circulant matrix.

[0017] In a possible implementation, the multiple parallel feature extraction layers respectively include different weight matrices.

[0018] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer. The second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.

[0019] In a possible implementation, the self-attention module and / or the multilayer perceptron further has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.

[0020] Compared to the first neural network layer, shortcuts do not modify the input features, meaning they do not perform feature transformations on them. Therefore, shortcuts can be considered a specialized feature processing method. By introducing parallel shortcuts on top of the first neural network layer, we can further enhance the diversity of features obtained by the self-attention network, thereby increasing the expressive power of features.

[0021] In one possible implementation, the model includes a computer vision model or a natural language processing model.

[0022] The second aspect of the present application provides a feature extraction method, including: obtaining data to be processed; inputting the data to be processed into one or more serially connected self-attention networks to obtain features of the data to be processed; wherein the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes multiple parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the multiple parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, the multi-layer perceptron includes multiple serial first fully connected layers, the first neural network layer is connected in parallel to the self-attention module and one or more of the multi-layer perceptrons, and the first neural network layer is used to perform feature transformation.

[0023] In one possible implementation, the self-attention network further includes: a second neural network layer, which is used to perform feature transformation; the first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multi-layer perceptron, or the second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multi-layer perceptron.

[0024] In one possible implementation, the first neural network layer includes a weight matrix and an activation function, the weight matrix is ​​used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.

[0025] In a possible implementation, the weight matrix includes multiple sub-matrices, and each sub-matrix is ​​a circulant matrix.

[0026] In a possible implementation, the multiple parallel feature extraction layers respectively include different weight matrices.

[0027] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer. The second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.

[0028] In a possible implementation, the self-attention module and / or the multilayer perceptron further has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.

[0029] In one possible implementation, the method is applied to computer vision tasks or natural language processing tasks.

[0030] The third aspect of the present application provides an image processing method, including: obtaining an image to be processed; inputting the image to be processed into an image processing model to extract image features through an attention model in the image processing model, wherein the attention model is the attention model described in the first aspect or any implementation method of the first aspect; and processing the image to be processed according to the image features.

[0031] In a possible implementation, the processing the image to be processed according to the image features includes: performing one or more of the following tasks on the image to be processed according to the image features: image recognition, target detection, semantic segmentation, and image generation.

[0032] The fourth aspect of the present application provides a natural language processing method, including: obtaining a text to be processed; inputting the text to be processed into a natural language processing model to extract text features through an attention model in the natural language processing model, wherein the attention model is the attention model described in the first aspect or any implementation method of the first aspect; processing the text to be processed according to the text features.

[0033] In one possible implementation, the processing of the text to be processed according to the text features includes: performing one or more of the following tasks on the text to be processed according to the text features: machine translation, public opinion monitoring, automatic summary generation, opinion extraction, text classification, question answering, and text semantic comparison.

[0034] The fifth aspect of the present application provides a feature extraction device, including: an acquisition unit and a processing unit; the acquisition unit is used to acquire data to be processed; the processing unit is used to input the data to be processed into one or more serially connected self-attention networks to obtain the features of the data to be processed; wherein the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes multiple parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the multiple parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, the multi-layer perceptron includes multiple serial first fully connected layers, the first neural network layer is connected in parallel to the self-attention module and one or more of the multi-layer perceptrons, and the first neural network layer is used to perform feature transformation.

[0035] In one possible implementation, the self-attention network further includes: a second neural network layer, the second neural network layer being configured to perform feature transformation;

[0036] The first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron,

[0037] or,

[0038] The second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.

[0039] In one possible implementation, the first neural network layer includes a weight matrix and an activation function, the weight matrix is ​​used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.

[0040] In a possible implementation, the weight matrix includes multiple sub-matrices, and each sub-matrix is ​​a circulant matrix.

[0041] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer. The second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.

[0042] In a possible implementation, the self-attention module and / or the multilayer perceptron further has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.

[0043] In the sixth aspect of the present application, there is provided an image processing device, comprising: an acquisition unit and a processing unit; the acquisition unit is used to acquire an image to be processed; the processing unit is used to input the image to be processed into an image processing model to extract image features through an attention model in the image processing model, and the attention model is the attention model described in the first aspect or any implementation method of the first aspect; the processing unit is also used to process the image to be processed according to the image features.

[0044] In a possible implementation, the processing unit is further configured to perform one or more of the following tasks on the image to be processed according to the image features: image recognition, object detection, semantic segmentation, and image generation.

[0045] The seventh aspect of the present application provides a natural language processing device, including: an acquisition unit and a processing unit; the acquisition unit is used to acquire a text to be processed; the processing unit is used to input the text to be processed into a natural language processing model to extract text features through an attention model in the natural language processing model, and the attention model is the attention model described in the first aspect or any implementation method of the first aspect; the processing unit is also used to process the text to be processed according to the text features.

[0046] In one possible implementation, the processing unit is further used to perform one or more of the following tasks on the text to be processed based on the text features: machine translation, public opinion monitoring, automatic summary generation, opinion extraction, text classification, question answering, and text semantic comparison.

[0047] In an eighth aspect, the present application provides an electronic device that may include a processor coupled to a memory, the memory storing program instructions. When the program instructions stored in the memory are executed by the processor, the method described in the first or second aspect is implemented. For details of the steps in each possible implementation of the second aspect performed by the processor, please refer to the second aspect and will not be repeated here.

[0048] In a ninth aspect, the present application provides a server that may include a processor coupled to a memory, wherein the memory stores program instructions. When the program instructions stored in the memory are executed by the processor, the method described in the second aspect is implemented. For details of the steps in each possible implementation of the second aspect performed by the processor, please refer to the second aspect and will not be repeated here.

[0049] In a tenth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer executes the method described in the second aspect.

[0050] In an eleventh aspect, the present application provides a circuit system, which includes a processing circuit, and the processing circuit is configured to execute the method described in the second aspect.

[0051] A twelfth aspect of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method described in the second aspect above.

[0052] In a thirteenth aspect of the present application, a chip system is provided, which includes a processor for supporting a server or a threshold value acquisition device to implement the functions involved in the first aspect, for example, sending or processing the data and / or information involved in the above method. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the server or communication device. The chip system can be composed of a chip, or it can include a chip and other discrete devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A structural diagram of the main framework of artificial intelligence;

[0054] Figure 2 A schematic diagram of a convolutional neural network provided in an embodiment of the present application;

[0055] Figure 3 A schematic diagram of a convolutional neural network provided in an embodiment of the present application;

[0056] Figure 4 A schematic diagram of a system architecture provided in an embodiment of the present application;

[0057] Figure 5 A schematic diagram of the structure of an attention model provided in an embodiment of the present application;

[0058] Figure 6a A schematic diagram of the structure of the attention model provided in an embodiment of the present application;

[0059] Figure 6b Another structural diagram of the attention model provided in an embodiment of the present application;

[0060] Figure 6c Another structural diagram of the attention model provided in an embodiment of the present application;

[0061] Figure 6d Another structural diagram of the attention model provided in an embodiment of the present application;

[0062] Figure 6e Another structural diagram of the self-attention network provided in an embodiment of the present application;

[0063] Figure 7 A schematic diagram of a feature processing provided in an embodiment of the present application;

[0064] Figure 8 A schematic diagram showing the performance comparison of different models provided in the embodiments of this application on the Imagenet dataset;

[0065] Figure 9 A schematic structural diagram of a feature extraction device provided in an embodiment of the present application;

[0066] Figure 10 A schematic diagram of the structure of an execution device provided in an embodiment of the present application;

[0067] Figure 11 A schematic diagram of the structure of the chip provided in an embodiment of the present application. DETAILED DESCRIPTION

[0068] The embodiments of the present application are described below in conjunction with the accompanying drawings. Those skilled in the art will appreciate that, with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0069] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0070] First, the overall workflow of the artificial intelligence system is described. Figure 1 , Figure 1 The following diagram illustrates a structural diagram of the AI ​​framework. This framework is explained below from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it encompasses the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed progression from "data-information-knowledge-wisdom." The "IT value chain," encompassing the entire process from the underlying infrastructure of human intelligence, information (provided and processed by technology), to the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.

[0071] (1) Infrastructure.

[0072] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.

[0073] (2)Data.

[0074] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0075] (3)Data processing.

[0076] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.

[0077] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.

[0078] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.

[0079] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.

[0080] (4) General ability.

[0081] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0082] (5) Smart products and industry applications.

[0083] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart electronic devices, smart transportation, smart medical care, autonomous driving, smart cities, etc.

[0084] The following describes the method provided by this application from the perspectives of model training and model application:

[0085] The model training method provided in the embodiments of the present application can be specifically applied to data processing methods such as data training, machine learning, and deep learning, and symbolizes and formalizes intelligent information modeling, extraction, preprocessing, and training of the training data to ultimately obtain a trained neural network model (such as the target neural network model in the embodiments of the present application); and the target neural network model can be used for model inference, specifically, input data can be input into the target neural network model to obtain output data.

[0086] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application are first introduced below.

[0087] (1) Neural network.

[0088] A neural network can be composed of neural units. A neural unit can refer to an operation unit that takes xs (i.e., input data) and intercept 1 as input. The output of the operation unit can be:

[0089] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal of the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer, and the activation function can be a sigmoid function. A neural network is a network formed by connecting multiple single neural units mentioned above, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.

[0090] (2) Convolutional Neural Network (CNN) is a deep neural network with a convolutional structure. The convolutional neural network contains a feature extractor consisting of a convolution layer and a subsampling layer. The feature extractor can be regarded as a filter, and the convolution process can be regarded as using a trainable filter to convolve with an input image or convolution feature plane (feature map). The convolution layer refers to the neuron layer in the convolutional neural network that performs convolution processing on the input signal (for example, the first convolution layer and the second convolution layer in this embodiment). In the convolution layer of the convolutional neural network, a neuron can only be connected to some neurons in the adjacent layers. A convolution layer usually contains several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights here are the convolution kernels. Shared weights can be understood as the way of extracting image information is independent of position. The implicit principle is that the statistical information of a part of the image is the same as that of other parts. This means that the image information learned in a certain part can also be used in another part. Therefore, we can use the same learned image information for all positions on the image. In the same convolutional layer, multiple convolution kernels can be used to extract different image information. Generally speaking, the more convolution kernels there are, the richer the image information reflected by the convolution operation.

[0091] Convolution kernels can be initialized as matrices of random size, and during the training process of the convolutional neural network, the convolution kernels can be learned to obtain reasonable weights. In addition, the direct benefit of shared weights is that they reduce the number of connections between the layers of the convolutional neural network, while also reducing the risk of overfitting.

[0092] Specifically, such as Figure 2 As shown, a convolutional neural network (CNN) 100 may include an input layer 110 , a convolutional layer / pooling layer 120 , wherein the pooling layer is optional, and a neural network layer 130 .

[0093] Among them, the structure composed of the convolution layer / pooling layer 120 and the neural network layer 130 can be the first convolution layer and the second convolution layer described in this application, the input layer 110 is connected to the convolution layer / pooling layer 120, the convolution layer / pooling layer 120 is connected to the neural network layer 130, the output of the neural network layer 130 can be input to the activation layer, and the activation layer can perform nonlinear processing on the output of the neural network layer 130.

[0094] Convolution layer / pooling layer 120. Convolution layer: such as Figure 2 The convolutional layer / pooling layer 120 shown may include layers 121-126, for example. In one implementation, layer 121 is a convolutional layer, layer 122 is a pooling layer, layer 123 is a convolutional layer, layer 124 is a pooling layer, layer 125 is a convolutional layer, and layer 126 is a pooling layer. In another implementation, layers 121 and 122 are convolutional layers, layer 123 is a pooling layer, layers 124 and 125 are convolutional layers, and layer 126 is a pooling layer. That is, the output of a convolutional layer can be used as the input of a subsequent pooling layer, or as the input of another convolutional layer to continue the convolution operation.

[0095] Taking convolution layer 121 as an example, convolution layer 121 can include many convolution operators, also known as kernels. Their role in image processing is equivalent to a filter that extracts specific information from the input image matrix. The convolution operator can essentially be a weight matrix, which is usually predefined. During the convolution operation on the image, the weight matrix is ​​usually processed horizontally on the input image one pixel at a time (or two pixels at a time... depending on the value of the stride), thereby completing the task of extracting specific features from the image. The size of the weight matrix should be related to the size of the image. It is important to note that the depth dimension of the weight matrix is ​​the same as the depth dimension of the input image. During the convolution operation, the weight matrix extends to the entire depth of the input image. Therefore, convolution with a single weight matrix produces a convolution output with a single depth dimension. However, in most cases, a single weight matrix is ​​not used, but multiple weight matrices of the same dimension are applied. The output of each weight matrix is ​​stacked to form the depth dimension of the convolved image. Different weight matrices can be used to extract different features in the image. For example, one weight matrix is ​​used to extract image edge information, another weight matrix is ​​used to extract specific colors of the image, and another weight matrix is ​​used to blur unwanted noise in the image... The multiple weight matrices have the same dimensions, and the feature maps extracted by the multiple weight matrices with the same dimensions also have the same dimensions. The multiple feature maps with the same dimensions extracted are then merged to form the output of the convolution operation.

[0096] The weight values ​​in these weight matrices need to be obtained through a lot of training in practical applications. The weight matrices formed by the weight values ​​obtained through training can extract information from the input image, thereby helping the convolutional neural network 100 to make correct predictions.

[0097] When the convolutional neural network 100 has multiple convolutional layers, the initial convolutional layer (for example, 121) often extracts more general features, which can also be called low-level features. As the depth of the convolutional neural network 100 increases, the features extracted by the later convolutional layers (for example, 126) become more and more complex, such as high-level semantic features. Features with higher semantics are more suitable for the problem to be solved.

[0098] Pooling layer: Since it is often necessary to reduce the number of training parameters, it is often necessary to periodically introduce a pooling layer after the convolution layer, such as Figure 2 The layers 121-126 in the example 120 can be a convolution layer followed by a pooling layer, or multiple convolution layers can be followed by one or more pooling layers.

[0099] Neural network layer 130: After being processed by the convolution layer / pooling layer 120, the convolution neural network 100 is not sufficient to output the required output information. Because as mentioned above, the convolution layer / pooling layer 120 only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other related information), the convolution neural network 100 needs to use the neural network layer 130 to generate one or a group of outputs of the required number of classes. Therefore, the neural network layer 130 may include multiple hidden layers (such as Figure 2 131, 132 to 13n) and the output layer 140 shown, the parameters contained in the multi-layer hidden layer can be pre-trained according to relevant training data of a specific task type, for example, the task type may include image recognition, image classification, image super-resolution reconstruction, etc.

[0100] After the multiple hidden layers in the neural network layer 130, that is, the last layer of the entire convolutional neural network 100 is the output layer 140, which has a loss function similar to the classification cross entropy, specifically for calculating the prediction error. Once the forward propagation of the entire convolutional neural network 100 (such as Figure 2 The propagation from 110 to 140 is forward propagation), and the reverse propagation (such as Figure 2 The propagation from 140 to 110 is back propagation) and then starts to update the weight values ​​and biases of the aforementioned layers to reduce the loss of the convolutional neural network 100 and the error between the result output by the convolutional neural network 100 through the output layer and the ideal result.

[0101] It should be noted that if Figure 2 The convolutional neural network 100 shown is only an example of a convolutional neural network. In specific applications, the convolutional neural network can also exist in the form of other network models, such as Figure 3 The multiple convolutional layers / pooling layers shown are operated in parallel, and the features extracted from each layer are input to the full neural network layer 130 for processing.

[0102] (3) Deep neural networks.

[0103] Deep Neural Network (DNN), also known as multi-layer neural network, can be understood as a neural network with many hidden layers. There is no special metric for "many" here. Based on the position of different layers in DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the coefficient W and the offset vector The definition of these parameters in DNN is as follows: Take the coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscripts correspond to the third layer index 2 of the output and the second layer index 4 of the input.

[0104] In summary, the coefficient from the kth neuron in the L-1th layer to the jth neuron in the Lth layer is defined as

[0105] It's important to note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters has higher complexity and greater "capacity," meaning it can handle more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrices for all layers of a trained deep neural network (a weight matrix formed by the vectors W across many layers).

[0106] (4) Loss function.

[0107] During the training of a deep neural network, because we want the output of the deep neural network to be as close as possible to the desired predicted value, we can compare the current network's predicted value with the desired target value and then update the weight vector of each layer of the neural network based on the difference between the two. (Of course, before the first update, there is usually an initialization process, which pre-configures the parameters for each layer in the deep neural network.) For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value. This adjustment is continued until the deep neural network can predict the desired target value or a value very close to the desired target value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the target value." This is the loss function (or objective function), which is an important equation used to measure the difference between the predicted value and the target value. For example, the loss function output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss.

[0108] (5) Back propagation algorithm.

[0109] Convolutional neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during training, reducing the reconstruction error loss of the super-resolution model. Specifically, the forward propagation of the input signal to the output generates an error loss. This error loss information is then backpropagated to update the parameters of the initial super-resolution model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal super-resolution model parameters, such as the weight matrix.

[0110] (6) Linear operation.

[0111] Linearity refers to the proportional, linear relationship between quantities. Mathematically, it can be understood as a function whose first-order derivative is a constant. Linear operations include, but are not limited to, addition, null operations, identity operations, convolution, batch normalization (BN), and pooling. Linear operations, also known as linear mappings, must meet two conditions: homogeneity and additivity. Failure of either condition constitutes nonlinearity.

[0112] Here, homogeneity means f(ax) = af(x); additivity means f(x + y) = f(x) + f(y); for example, f(x) = ax is linear. It should be noted that x, a, and f(x) here are not necessarily scalars, but can be vectors or matrices, forming a linear space of any dimension. If x and f(x) are n-dimensional vectors, when a is a constant, they are equivalent to satisfying homogeneity, and when a is a matrix, they are equivalent to satisfying additivity. Relatively speaking, a function graph that is a straight line does not necessarily conform to a linear mapping. For example, f(x) = ax + b neither satisfies homogeneity nor additivity, and therefore belongs to a nonlinear mapping.

[0113] In the embodiment of the present application, the combination of multiple linear operations may be referred to as a linear operation, and each linear operation included in a linear operation may also be referred to as a sub-linear operation.

[0114] (7)Attention model.

[0115] An attention model is a neural network that uses the attention mechanism. In deep learning, the attention mechanism can be broadly defined as a weight vector that describes the importance of an element: this weight vector is used to predict or infer an element. For example, for a pixel in an image or a word in a sentence, the attention vector can be used to quantitatively estimate the correlation between the target element and other elements, and the weighted sum of the attention vectors can be used as an approximation of the target value.

[0116] The attention mechanism in deep learning simulates the attention mechanism of the human brain. For example, when a person looks at a painting, although their eyes can see the entire painting, when they look closely, their eyes actually focus on only a small part of the pattern. At this time, the human brain focuses primarily on this small part. In other words, when a person observes an image carefully, the human brain's attention is not evenly distributed across the entire image; instead, it is weighted differently. This is the core idea of ​​the attention mechanism.

[0117] Simply put, the human visual processing system tends to selectively focus on certain parts of an image while ignoring other irrelevant information, thereby facilitating human perception. Similarly, in deep learning's attention mechanism, in some problems involving language, speech, or vision, certain parts of the input may be more relevant than others. Therefore, through the attention mechanism in the attention model, the attention model can dynamically focus on only the parts of the input that help effectively perform the task at hand.

[0118] (8) Self-attention network.

[0119] The self-attention network is a neural network that uses the self-attention mechanism. The self-attention mechanism is an extension of the attention mechanism. It essentially associates different positions in a single sequence to compute a representation of the same sequence. The self-attention mechanism plays a key role in machine reading, abstractive summarization, and image description generation.

[0120] Taking the application of self-attention networks in natural language processing as an example, self-attention networks process input data of arbitrary length and generate new feature representations of the input data, which are then converted into target words. The self-attention network layer in the self-attention network uses the attention mechanism to capture the relationships between all other words, thereby generating a new feature representation for each word. The advantage of self-attention networks is that the attention mechanism can directly capture the relationships between all words in a sentence, regardless of their position.

[0121] Figure 4 This is a schematic diagram of a system architecture provided by an embodiment of the present application. Figure 4 In the embodiment, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with an external device. A user can input data to the I / O interface 112 through a client device 140 .

[0122] When the execution device 120 preprocesses the input data, or when the computing module 111 of the execution device 120 performs calculations and other related processing (such as implementing the functions of the neural network in this application), the execution device 120 can call the data, code, etc. in the data storage system 150 for the corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 150.

[0123] Finally, the I / O interface 112 returns the processing result to the client device 140 so as to provide it to the user.

[0124] Optionally, the client device 140 may be, for example, a control unit in an autonomous driving system or a functional algorithm module in a mobile electronic device, and the functional algorithm module may be used to implement related tasks.

[0125] It is worth noting that the training device 120 can generate corresponding target models / rules (such as the target neural network model in this embodiment) based on different training data for different goals or different tasks. The corresponding target models / rules can be used to achieve the above goals or complete the above tasks, thereby providing users with the desired results.

[0126] exist Figure 4In the case shown, the user can manually input data, which can be operated through the interface provided by I / O interface 112. In another case, client device 140 can automatically send input data to I / O interface 112. If the automatic transmission of input data by client device 140 requires user authorization, the user can set the corresponding permissions in client device 140. The user can view the results output by execution device 110 on client device 140, which can be displayed, sounded, or displayed in a specific form. Client device 140 can also serve as a data acquisition terminal, collecting input data input into I / O interface 112 and output results from I / O interface 112 as new sample data, and storing them in database 130. Of course, the collection can also be performed without going through client device 140, and instead the input data input into I / O interface 112 and output results from I / O interface 112 as new sample data can be directly stored in database 130 by I / O interface 112.

[0127] It is worth noting that Figure 4 This is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, Figure 4 In the embodiment, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110.

[0128] The attention model and feature extraction method provided in the embodiments of the present application can be applied to electronic devices, especially electronic devices that need to perform data processing tasks based on a self-attention network. For example, the electronic device can be a server, a mobile phone, a personal computer (PC), a laptop, a tablet computer, a smart TV, a mobile internet device (MID), a wearable device, a virtual reality (VR) device, an augmented reality (AR) device, a wireless electronic device in industrial control, a wireless electronic device in self-driving, a wireless electronic device in remote medical surgery, a wireless electronic device in a smart grid, a wireless electronic device in transportation safety, a wireless electronic device in a smart city, a wireless electronic device in a smart home, etc.

[0129] The above introduces the devices to which the attention model and feature extraction method provided in the embodiments of the present application are applied. The following introduces the scenarios to which the attention model and feature extraction method provided in the embodiments of the present application are applied.

[0130] The attention model and feature extraction method provided in the embodiments of the present application can be applied to computer vision or natural language processing. That is, electronic devices can perform computer vision tasks or natural language processing tasks through the above-mentioned attention model and feature extraction method.

[0131] Natural language processing (NLP) is a key area of ​​research in computer science and artificial intelligence. It studies the theories and methods that enable effective communication between humans and computers using natural language. Generally, NLP tasks include machine translation, public opinion monitoring, automatic summarization, opinion extraction, text classification, question answering, text semantic comparison, and speech recognition.

[0132] Computer vision is the science of teaching machines to see. Specifically, it refers to machine vision techniques that use cameras and computers to replace the human eye in identifying, tracking, and measuring objects. Further image processing is performed to make the resulting images more suitable for human observation or for transmission to instrumentation. Generally speaking, computer vision tasks include image classification, object detection, semantic segmentation, and image generation.

[0133] Image recognition is a common classification problem, often also referred to as image classification. Specifically, in image recognition tasks, a neural network takes image data as input and outputs the probability that the image data belongs to each category. The category with the highest probability is typically selected as the predicted category. Image recognition was one of the earliest successful applications of deep learning, with classic network models including the VGG series, the Inception series, and the ResNet series.

[0134] Object detection refers to the automatic detection of the approximate location of common objects in an image through algorithms. A bounding box is usually used to represent the approximate location of the object and to classify the category information of the object in the bounding box.

[0135] Semantic segmentation refers to the automatic segmentation and identification of image content through algorithms. Semantic segmentation can be understood as a per-pixel classification problem, that is, analyzing the category of object to which each pixel belongs.

[0136] Image generation involves learning the distribution of real images and sampling from that distribution to generate realistic images. For example, generating a clear image from a blurry image or generating a dehazed image from a foggy image.

[0137] The above introduces the scenarios in which the attention model and feature extraction method provided in the embodiments of the present application are applied. The following will introduce the specific structure of the model provided in the embodiments of the present application.

[0138] The attention model provided in the embodiment of the present application includes a self-attention network, or multiple self-attention networks connected in series. When the attention model includes multiple self-attention networks connected in series, the structure of each self-attention network in the attention model is the same, but the weight parameters in different self-attention networks may be different. Figure 5 As shown, Figure 5 A schematic diagram of the structure of an attention model provided in an embodiment of the present application. Figure 5 In [1], the attention model includes N serially connected self-attention networks, namely self-attention network 1, self-attention network 2, ..., self-attention network N. The structures of these N networks can be identical, but the weight parameters in self-attention networks 1 to N can be different. The input of self-attention network 1 is the input data of the attention model, the input of self-attention network 2 is the output of self-attention network 1, and the input of self-attention network N is the output of self-attention network N-1.

[0139] Specifically, in the attention model, each self-attention network includes a self-attention module, a multi-layer perceptron, and a first neural network layer. The self-attention module includes multiple parallel feature extraction layers and a fusion layer, each of which is connected to the multiple parallel feature extraction layers and is used to fuse the features output by the multiple parallel feature extraction layers. The self-attention module is a network that uses the self-attention mechanism, which can associate different positions in the input sequence to calculate the representation of the same sequence.

[0140] The multilayer perceptron (MLP) is serially connected to the self-attention module, and the multilayer perceptron includes a plurality of serial first fully connected layers. Specifically, the multilayer perceptron can also be called a fully connected neural network, and the multilayer perceptron includes an input layer, a hidden layer, and an output layer, and the number of hidden layers can be one or more layers. Among them, the network layers in the multilayer perceptron are all fully connected layers. That is, the input layer and the hidden layer of the multilayer perceptron are fully connected, and the hidden layer and the output layer of the multilayer perceptron are also fully connected. Among them, the fully connected layer means that each neuron in the fully connected layer is connected to all neurons in the previous layer, which is used to integrate the features extracted from the previous layer.

[0141] The first neural network layer is connected in parallel to the self-attention module and one or more of the multilayer perceptrons, and the first neural network layer is used to perform feature transformation.

[0142] In this embodiment, the input of the attention model is data in the form of a sequence. That is, the input data of the attention model is sequence data. For example, the input data of the attention model can be a sentence sequence consisting of multiple consecutive words; for another example, the input data of the attention model can be an image block sequence consisting of multiple consecutive image blocks, where the multiple consecutive image blocks are obtained by segmenting a complete image.

[0143] For ease of understanding, the following will introduce in detail various implementation methods of the above-mentioned self-attention network with reference to the accompanying drawings.

[0144] Implementation method 1, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, and the first neural network layer is connected in parallel to the self-attention module.

[0145] See Figure 6a , Figure 6a A structural diagram of the attention model provided in the embodiment of the present application. Figure 6a As shown, in the self-attention network, the self-attention module is connected in series with the multilayer perceptron, and the first neural network layer is connected in parallel with the self-attention module. During the operation of the self-attention network, the self-attention module and the first neural network layer process the input data to the self-attention network in parallel, and the output of the self-attention module and the output of the first neural network layer are added together to form the input of the multilayer perceptron. Finally, after the output of the self-attention module and the output of the first neural network layer are added together, they are further processed by the multilayer perceptron to obtain the output data of the self-attention network.

[0146] It is understandable that the attention model includes multiple serial self-attention networks, and Figure 6aWhen the self-attention network shown is the first self-attention network in the attention model, Figure 6a The data to be processed in the self-attention network is the original sequence data, such as the text data to be processed or the image data to be processed. Figure 6a When the self-attention network shown is not the first self-attention network in the attention model, Figure 6a The data to be processed input into the self-attention network is the feature data output by the previous self-attention network.

[0147] In this solution, a parallel first neural network layer is introduced on top of the self-attention module. This parallel first neural network layer performs feature transformation on the input features to generate transformed features. These transformed features are then added to the output features of the self-attention module to increase the diversity of features output by the intermediate layer of the self-attention network, enhancing the expressiveness of the features and thus improving the performance of the attention model.

[0148] Implementation method 2, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, and the first neural network layer is connected in parallel to the multi-layer perceptron.

[0149] See Figure 6b , Figure 6b Another structural diagram of the attention model provided in the embodiment of the present application. Figure 6b As shown, in the self-attention network, the self-attention module is connected in series with the multilayer perceptron, and the first neural network layer is connected in parallel with the multilayer perceptron. During the operation of the self-attention network, the self-attention module processes the input data to the self-attention network, and the output of the self-attention module serves as the input of both the multilayer perceptron and the first neural network layer. The multilayer perceptron and the first neural network layer process the output of the self-attention module in parallel. Finally, the output of the multilayer perceptron and the output of the first neural network layer are added together to obtain the output data of the self-attention network.

[0150] This approach introduces a parallel first neural network layer based on a multilayer perceptron. This parallel first neural network layer performs a feature transformation on the input features of the multilayer perceptron to generate transformed features. These transformed features are then added to the output features of the multilayer perceptron to increase the diversity of features output by the intermediate layers of the self-attention network, enhancing the expressiveness of the features and thus improving the performance of the attention model.

[0151] Implementation method 3, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, and the first neural network layer is parallel to the self-attention module and the multi-layer perceptron.

[0152] See Figure 6c , Figure 6c Another structural diagram of the attention model provided in the embodiment of the present application. Figure 6c As shown, in the self-attention network, the self-attention module is connected in series with the multilayer perceptron, and the first neural network layer is connected in parallel with the multilayer perceptron. During the operation of the self-attention network, the self-attention module processes the input data to the self-attention network, and the output of the self-attention module serves as the input of both the multilayer perceptron and the first neural network layer. The multilayer perceptron and the first neural network layer process the output of the self-attention module in parallel. Finally, the output of the multilayer perceptron and the output of the first neural network layer are added together to obtain the output data of the self-attention network.

[0153] This approach introduces a parallel first neural network layer based on the self-attention module and the multilayer perceptron. This parallel first neural network layer performs a feature transformation operation on the input features of the self-attention network to obtain transformed features. Furthermore, these transformed features are added to the output features of the multilayer perceptron to increase the diversity of features output by the intermediate layers of the self-attention network, enhance the expressiveness of the features, and thus improve the performance of the attention model.

[0154] Implementation 4: The self-attention network includes a self-attention module, a multi-layer perceptron, a first neural network layer, and a second neural network layer. The first neural network layer is parallel to the self-attention module, and the second neural network layer is parallel to the multi-layer perceptron; alternatively, the second neural network layer is parallel to the self-attention module, and the first neural network layer is parallel to the multi-layer perceptron. Both the first neural network layer and the second neural network layer are used to perform feature transformation. Furthermore, the network structures of the first neural network layer and the second neural network layer can be the same, but the weight parameters in the first neural network layer and the second neural network layer can be different.

[0155] See Figure 6d , Figure 6d Another structural diagram of the attention model provided in the embodiment of the present application. Figure 6dAs shown, in the self-attention network, the self-attention module is connected in series with the multi-layer perceptron, and the first neural network layer is connected in parallel with the self-attention module, and the second neural network layer is connected in parallel with the multi-layer perceptron. During the operation of the self-attention network, the self-attention module and the first neural network layer process the data to be processed input to the self-attention network, and the result obtained by adding the output of the self-attention module and the output of the first neural network layer serves as the input of the multi-layer perceptron and the second neural network layer. In other words, the result obtained by adding the output of the self-attention module and the output of the first neural network layer serves as the input of the multi-layer perceptron and the second neural network layer, and is processed in parallel by the multi-layer perceptron and the second neural network layer. Finally, after performing the addition operation on the output of the multi-layer perceptron and the output of the second neural network layer, the output data of the self-attention network is obtained.

[0156] This approach introduces a first neural network layer and a second neural network layer, running in parallel with the self-attention module and the multi-layer perceptron, respectively. These two parallel neural network layers perform feature transformation operations on the input features to generate transformed features. These transformed features are then added to the output features of the self-attention module and the multi-layer perceptron, respectively, to increase the diversity of features output by the intermediate layers of the self-attention network, enhance the expressiveness of the features, and thus improve the performance of the attention model.

[0157] Optionally, based on the aforementioned four implementation methods, in addition to introducing a neural network layer parallel to the self-attention module and / or multi-layer perceptron, the self-attention network can also include multiple neural network layers, which are parallel to the aforementioned first neural network layer and / or second neural network layer.

[0158] For example, in the aforementioned implementation 1, in addition to the first neural network layer connected in parallel with the self-attention module, the self-attention network may also include multiple neural network layers. The multiple neural network layers are all connected in parallel with the first neural network layer, that is, the self-attention module is simultaneously connected in parallel with the first neural network layer and the multiple neural network layers. The multiple neural network layers are all used to perform feature transformation. Moreover, the network structure of the multiple neural network layers may be the same as the network structure of the first neural network layer, but the weight parameters in the first neural network layer and the multiple neural network layers may be different.

[0159] Taking the above-mentioned implementation method 4 as an example, you can refer to Figure 6e , Figure 6e Another structural diagram of the self-attention network provided in the embodiment of this application. Figure 6eAs shown, the self-attention network includes a self-attention module and a multi-layer perceptron connected in series, as well as n neural network layers connected in parallel to the self-attention module and n neural network layers connected in parallel to the multi-layer perceptron. Specifically, in the self-attention network, the self-attention module and the multi-layer perceptron are connected in series. The self-attention module has n neural network layers connected in parallel, namely neural network layer A1, neural network layer A2, ... neural network layer An. The multi-layer perceptron also has n neural network layers connected in parallel, namely neural network layer B1, neural network layer B2, ... neural network layer Bn.

[0160] In this scheme, by introducing multiple parallel neural network layers, the diversity of features output by the self-attention network can be further enriched.

[0161] In one possible embodiment, based on the four aforementioned implementations, the self-attention module and / or the multilayer perceptron in the self-attention network further includes a shortcut in parallel, where the input features and output features of the shortcut are the same. In other words, the self-attention module and / or the multilayer perceptron also include a shortcut to the first neural network layer and the identity mapping in parallel.

[0162] For example, taking the above implementation method 1 as an example, you can refer to Figure 7 , Figure 7 This is a schematic diagram of a feature processing provided in an embodiment of the present application. Figure 7 As shown in the figure, the self-attention module in the self-attention network has a shortcut and a first neural network layer in parallel. The self-attention module, shortcut, and first neural network layer process the same input features respectively. Then, the output features of the self-attention module, the output features of the shortcut, and the output features of the first neural network layer are added together to obtain the feature addition result, which is used as the input of the subsequent multi-layer perceptron. Figure 7 It can be seen that the input features and output features of the shortcut are the same. After the feature transformation of the first neural network layer, the input features and output features of the first neural network layer are different.

[0163] Compared to the first neural network layer, shortcuts do not modify the input features, meaning they do not perform feature transformations on them. Therefore, shortcuts can be considered a specialized feature processing method. By introducing parallel shortcuts on top of the first neural network layer, we can further enhance the diversity of features obtained by the self-attention network, thereby increasing the expressive power of features.

[0164] The above introduces various implementation methods of self-attention networks. The following will introduce the specific modules in the self-attention network in detail.

[0165] In one possible embodiment, the first neural network layer includes a weight matrix and an activation function, wherein the weight matrix is ​​used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix. In other words, in the process of processing the input features by the first neural network layer, the input features are first multiplied by the weight matrix in the first neural network layer to obtain a multiplication result; then, the multiplication result is processed by the activation function in the first neural network layer to obtain the output features of the first neural network layer.

[0166] Exemplarily, the process of processing the input features through the first neural network layer to obtain output features can be expressed by Formula 1.

[0167]

[0168] in, Represents the output features of the first neural network layer; Z l represents the input features of the first neural network layer; Θ li Represents the weight matrix; σ represents the activation function; i represents the number of blocks of the weight matrix, and the weight matrix can be divided into multiple sub-matrices to perform multiplication operations with the input features; i∈[1, 2, …, T].

[0169] An activation function is a function that runs on neurons in a neural network model, mapping the neuron's input to its output. Activation functions are crucial for neural network models to learn and understand complex and nonlinear functions. They can introduce nonlinear characteristics into neural network models. For example, in a neuron, after the input features are weighted and summed using a weight matrix, a function called the activation function is applied. Activation functions are introduced to increase the nonlinearity of the neural network model. In a neural network model, without an activation function, each network layer is equivalent to a matrix multiplication. Therefore, without an activation function, the output obtained by stacking several network layers is still the result of a matrix multiplication.

[0170] Simply put, if a neural network doesn't use an activation function, the output of each layer is a linear function of the inputs from the previous layer. Regardless of the number of layers in the network, the output is a linear combination of the inputs. However, if an activation function is used, it introduces nonlinearity into the neurons, allowing the network to approximate any nonlinear function. This allows neural networks to be applied to a wide range of nonlinear models.

[0171] For example, the activation function may be a nonlinear function such as a Sigmoid function, a Tanh function, or a ReLU function. This embodiment does not specifically limit the specific type of the activation function.

[0172] It can be understood that the structure of the second neural network layer is the same as that of the first neural network layer, that is, the second neural network layer may also include a weight matrix and an activation function.

[0173] Specifically, when the self-attention module in the self-attention network is run in parallel with the first neural network layer and the shortcut, the output features of the self-attention module, the output features of the first neural network layer, and the output features of the shortcut need to be added together to obtain the added output features. The added output features can be used as the input of the subsequent multi-layer perceptron.

[0174] For example, the process of adding the output features of the self-attention module, the output features of the first neural network layer, and the output features of the shortcut to obtain the added output features can be expressed by Formula 2.

[0175]

[0176] Among them, AugMSA(Z l ) represents the output feature after addition; Z l Represents input features; MSA(Z l ) represents the features obtained based on the self-attention module; Represents the i-th neural network layer parallel to the self-attention module. There are T neural network layers parallel to the self-attention module. The first neural network layer is one of the neural network layers parallel to the self-attention module.

[0177] It is understandable that after the introduction of the parallel first neural network layer, the process of processing features through the parallel first neural network layer will involve matrix multiplication operations. If the matrix multiplication operation is directly performed on the input features and the weight matrix, it will bring huge computational overhead. Therefore, this embodiment proposes to use a block circulant matrix to implement the calculation of the first neural network layer, thereby reducing the computational overhead caused by the introduction of the parallel first neural network layer.

[0178] Specifically, in the first neural network layer described above, which is run in parallel with the self-attention module and / or the multi-layer perceptron, the weight matrix of the first neural network layer includes multiple sub-matrices, and each sub-matrix is ​​a circulant matrix. In simple terms, the weight matrix can be composed of multiple sub-matrices, and each sub-matrix that constitutes the weight matrix is ​​a circulant matrix.

[0179] A circulant matrix is ​​a special form of matrix. Each element of a circulant matrix's row vector is the result of shifting each element of the previous row vector right by one position. Circulant matrices can be efficiently computed using Fourier transforms, saving the computational overhead of transforming features across parallel neural network layers.

[0180] Since each submatrix is ​​a circulant matrix, in practical applications, each submatrix can be represented by a set of vectors. For any set of vectors, the circulant matrix corresponding to that set of vectors can be obtained by right-shifting the elements in the set of vectors to obtain each row in the circulant matrix. Thus, during the training of the self-attention network, when adjusting the weight parameters of the weight matrix in the self-attention network based on the loss function, we are actually adjusting the vectors corresponding to the submatrices in the weight matrix. After the vectors corresponding to the submatrices are adjusted, the adjusted weight matrix can be obtained based on the adjusted vectors.

[0181] Exemplarily, the weight matrix may be represented by the following formula 3.

[0182]

[0183] Among them, Θ represents the weight matrix; the weight matrix Θ is divided into b*b sub-matrices in total, and each sub-matrix is ​​a circulant matrix.

[0184] For any submatrix C in the weight matrix ij , submatrix C ij It can be obtained by looping through the elements in a vector. Specifically, the submatrix C ij It can be expressed based on the following formula 4.

[0185]

[0186] From formula 4, we can know that the sub-matrix C ij It can be achieved by looping vector [C1 ij , C2 ij ,…,C d-1 ij , C d ij ] in the elements to get.

[0187] Exemplarily, when the weight matrix of the first neural network layer is divided into multiple sub-matrices and each sub-matrix is ​​a circulant matrix, multiplying the input features of the first neural network layer by the weight matrix in the first neural network layer can specifically include the following steps.

[0188] First, the electronic device performs Fourier transform on multiple sub-matrices of the first neural network layer and the input features of the first neural network layer to obtain multiple transformed sub-matrices and transformed features.

[0189] Then, the electronic device performs element-by-element multiplication operations on the transformed features and the multiple transformed sub-matrices to obtain multiple sub-features, wherein the element-by-element multiplication operation refers to multiplying the elements at the same position in the two matrices one by one.

[0190] Finally, the electronic device performs inverse Fourier transform on the multiple sub-features, and accumulates the multiple sub-features after the inverse Fourier transform to obtain the processed first feature.

[0191] In this scheme, based on the fact that the weight matrix can be divided into multiple circulant matrices, Fourier transform is used to achieve efficient calculation between input features and weight features, thereby saving the computational overhead of performing feature transformation on features through parallel network layers.

[0192] Furthermore, to facilitate parallel processing of the input features of the first neural network layer on the same hardware, the electronic device may further divide the input features into multiple feature matrices and perform the aforementioned processing on each feature matrix based on the multiple sub-matrices. Finally, the multiple output features obtained are concatenated to obtain features processed based on the weight matrix in the first neural network layer.

[0193] Exemplarily, the process of processing the input features based on the weight matrix of the first neural network layer can be shown as Formula 5 and Formula 6.

[0194]

[0195] in, Represents the features processed based on the weight matrix in the first neural network layer; Represents some features after processing based on the weight matrix in the first neural network layer; Z represents the input features, Z is a matrix of size N*d, N represents the number of image blocks, d represents the dimension of each image block; Z j Represents the j-th feature matrix of the input matrix Z, with a size of N*(d / b), where b is the number of feature matrices divided by the input features; c ij represents a vector of length d / b; FFT represents fast Fourier transform; IFFT represents inverse fast Fourier transform; Represents element-wise multiplication.

[0196] The above introduces the process of processing features by the first neural network layer provided in the embodiment of the present application. The following details the process of processing features by the self-attention module provided in the embodiment of the present application.

[0197] In one possible embodiment, the self-attention module includes multiple parallel feature extraction layers and fusion layers, wherein the fusion layers are respectively connected to the multiple parallel feature extraction layers, and each of the multiple parallel feature extraction layers includes a different weight matrix. Specifically, the input features of the self-attention module are processed by the multiple parallel feature extraction layers to obtain multiple extracted features; then, the fusion layer in the self-attention module fuses the multiple extracted features to obtain the output features of the self-attention module.

[0198] For example, the self-attention module may include three parallel feature extraction layers, each of which includes a weight matrix W Q , weight matrix W K and the weight matrix W V For the input feature of the self-attention module, the input feature includes multiple sub-features. Each sub-feature in the input feature can be first compared with the weight matrix W Q , weight matrix W K and the weight matrix W V Multiply them together to get feature Q, feature K, and feature V. Among them, feature Q, feature K, and feature V are the inputs of the fusion layer of the self-attention module, and are further processed by the fusion layer.

[0199] Specifically, the fusion layer performs a dot product operation on feature Q and feature K to obtain the feature value score. The fusion layer then processes the feature value score through the softmax activation function to obtain the feature value softmax. The fusion layer then performs a dot product operation on the feature value softmax to obtain the score value v corresponding to each sub-feature. Finally, the score value v corresponding to each sub-feature is added together to obtain the output feature of the self-attention module.

[0200] Optionally, the self-attention module can be a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer, the second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.

[0201] For example, assuming that the multi-head self-attention module includes 8 parallel self-attention units, the input features of the multi-head self-attention module can be input into the 8 parallel self-attention units respectively to obtain the feature matrix Z output by the 8 parallel self-attention units. i , i∈{1, 2, ..., 8}. Then, the 8 feature matrices Z iThe columns are spliced ​​into a total feature matrix, and the total feature matrix is ​​processed by the second fully connected layer to obtain the output features of the multi-head self-attention module.

[0202] In order to facilitate verification of the advantages of the attention model provided in the embodiment of the present application in data processing, the beneficial effects of the attention model provided in the embodiment of the present application will be illustrated below based on specific experiments.

[0203] In the embodiments of the present application, extensive experiments are conducted on a large-scale dataset based on existing network models and the attention model provided in the embodiments of the present application to empirically study the proposed method. The large-scale dataset can be the Imagenet (ILSVRC-2012) dataset, which contains 1.28 million training images and validation images from 1,000 categories.

[0204] See Figure 8 , Figure 8 This is a schematic diagram showing the performance comparison of different models on the Imagenet dataset provided in the embodiments of this application. Figure 8 In the figure, Aug-ViT, Aug-PVT, and Aug-T2T represent image processing models that use the attention model provided by the embodiments of this application. Resolution represents image resolution; Top-1 Accuracy represents accuracy; Params represents the number of parameters; and FLOPs represents the number of computations (multiplications and additions).

[0205] Depend on Figure 8 It can be seen that compared with the image processing model in the prior art, the image processing model using the attention model provided in the embodiment of the present application can achieve higher accuracy while keeping the amount of calculation and the amount of parameters basically unchanged.

[0206] An embodiment of the present application also provides a feature extraction method, including: an electronic device obtains data to be processed; the electronic device inputs the data to be processed into one or more serially connected self-attention networks to obtain features of the data to be processed; wherein the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes multiple parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the multiple parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, the multi-layer perceptron includes multiple serial first fully connected layers, the first neural network layer is connected in parallel to the self-attention module and one or more of the multi-layer perceptrons, and the first neural network layer is used to perform feature transformation.

[0207] In one possible implementation, the self-attention network further includes: a second neural network layer, which is used to perform feature transformation; the first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multi-layer perceptron, or the second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multi-layer perceptron.

[0208] In one possible implementation, the first neural network layer includes a weight matrix and an activation function, the weight matrix is ​​used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.

[0209] In a possible implementation, the weight matrix includes multiple sub-matrices, and each sub-matrix is ​​a circulant matrix.

[0210] In a possible implementation, the multiple parallel feature extraction layers respectively include different weight matrices.

[0211] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer. The second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.

[0212] In a possible implementation, the self-attention module and / or the multilayer perceptron further has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.

[0213] In one possible implementation, the method is applied to computer vision tasks or natural language processing tasks.

[0214] An embodiment of the present application also provides an image processing method, including: obtaining an image to be processed; inputting the image to be processed into an image processing model to extract image features through an attention model in the image processing model, wherein the attention model is the attention model described in the aforementioned embodiment; and processing the image to be processed according to the image features.

[0215] In a possible implementation, the processing the image to be processed according to the image features includes: performing one or more of the following tasks on the image to be processed according to the image features: image recognition, target detection, semantic segmentation, and image generation.

[0216] An embodiment of the present application also provides a natural language processing method, including: obtaining a text to be processed; inputting the text to be processed into a natural language processing model to extract text features through an attention model in the natural language processing model, wherein the attention model is the attention model described in the aforementioned embodiment; and processing the text to be processed according to the text features.

[0217] In one possible implementation, the processing of the text to be processed according to the text features includes: performing one or more of the following tasks on the text to be processed according to the text features: machine translation, public opinion monitoring, automatic summary generation, opinion extraction, text classification, question answering, and text semantic comparison.

[0218] In a possible embodiment, the present application provides a feature extraction device. Figure 9 , Figure 9 A structural diagram of a feature extraction device provided in an embodiment of the present application. The feature extraction device 900 includes: an acquisition unit 901 and a processing unit 902; the acquisition unit 901 is used to acquire data to be processed; the processing unit 902 is used to input the data to be processed into one or more serially connected self-attention networks to obtain the features of the data to be processed; wherein the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer, the self-attention module includes multiple parallel feature extraction layers and fusion layers, the fusion layers are respectively connected to the multiple parallel feature extraction layers, the multi-layer perceptron is serially connected to the self-attention module, the multi-layer perceptron includes multiple serial first fully connected layers, the first neural network layer is parallelly connected to the self-attention module and one or more of the multi-layer perceptrons, and the first neural network layer is used to perform feature transformation.

[0219] In one possible implementation, the self-attention network further includes: a second neural network layer, the second neural network layer being configured to perform feature transformation;

[0220] The first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron, or,

[0221] The second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.

[0222] In one possible implementation, the first neural network layer includes a weight matrix and an activation function, the weight matrix is ​​used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.

[0223] In a possible implementation, the weight matrix includes multiple sub-matrices, and each sub-matrix is ​​a circulant matrix.

[0224] In one possible implementation, the self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer. The second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.

[0225] In a possible implementation, the self-attention module and / or the multilayer perceptron further has a shortcut in parallel, and the input feature and output feature of the shortcut are the same.

[0226] In a possible embodiment, the embodiment of the present application also provides an image processing device, including: an acquisition unit and a processing unit; the acquisition unit is used to acquire the image to be processed; the processing unit is used to input the image to be processed into an image processing model to extract image features through an attention model in the image processing model, and the attention model is the attention model described in the first aspect or any implementation method of the first aspect; the processing unit is also used to process the image to be processed according to the image features.

[0227] In a possible implementation, the processing unit is further configured to perform one or more of the following tasks on the image to be processed according to the image features: image recognition, object detection, semantic segmentation, and image generation.

[0228] In a possible embodiment, the embodiment of the present application also provides a natural language processing device, including: an acquisition unit and a processing unit; the acquisition unit is used to acquire the text to be processed; the processing unit is used to input the text to be processed into a natural language processing model to extract text features through the attention model in the natural language processing model, and the attention model is the attention model described in the first aspect or any implementation method of the first aspect; the processing unit is also used to process the text to be processed according to the text features.

[0229] In one possible implementation, the processing unit is further used to perform one or more of the following tasks on the text to be processed based on the text features: machine translation, public opinion monitoring, automatic summary generation, opinion extraction, text classification, question answering, and text semantic comparison.

[0230] Next, we will introduce an execution device provided by the embodiment of the present application. Figure 10 , Figure 10This is a structural diagram of an execution device provided in an embodiment of the present application. The execution device 1000 can be specifically manifested as a mobile phone, a tablet, a laptop, a smart wearable device, a server, etc., which is not limited here. Among them, the execution device 1000 can be deployed with Figure 10 The data processing device described in the corresponding embodiment is used to implement Figure 10 The data processing function in the corresponding embodiment. Specifically, the execution device 1000 includes: a receiver 1001, a transmitter 1002, a processor 1003 and a memory 1004 (wherein the number of the processor 1003 in the execution device 1000 can be one or more, Figure 10 (taking one processor as an example), the processor 1003 may include an application processor 10031 and a communication processor 10032. In some embodiments of the present application, the receiver 1001, the transmitter 1002, the processor 1003 and the memory 1004 may be connected via a bus or other means.

[0231] The memory 1004 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1003. A portion of the memory 1004 may also include non-volatile random access memory (NVRAM). The memory 1004 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.

[0232] Processor 1003 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.

[0233] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1003. Processor 1003 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1003. The above processor 1003 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1003 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 1004, and processor 1003 reads the information in memory 1004 and, in conjunction with its hardware, completes the steps of the above method.

[0234] Receiver 1001 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1002 can be used to output digital or character information through the first interface. Transmitter 1002 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1002 can also include a display device such as a display screen.

[0235] An embodiment of the present application also provides a computer program product, which, when running on a computer, enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.

[0236] A computer-readable storage medium is also provided in an embodiment of the present application, which stores a program for signal processing. When the computer-readable storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.

[0237] The execution device, training device or electronic device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the image processing method described in the above embodiment, or so that the chip in the training device executes the image processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0238] For details, please refer to Figure 11 , Figure 11 This is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1100. NPU 1100 is mounted on the host CPU as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is the arithmetic circuit 1103, which is controlled by a controller 1104 to extract matrix data from memory and perform multiplication operations.

[0239] In some implementations, the arithmetic circuit 1103 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1103 is a two-dimensional systolic array. The arithmetic circuit 1103 may also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1103 is a general-purpose matrix processor.

[0240] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1102 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1101 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1108.

[0241] Unified memory 1106 is used to store input and output data. Weight data is directly transferred to weight memory 1102 through the Direct Memory Access Controller (DMAC) 1105. Input data is also transferred to unified memory 1106 through the DMAC.

[0242] BIU stands for Bus Interface Unit, i.e., bus interface unit 1111 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1109 .

[0243] The bus interface unit 1111 (BIU) is used for the instruction fetch memory 1109 to obtain instructions from the external memory, and is also used for the storage unit access controller 1105 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0244] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1106 or to move weight data to the weight memory 1102 or to move input data to the input memory 1101.

[0245] The vector calculation unit 1107 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1103, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.

[0246] In some implementations, the vector calculation unit 1107 can store the processed output vector in the unified memory 1106. For example, the vector calculation unit 1107 can apply a linear function or a nonlinear function to the output of the operation circuit 1103, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values ​​to generate an activation value. In some implementations, the vector calculation unit 1107 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1103, for example, for use in subsequent layers in a neural network.

[0247] An instruction fetch buffer 1109 connected to the controller 1104 is used to store instructions used by the controller 1104;

[0248] Unified memory 1106, input memory 1101, weight memory 1102, and instruction fetch memory 1109 are all on-chip memories. External memories are private to the NPU hardware architecture.

[0249] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.

[0250] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.

[0251] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.

[0252] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0253] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

Claims

1. A feature extraction method, characterized in that: include: Acquiring data to be processed, wherein the data to be processed is an image or text; Inputting the data to be processed into one or more serially connected self-attention networks to obtain features of the data to be processed; Among them, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer. The self-attention module includes multiple parallel feature extraction layers and fusion layers. The multi-layer perceptron is serially connected to the self-attention module. The multi-layer perceptron includes multiple serial first fully connected layers. The first neural network layer is connected in parallel to the self-attention module. The first neural network layer is used to perform feature transformation.

2. The method according to claim 1, characterized in that The first neural network layer includes a weight matrix and an activation function. The weight matrix is ​​used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.

3. The method according to claim 2, characterized in that The weight matrix includes multiple sub-matrices, and each sub-matrix is ​​a circulant matrix.

4. The method according to claim 1 or 2, characterized in that The self-attention module also has a shortcut in parallel, and the input features and output features of the shortcut are the same.

5. The method according to any one of claims 1 to 4, characterized in that: The self-attention network further includes: a second neural network layer, the second neural network layer being configured to perform feature transformation; The first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron, or, The second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.

6. The method according to any one of claims 1 to 5, characterized in that: The self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer. The second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.

7. The method according to any one of claims 1 to 6, characterized in that: When the data to be processed is an image, the feature of the data to be processed is an image feature, and the method further includes: The image is processed according to the image features.

8. The method according to any one of claims 1 to 6, characterized in that: When the data to be processed is text, the feature of the data to be processed is a text feature, and the method further includes: The text is processed according to the text features.

9. A feature extraction device, characterized in that: include: An acquisition unit, configured to acquire data to be processed, wherein the data to be processed is an image or text; A processing unit, configured to input the data to be processed into one or more serially connected self-attention networks to obtain features of the data to be processed; Among them, the self-attention network includes a self-attention module, a multi-layer perceptron and a first neural network layer. The self-attention module includes multiple parallel feature extraction layers and fusion layers. The multi-layer perceptron is serially connected to the self-attention module. The multi-layer perceptron includes multiple serial first fully connected layers. The first neural network layer is connected in parallel to the self-attention module. The first neural network layer is used to perform feature transformation.

10. The device according to claim 9, characterized in that The first neural network layer includes a weight matrix and an activation function. The weight matrix is ​​used to multiply the input features of the first neural network layer, and the activation function is used to process the multiplication result of the input features and the weight matrix.

11. The device according to claim 10, characterized in that The weight matrix includes multiple sub-matrices, and each sub-matrix is ​​a circulant matrix.

12. The device according to claim 9 or 10, characterized in that The self-attention module also has a shortcut in parallel, and the input features and output features of the shortcut are the same.

13. The device according to any one of claims 9 to 12, characterized in that: The self-attention network further includes: a second neural network layer, the second neural network layer being configured to perform feature transformation; The first neural network layer is connected in parallel to the self-attention module, and the second neural network layer is connected in parallel to the multilayer perceptron, or, The second neural network layer is connected in parallel to the self-attention module, and the first neural network layer is connected in parallel to the multilayer perceptron.

14. The device according to any one of claims 9 to 13, characterized in that: The self-attention module is a multi-head self-attention module, which includes multiple parallel self-attention units and a second fully connected layer. The second fully connected layer is respectively connected to the multiple parallel self-attention units, and the multiple parallel self-attention units each include multiple parallel feature extraction layers and fusion layers.

15. The device according to any one of claims 9 to 14, characterized in that: When the data to be processed is an image, the feature of the data to be processed is an image feature, and the processing unit is further configured to: The image is processed according to the image features.

16. The device according to any one of claims 9 to 14, characterized in that: When the data to be processed is text, the feature of the data to be processed is a text feature, and the processing unit is further configured to: The text is processed according to the text features.

17. An electronic device, characterized in that: The electronic device comprises a memory and a processor; the memory stores codes, the processor is configured to execute the codes, and when the codes are executed, the electronic device executes the method according to any one of claims 1 to 8.

18. A computer storage medium, characterized in that The computer storage medium stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 8.

19. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, cause the computer to implement the method according to any one of claims 1 to 8.