Data processing method and device, electronic equipment and storage medium

By calculating the similarity score between key features and target query features in a high-dimensional space, the problem of self-attention mechanisms failing to capture implicit relationships is solved, thereby improving the training efficiency of deep learning models and reducing training costs.

CN120805981APending Publication Date: 2025-10-17TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410432567.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing self-attention mechanisms struggle to capture the implicit correlation between the positional features of the query vector and the key vector when calculating weights, leading to increased training costs for deep learning models.

Method used

By performing a linear transformation on the input sequence, query feature sequence, key feature sequence, and value feature sequence are obtained. The similarity score between the key feature and the target query feature is calculated in a high-dimensional space. Based on these scores, a self-attention representation is calculated and substituted into a deep learning model.

Benefits of technology

It improves the training efficiency of deep learning models, reduces training costs, can capture more implicit relationships between features, and reduces the learning difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805981A_ABST
    Figure CN120805981A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the linear transformation of an input sequence, and obtaining a query feature sequence, a key feature sequence and a value feature sequence corresponding to the input sequence; the query feature sequence, the key feature sequence and the value feature sequence have n features; for each key feature in the key feature sequence, calculating a similarity score of the key feature and the target query feature in a high-dimensional space; the target query feature is any query feature in the query feature sequence; and based on the n similarity scores and n value features in the value feature sequence, calculating a self-attention representation corresponding to the target query feature, so that the self-attention representation is substituted into the deep learning model. In the embodiment of the application, the similarity score of the key feature and the target query feature can be calculated in a high-dimensional space, a more implicit association relationship between the features can be captured, and the training cost of a deep learning model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the computer field, and in particular to a data processing method and device, electronic equipment and storage medium. BACKGROUND

[0002] The self-attention mechanism is a method for capturing the internal dependency of sequence data in a deep learning model, which captures the internal dependency of sequence data by assigning weights, and is used to improve the expression ability of the model. In the existing self-attention mechanism, when calculating the weight, the inner product of the query vector in the query vector sequence and the key vector in the key vector sequence is often obtained.

[0003] However, if there is a relatively implicit association relationship between the features corresponding to the positions of the query vector and the key vector, the existing weight calculation method often fails to capture the above relationship, resulting in an increased training cost of the deep learning model. SUMMARY

[0004] The embodiments of the present application provide a data processing method and device, electronic equipment and storage medium, which can improve the problem of high training cost of the deep learning model in the prior art.

[0005] The embodiments of the present application provide a data processing method for processing an input sequence, the input sequence comprising n features, n being a positive integer; the method comprising: performing linear transformation on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence; wherein the number of features in the query feature sequence, the key feature sequence and the value feature sequence is n; for each key feature in the key feature sequence, calculating a similarity score of the key feature and a target query feature in a high-dimensional space; wherein the target query feature is any one of the query features in the query feature sequence; based on n similarity scores and n value features in the value feature sequence, calculating a self-attention representation corresponding to the target query feature, so as to substitute the self-attention representation into a deep learning model.

[0006] The embodiments of the present application provide a data processing device for processing an input sequence, the input sequence comprising n features, n being a positive integer; the device comprising:

[0007] a linear transformation unit configured to perform linear transformation on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence, wherein the query feature sequence, the key feature sequence, and the value feature sequence each have n features;

[0008] a high-dimensional space unit configured to calculate, for each key feature in the key feature sequence, a similarity score between the key feature and a target query feature in a high-dimensional space, wherein the target query feature is any one of the query feature sequence;

[0009] a self-attention unit configured to calculate, based on the n similarity scores and n value features in the value feature sequence, a self-attention representation corresponding to the target query feature, so as to substitute the self-attention representation into a deep learning model.

[0010] In an embodiment, the high-dimensional space unit comprises:

[0011] an inner product calculation sub-unit configured to calculate an inner product of a mapping value of the key feature in the high-dimensional space and a mapping value of the target query feature in the high-dimensional space;

[0012] a score calculation sub-unit configured to calculate the similarity score based on the inner product and a dimension value of the query feature sequence.

[0013] In an embodiment, the inner product calculation sub-unit is specifically configured to substitute the key feature and the target query feature into a kernel function to obtain a function value of the kernel function, wherein the function value is the inner product of the mapping value of the key feature in the high-dimensional space and the mapping value of the target query feature in the high-dimensional space.

[0014] In an embodiment, the inner product calculation sub-unit comprises:

[0015] a first intermediate sub-sub-unit configured to calculate a product of a transpose of the target query feature and the key feature to obtain a first intermediate result;

[0016] a first function sub-sub-unit configured to calculate the function value based on the first intermediate result and at least one hyperparameter.

[0017] In an embodiment, the inner product calculation sub-unit comprises:

[0018] a second intermediate sub-sub-unit configured to calculate an absolute value of a difference between the target query feature and the key feature to obtain a second intermediate result;

[0019] The second function subunit is configured to calculate the function value based on the second intermediate result and at least one hyperparameter.

[0020] In an embodiment, the kernel function is any one of a linear kernel function, a polynomial kernel function, a Laplacian kernel function, a sigmoid kernel function, and a Gaussian kernel function.

[0021] In an embodiment, the score calculation subunit comprises:

[0022] The square root subunit is configured to perform square root operation on the dimension value to obtain a square root result.

[0023] The ratio subunit is configured to calculate a ratio of the inner product and the square root result, and the ratio is the similarity score.

[0024] In an embodiment, the self-attention unit comprises:

[0025] The normalization subunit is configured to perform normalization processing on n similarity scores to obtain n weight values, and the n weight values correspond to n value features in the value feature sequence one by one.

[0026] The product subunit is configured to calculate a product of each value feature and a corresponding weight value.

[0027] The summation subunit is configured to calculate a summation of n products, and the summation is a self-attention representation corresponding to the target query feature.

[0028] In an embodiment, the normalization subunit comprises:

[0029] The exponentiation subunit is configured to calculate an exponentiation result of each similarity score.

[0030] The exponentiation summation subunit is configured to calculate a summation of n exponentiation results of the similarity scores.

[0031] The weight subunit is configured to calculate a ratio of each exponentiation result and the summation of the exponentiation results to obtain a corresponding weight value.

[0032] In an embodiment, the linear transformation unit comprises:

[0033] The query weight matrix subunit is configured to transform the input sequence by using a query weight matrix to obtain the query feature sequence.

[0034] The key weight matrix subunit is configured to transform the input sequence by using a key weight matrix to obtain the key feature sequence.

[0035] A value weight matrix subunit is configured to convert the input sequence by using a value weight matrix to obtain the value feature sequence.

[0036] In the data processing method provided in the embodiments of the present application, linear transformation is performed on the input sequence to obtain the query feature sequence corresponding to the input sequence, the key feature sequence corresponding to the input sequence, and the value feature sequence corresponding to the input sequence. The query feature sequence includes n query features, the key feature sequence includes n key features, and the value feature sequence includes n value features. The n query features, the n key features, and the n value features each correspond to one of the n features of the input sequence. For each key feature of the n key features, a similarity score of the key feature and a target query feature in a high-dimensional space is calculated, and n similarity scores are obtained. The target query feature is any one of the n query features. Subsequently, a self-attention representation corresponding to the target query feature is calculated based on the n similarity scores and the n value features, and the self-attention representation is substituted into the deep learning model, so as to reduce the learning difficulty of the deep learning model.

[0037] In the embodiments of the present application, the similarity score of the key feature and the target query feature in the high-dimensional space can be calculated. Compared with directly calculating the inner product of the key feature and the query feature as the similarity score in the prior art, the calculation in the high-dimensional space can capture more implicit correlation between the features, thereby improving the training efficiency of the deep learning model and reducing the training cost of the deep learning model. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0039] Figure 1a is an application scenario diagram of the data processing method provided by the present application;

[0040] Figure 1b is a flowchart of the data processing method provided by the embodiments of the present application;

[0041] Figure 1c shows a schematic structural diagram of the operation process of the improved self-attention mechanism;

[0042] Figure 1d shows a schematic structural block diagram of the specific operation of the dashed box position in Figure 1c

[0043] Figure 1e ​A schematic structural diagram showing an improved self-attention mechanism layer applied to a text feature extraction task is shown.

[0044] Figure 2 is a flowchart of a data processing method provided by an embodiment of the present application;

[0045] Figure 3 is a structural schematic diagram of a data processing device provided by an embodiment of the present application;

[0046] Figure 4 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0047] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, any other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0048] The embodiments of the present application provide a data processing method and device, an electronic device and a storage medium.

[0049] The data processing device can be integrated in an electronic device, which can be a terminal, a server or the like. The terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a notebook computer or a personal computer (PC) or the like. The server can be a single server or a server cluster composed of multiple servers.

[0050] In some embodiments, the data processing device can also be integrated in multiple electronic devices, for example, the data processing device can be integrated in multiple servers to implement the data processing method of the present application.

[0051] In some embodiments, the server can also be implemented in the form of a terminal.

[0052] For details, see Figure 1a The method provided by the embodiments of the present application can include: performing linear transformation on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence and a value feature sequence corresponding to the input sequence; the number of features of the query feature sequence, the key feature sequence and the value feature sequence is n; for each key feature in the key feature sequence, the similarity score of the key feature and a target query feature in a high-dimensional space is calculated; the target query feature is any query feature in the query feature sequence, without loss of generalityFigure 1a Taking query feature 2 as an example, a self-attention representation corresponding to the target query feature is calculated based on the n similarity scores and the n value features in the value feature sequence, so as to substitute the self-attention representation into the deep learning model.

[0053] In this method, the similarity score between key features and target query features can be calculated in a high-dimensional space. Compared to directly calculating the inner product of key and query features as the similarity score, calculating in a high-dimensional space can capture more implicit relationships between features, thereby improving the training efficiency of deep learning models and reducing their training costs.

[0054] Furthermore, the applicant discovered that directly calculating the inner product of the key feature and the query feature as the similarity score is a case of encoding two input quantities of different dimensions into a single value. This approach increases the learning difficulty of deep learning models and affects their learning and expression. Calculating the similarity score between the key feature and the target query feature in a high-dimensional space shifts the weight information at each position in the input sequence from direct encoding to implicit encoding, which helps to improve the above problem and thus reduces the learning difficulty of deep learning models.

[0055] The data processing method provided in the embodiment of the present application is an improvement to the original self-attention mechanism operation process. The improved self-attention mechanism can be widely used in the fields of deep learning and natural language processing (NLP). For example, machine translation: the self-attention mechanism is widely used in machine translation tasks in the Transformer model, which can capture the long-distance dependencies between the source language and the target language, thereby improving the translation quality. Text classification: the self-attention mechanism can be used for text classification tasks, such as sentiment analysis, topic classification, etc. By capturing the correlation information between different words in the text, the self-attention mechanism helps to improve classification accuracy. Question and answer system: the self-attention mechanism plays an important role in reading comprehension and question and answer systems. By focusing on the correlation information between questions and articles, the self-attention mechanism helps to extract the correct answer. Text generation: in text generation tasks, such as summary generation, dialogue generation, etc., the self-attention mechanism can capture the dependencies in the input sequence, thereby generating more coherent and natural text. Speech recognition: the self-attention mechanism can also be applied to speech recognition tasks, by focusing on the temporal dependencies in the speech signal to improve recognition accuracy.

[0056] The improved self-attention mechanism can also be widely applied in the field of computer vision. Through the self-attention mechanism, the model can capture local and global dependencies in images or videos, thereby improving the performance of the model. Related applications include, for example, image classification: in the image classification task, the self-attention mechanism can help the model focus on important areas related to classification. For example, the ViT (Vision Transformer) model divides the image into multiple small patches, and then uses the self-attention mechanism to capture the relationship between the patches, thereby improving the classification accuracy. Object detection: in the object detection task, the self-attention mechanism can help the model focus on the target objects and their context information in the image. For example, the DETR (Detection Transformer) model combines the traditional convolutional neural network (CNN) with the Transformer, using the self-attention mechanism to capture the relationship between the target objects and the background, thereby improving the detection performance. Semantic segmentation: in the semantic segmentation task, the self-attention mechanism can help the model focus on local and global information related to pixel classification in the image. Video understanding: in the video understanding task, the self-attention mechanism can help the model capture the temporal dependency between video frames, thereby improving the video classification and prediction performance.

[0057] It can be understood that in the embodiments of the present application, data related to user information and the like are involved, and when the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of countries and regions.

[0058] The following will be described in detail. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.

[0059] In this embodiment, a data processing method is provided. As shown in the figure, the data processing method is applied to an electronic device. The specific process of the method can include the following steps 110 to 130: Figure 1b

[0060] 110, performing linear transformation on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence.

[0061] Among them, the number of features of the query feature sequence, the key feature sequence and the value feature sequence is n.

[0062] ​Input sequence refers to sequence data that needs to be processed in a self-attention mechanism. For example, in a text classification task, the input sequence can be a sequence of word vectors in a piece of text, or a sequence of sentence vectors / paragraph vectors in a piece of text; in a machine translation task, the input sequence is a sequence of word vectors in a source language; in a time series prediction task, the input sequence is a time series composed of a series of observation values. The input sequence is a collection of data that needs to be processed and learned by a deep learning model.

[0063] In a sequence modeling task, each position of the input sequence corresponds to a feature, which is used to represent the information of the corresponding position. Let the input sequence include n features, n being a positive integer, and the dimension of each feature be D x , for details, see Figure 1c . The feature can be a feature vector.

[0064] Continuing the example above, if the input sequence is a sequence of word vectors in a piece of text, or a sequence of sentence vectors / paragraph vectors in a piece of text, then correspondingly, the feature vector is a word vector, a sentence vector, or a paragraph vector in the text. If the input sequence is a sequence of word vectors in a source language, then correspondingly, the feature vector is a word vector. If the input sequence is a time series composed of a series of observation values, then correspondingly, the feature vector is an observation value vector in the time series data.

[0065] The process of linear transformation is usually achieved by multiplying the input sequence with three weight matrices. The three weight matrices can specifically include a query weight matrix W q , a key weight matrix W k , and a value weight matrix W v . Optionally, in a specific implementation, the specific process of linear transformation can include steps 111 to 113 as follows:

[0066] 111. Transform the input sequence using the query weight matrix to obtain the query feature sequence.

[0067] Optionally, the input sequence can be multiplied by the query weight matrix W q to obtain the query feature sequence, for details, see Figure 1c . The query feature sequence can specifically include n query features: query feature 1, query feature 2, …, and query feature n, for details, see Figure 1a . The dimension of each query feature can be denoted as D k .

[0068] 112. Transform the input sequence using the key weight matrix to obtain the key feature sequence.

[0069] Optionally, the input sequence can be multiplied by the key weight matrix W k to obtain a key feature sequence, details of which are described below. Figure 1c The key feature sequence can specifically include n key features: key feature 1, key feature 2, …, key feature n, details of which are described below. Figure 1a The dimension of the key feature is equal to the dimension of the query feature. Continuing the example above, the dimension of the key feature is also represented by D k .

[0070] 113. The input sequence is converted using a value weight matrix to obtain the value feature sequence.

[0071] Optionally, the input sequence can be multiplied by the value weight matrix W v to obtain a value feature sequence, details of which are described below. Figure 1c The value feature sequence can specifically include n value features: value feature 1, value feature 2, …, value feature n, details of which are described below. Figure 1a The dimension of each value feature is represented by D v .

[0072] In the above embodiments, the query feature sequence, the key feature sequence, and the value feature sequence can be obtained by multiplying the input sequence with the query weight matrix W q , the key weight matrix W k , and the value weight matrix W v , respectively, which lays a good foundation for subsequent operations.

[0073] 120. For each key feature in the key feature sequence, a similarity score between the key feature and a target query feature in a high-dimensional space is calculated.

[0074] The target query feature is any one of the query features in the query feature sequence. For ease of description, the target query feature is represented by q i , where i is a positive integer less than or equal to n.

[0075] Since the same calculation process is performed for each key feature, the jth key feature in the key feature sequence is taken as an example and represented by k j , where j is a positive integer less than or equal to n. In the following, the similarity score between the target query feature q i and the key feature k j in a high-dimensional space is calculated.

[0076] Optionally, in a specific embodiment, the step of “calculating a similarity score between the key feature and a target query feature in a high-dimensional space” can specifically include steps 121 and 122 as follows:

[0077] 121、compute an inner product of the mapping value of the key feature in the high-dimensional space and the mapping value of the target query feature in the high-dimensional space.

[0078] With the above example continuing to illustrate, it can be assumed that the key feature k j is mapped to Φ(k j ) in the high-dimensional space, and the target query feature q i is mapped to Φ(q i ) in the high-dimensional space. i Correspondingly, the inner product of the two can be expressed as Φ(q j )·Φ(k i ).

[0079] Optionally, in a specific implementation, the step 121 can specifically include: substituting the key feature and the target query feature into a kernel function to obtain a function value of the kernel function. The function value is an inner product of the mapping value of the key feature in the high-dimensional space and the mapping value of the target query feature in the high-dimensional space.

[0080] Kernel function is an important concept frequently used in machine learning, such as support vector machine (SVM) and kernel method (Kernel Methods) algorithms. Kernel function can be used to map input data into a higher-dimensional feature space without explicitly calculating or storing the mapped data.

[0081] For convenience, the query feature sequence is denoted as Q, and the key feature sequence is denoted as K. For details, please refer to Figure 1d For convenience, it is assumed that i is 1 and j is 1, and the kernel function is denoted as Ke(·). The calculation process of Φ(q Figure 1d 1)·Φ(k Figure 1d 1) is as shown in Figure 1c The schematic structural block diagram of the specific operation of the dashed box position in

[0082] In the above embodiment, the inner product Φ(q i )·Φ(k j ) of the mapping value of the key feature in the high-dimensional space and the mapping value of the target query feature in the high-dimensional space can be calculated by substituting the key feature and the target query feature into the kernel function to obtain the function value of the kernel function, without separately calculating the mapping value Φ(k j ) of the key feature in the high-dimensional space and the mapping value Φ(q i ) of the target query feature in the high-dimensional space. The above calculation process can simplify the calculation process, thereby improving the calculation efficiency.

[0083] Optionally, the kernel function may be a linear kernel function (Linear Kernel): Ke(x,y)=x T y+c, Polynomial Kernel: Ke(x,y)=(αx T y+c) d , Laplacian Kernel: Sigmoid kernel function: Ke(x,y)=tanh(αx T y+c), Gaussian Radial Basis Function (RBF): Any of the above. T represents the transpose, c, d, α, σ, are all hyperparameters.

[0084] In one embodiment, the kernel function may be a linear kernel function, a polynomial kernel function, or a sigmoid kernel function; accordingly, the step of "substituting the key feature and the target query feature into the kernel function to obtain a function value of the kernel function" may specifically include the following steps S1 to S2:

[0085] S1. Calculate the product of the transpose of the target query feature and the key feature to obtain a first intermediate result.

[0086] Continuing with the above example, let’s take the target query feature q i With key feature k j Substituting in, we can get the first intermediate result: q i T k j .

[0087] S2. Calculate the function value based on the first intermediate result and at least one hyperparameter.

[0088] Optionally, a first intermediate result q may be calculated i T k j The sum of the hyperparameter c gives the function value of the linear kernel function: Ke(q i ,k j )=q i T k j +c; you can also do the first intermediate result q i T k j Perform scaling operation to obtain the scaling result αq i T k j , and calculate the scaling result αq i T k jthe sum of the first intermediate result q i T k j +c and the hyperparameter c, and then perform an exponential operation on the sum αq i T k j +c, to obtain a function value Ke(q i ,k j ) of the polynomial kernel function. i T k j +c) d ; a scaling operation can also be performed on the first intermediate result q i T k j to obtain a scaling result αq i T k j , and the scaling result αq i T k j and the hyperparameter c, and then perform a hyperbolic tangent function operation on the sum αq i T k j +c, to obtain a function value Ke(q i T k j +c) of the sigmoid kernel function. i ,k j ) = tanh(αq i T k j +c).

[0089] In the above embodiments, the first intermediate result can be calculated first, and then based on the first intermediate result, at least one hyperparameter, and an operation process corresponding to the hyperparameter, the function value operation of the linear kernel function, the polynomial kernel function, or the sigmoid kernel function can be implemented.

[0090] In another embodiment, the kernel function can be a Laplace kernel function or a Gaussian kernel function; accordingly, the step of “substituting the key feature and the target query feature into the kernel function to obtain a function value of the kernel function” can specifically include the following steps X1 to X2:

[0091] X1. Calculate an absolute value of a difference between the target query feature and the key feature to obtain a second intermediate result.

[0092] Continuing with the above example, when the target query feature q i and the key feature k j are substituted, the second intermediate result is ||q i -kj ||.

[0093] X2, based on the second intermediate result and at least one hyperparameter, calculating the function value of the function.

[0094] Optionally, the second intermediate result ||q i -k j || and the hyperparameter -σ and the ratio is subjected to an exponential operation to obtain the function value of the Laplace kernel function: The second intermediate result ||q i -k j || can also be calculated i -k j || 2 Then, the square result ||q i -k j || 2 and -2σ 2 is calculated Then, the ratio is subjected to an exponential operation to obtain the function value of the Gaussian kernel function:

[0095] In the above embodiments, the second intermediate result can be calculated first, and then based on the second intermediate result, at least one hyperparameter, and the operation process corresponding to the above hyperparameter, the function value operation of the Laplace kernel function or the Gaussian kernel function is realized.

[0096] 122, based on the inner product and the dimension value of the query feature sequence, calculating the similarity score.

[0097] The similarity score is used to reflect the correlation degree of q i and k j .

[0098] Continuing the above example, the inner product is: substituting q i and k j into the function value Ke(q i , k j ) of the kernel function; the dimension value of the query feature sequence and the dimension value of the key feature sequence are consistent, both are D k . Accordingly, based on the inner product Ke(q i , k j ) and the dimension value D k of the query feature sequence, the process of calculating the similarity score can be as follows: steps 1221 to 1222:

[0099] 1221, square root operation is performed on the dimension value to obtain a square root result.

[0100] By way of continuation of the foregoing example, the dimension value D k is subjected to a square root operation to obtain a square root result

[0101] 1222、Calculate the ratio of the inner product and the square root result, and the ratio is the similarity score.

[0102] By way of continuation of the foregoing example, the inner product Ke(q i ,k j ) and the square root result is calculated to obtain the similarity score:

[0103] 130、Based on n similarity scores and n value features in the value feature sequence, calculate the self-attention representation corresponding to the target query feature, so as to substitute the self-attention representation into the deep learning model.

[0104] The self-attention representation corresponding to the target query feature is used to reflect the degree of association of the target query feature with the n features of the input sequence. By way of continuation of the foregoing example, the target query feature is q i , that is, the i-th query feature in the query feature sequence. Since the n query features in the query feature sequence are one-to-one corresponding to the n features of the input sequence, the target query feature q i is the i-th feature in the n features: feature 1, feature 2, …, feature n, that is, feature i. Correspondingly, the self-attention representation corresponding to the target query feature is used to reflect the degree of association of feature i with the n features: feature 1, feature 2, …, feature n. The self-attention representation can specifically reflect the degree of association between features through the numerical value. For example, the larger the numerical value, the higher the degree of association between two features; the smaller the numerical value, the lower the degree of association between two features.

[0105] The target query feature is any one of the query features in the query feature sequence (that is, any column in Q shown in Figure 1c , for each query feature in the query feature sequence, the above steps 110 to 130 are performed, and the output quantity H can be obtained, for details, please refer to Figure 1c , the output quantity H includes n features, and the dimension of each feature is D v . Then, the H is subjected to scaling processing to obtain the output quantity Z, and the output quantity Z includes n features, and the dimension of each feature is D x , which is the same as the dimension of each feature in the input sequence.

[0106] The self-attention representation can be applied to a training phase of a deep learning model. Due to the operation process of the self-attention representation, more implicit correlation between features in a high-dimensional space can be captured, so that the learning difficulty of the deep learning model can be reduced, and thus the training cost of the deep learning model can be reduced.

[0107] The self-attention representation can also be applied to a prediction phase of the deep learning model. Compared with the operation process of the self-attention representation in the prior art, the operation process of the self-attention representation in the embodiment of the present application can capture more implicit correlation that cannot be captured by the prior art, so that the prediction result of the deep learning model can be more accurate.

[0108] Optionally, in a specific implementation, step 130 can specifically include steps 131 to 133 as follows:

[0109] 131. Normalizing the n similarity scores to obtain n weight values.

[0110] The n weight values correspond one-to-one to n value features in the value feature sequence.

[0111] Normalization is used to scale multiple data in proportion, so that the multiple data falls within a specific range. In the embodiment of the present application, since the weight values are calculated, the n similarity scores can be limited to between 0 and 1, and the sum of the n weight values is 1.

[0112] In the above implementation, the weight values are obtained by normalizing the similarity scores. The operation of the similarity scores involves the operation of the key features and the target query features in the high-dimensional space, as shown in Figure 1c and Figure 1d The above operation process changes the weight values between the two features in the input sequence X from direct encoding to implicit encoding, reducing the learning difficulty of the deep learning model.

[0113] Optionally, in a specific implementation, the specific calculation process of the normalization can include steps 1311 to 1313 as follows:

[0114] 1311. For each similarity score, calculate the exponential result of the similarity score.

[0115] Since the similarity scores of the n key features in the key feature sequence: key feature 1 (k1), key feature 2 (k2), …, key feature n (kn) and the target query feature q are calculated, n similarity scores Attention(q, k) are obtained. n ) are calculated, n similarity scores Attention(q i , k i ) are obtained. j), j e [1, n]. Exponentiating each of the n similarity scores, the exponentiation results of the n similarity scores are obtained: exp[Attention(q i ,k j )], j e [1, n].

[0116] 1312、Computing the sum of the exponentiation results of the n similarity scores.

[0117] Continuing the example above, the sum of the exponentiation results of the n similarity scores can be calculated by the formula

[0118] 1313、Computing the ratio of each of the exponentiation results to the sum of the exponentiation results to obtain a corresponding weight value.

[0119] Computing the ratio of each of the n exponentiation results to the sum of the above, the n weight values can be obtained.

[0120] Continuing the example above, the n weight values are: where j e [1, n]; x i is the i-th feature in the input sequence, i.e., the target query feature q i is the corresponding feature of the input sequence; x j is the j-th feature in the input sequence, i.e., the key feature j(k j ) is the corresponding feature of the input sequence; Attention_Weights(x i ,x j ) is used to reflect the dependency relationship between x i and x j .

[0121] The n weight values correspond one-to-one to the n value features in the value feature sequence: value feature 1 (i.e., V1), value feature 2 (i.e., V2), …, value feature n (i.e., V n ).

[0122] 132、For each of the value features, computing the product of the value feature and the corresponding weight value.

[0123] Continuing the example above, the product of the value feature and the corresponding weight value can be represented as: Attention_Weights(x i ,x j ) · V j , where j e [1, n].

[0124] 133、Computing the sum of the n products, and the sum is the self-attention representation corresponding to the target query feature. ​

[0125] By way of continuation of the above example, the sum of n products can be expressed as: is the self-attention representation corresponding to the target query feature.

[0126] In the data processing method provided by the embodiments of the present application, linear transformation can be performed on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence. The query feature sequence includes n query features, the key feature sequence includes n key features, and the value feature sequence includes n value features. The n query features, the n key features, and the n value features each correspond to one of the n features of the input sequence. For each key feature of the n key features, a similarity score of the key feature and a target query feature in a high-dimensional space is calculated, and n similarity scores are obtained. The target query feature is any one of the n query features. Subsequently, based on the n similarity scores and the n value features, a self-attention representation corresponding to the target query feature is calculated, and the self-attention representation is substituted into the deep learning model, so as to reduce the learning difficulty of the deep learning model. In the embodiments of the present application, the similarity score of the key feature and the target query feature in the high-dimensional space can be calculated. Compared with the prior art of directly calculating the inner product of the key feature and the query feature as the similarity score, the calculation in the high-dimensional space can capture more implicit relationships between the features, thereby improving the training efficiency of the deep learning model.

[0127] The embodiments of the present application can reduce the training cost of the deep learning model.

[0128] The data processing method provided by the embodiments of the present application describes a calculation process of an improved self-attention mechanism, which can be applied to various self-attention networks. In an embodiment, see Figure 1e , Figure 1e An illustrative structural diagram of applying the improved self-attention mechanism layer to a text feature extraction task is shown. After the input text is segmented, an input sequence X' = {x1, x2} is obtained. As shown in Figure 1c The operation process of the self-attention mechanism shown in Figure 1e The improved self-attention mechanism layer shown in

[0129] For details, see Figure 1e, the x1 in the input sequence is processed by positional encoding to obtain x1'; the x2 in the input sequence is processed by positional encoding to obtain x2'. Positional encoding is one of the commonly used techniques when using the self-attention mechanism in the Transformer model. Since the self-attention mechanism itself does not have the ability to process the order information of the elements in the sequence, it is necessary to introduce positional encoding to provide the position information of the elements in the sequence. In the Transformer model, positional encoding is usually achieved by encoding the position information into a vector. A common method is to use sine and cosine functions to encode the information of different positions, so that the encodings of different positions can be distinguished from each other and some position information is preserved.

[0130] The x1' is processed by the improved self-attention mechanism layer to obtain the output z1 corresponding to the x1'; the x2' is processed by the improved self-attention mechanism layer to obtain the output z2 corresponding to the x2'. The specific calculation process of the improved self-attention mechanism layer is the same as that shown in the calculation process, which will not be repeated here. Figure 1c

[0131] The z1 and z2 are input into the first addition & normalization layer together with the x1' and the x2' to obtain the output results z1' and z2'. The addition & normalization layer is a common interlayer connection and normalization operation in the Transformer model. In the addition operation, the output z1 of the improved self-attention mechanism layer is added to the corresponding input x1' to obtain the addition result 1 (not shown in the figure); the output z2 of the improved self-attention mechanism layer is also added to the corresponding input x2' to obtain the addition result 2 (not shown in the figure).

[0132] Subsequently, in the normalization operation, the addition result 1 and the addition result 2 are normalized to obtain the normalized results z1' and z2'. Common normalization methods include Layer Normalization or Instance Normalization.

[0133] The normalized result z1' is processed by the first feedforward layer to obtain the first feedforward result (not shown in the figure); the normalized result z2' is processed by the second feedforward layer to obtain the second feedforward result (not shown in the figure). The feedforward layer is usually a fully connected feedforward neural network, which consists of two linear transformations and a nonlinear activation function.

[0134] The first feedforward result and the second feedforward result are input into the second addition & normalization layer together with the z1' and the z2' to obtain the output results y1 and y2.

[0135] ​The above operation process occurs in a Transformer Block. In the above embodiment, N Transformer Blocks with the same structure (N is a positive integer) are connected in series to form a Transformer model. In this embodiment, the feature in the first position of the output result of the Nth Transformer Block is extracted as the encoding feature of the input text, which is used for subsequent tasks such as text classification.

[0136] The improved self-attention mechanism described in the embodiments of the present application can be widely applied to various Transformer networks to process various tasks such as image classification, video classification, and speech classification, in addition to the above network.

[0137] In this embodiment, the dimension value of the feature of the query feature sequence is equal to the dimension value of the feature of the key feature sequence, both of which are D k For example, the method of the embodiments of the present application is described in detail.

[0138] As shown in Figure 2 , a data processing method specifically includes the following steps:

[0139] 201. The input sequence is converted using a query weight matrix to obtain the query feature sequence, the number of features of the query feature sequence is n, and the dimension value of each feature is D k .

[0140] 202. The input sequence is converted using a key weight matrix to obtain the key feature sequence, the number of features of the key feature sequence is n, and the dimension value of each feature is D k .

[0141] 203. The input sequence is converted using a value weight matrix to obtain the value feature sequence, the number of features of the value feature sequence is n, and the dimension value of each feature is D v .

[0142] 204. For each key feature in the key feature sequence, the key feature and the target query feature are substituted into the kernel function to obtain the function value of the kernel function, wherein the function value is the inner product of the mapping value of the key feature in the high-dimensional space and the mapping value of the target query feature in the high-dimensional space, and the target query feature is any query feature in the query feature sequence.

[0143] In one embodiment, substituting the key feature and the target query feature into a kernel function to obtain a function value of the kernel function includes: calculating the product of the transpose of the target query feature and the key feature to obtain a first intermediate result; and calculating the function value based on the first intermediate result and at least one hyperparameter.

[0144] In another embodiment, substituting the key feature and the target query feature into a kernel function to obtain a function value of the kernel function includes: calculating the absolute value of the difference between the target query feature and the key feature to obtain the second intermediate result; and calculating the function value based on the second intermediate result and at least one hyperparameter.

[0145] Optionally, the kernel function is any one of a linear kernel function, a polynomial kernel function, a Laplace kernel function, a sigmoid kernel function, and a Gaussian kernel function.

[0146] 205. Dimension value D of query feature sequence k Perform a square root operation to obtain the square root result.

[0147] 206. Calculate a ratio of the inner product to the square root result, where the ratio is the similarity score.

[0148] 207. Normalize the n similarity scores to obtain n weight values; the n weight values ​​correspond one-to-one to the n value features in the value feature sequence.

[0149] Optionally, in a specific embodiment, the n similarity scores are normalized to obtain n weight values, including: for each similarity score, calculating the exponential result of the similarity score; calculating the sum of the exponential results of the n similarity scores; and calculating the ratio of each exponential result to the sum of the exponential results to obtain the corresponding weight value.

[0150] 208. For each value feature, calculate the product of the value feature and the corresponding weight value.

[0151] 209. Calculate the sum of n products, where the sum is the self-attention representation corresponding to the target query feature, so as to substitute the self-attention representation into a deep learning model.

[0152] The specific execution process of steps 201 to 209 has been described in detail above and will not be repeated here.

[0153] In the data processing method provided in the embodiment of the present application, the input sequence can be linearly transformed to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence. Among them, the query feature sequence includes n query features, the key feature sequence includes n key features, and the value feature sequence includes n value features. The n query features, n key features, and n value features all correspond one-to-one to the n features of the input sequence. For each of the n key features, the similarity score between the key feature and the target query feature in the high-dimensional space is calculated, and a total of n similarity scores are obtained. Among them, the target query feature is any query feature among the n query features. Subsequently, based on the n similarity scores and the n value features, the self-attention representation corresponding to the target query feature is calculated, and the self-attention representation is substituted into the deep learning model to reduce the learning difficulty of the deep learning model. In the embodiment of the present application, the similarity score between the key feature and the target query feature can be calculated in the high-dimensional space. Compared with the existing technology of directly calculating the inner product of key features and query features as the similarity score, calculation in high-dimensional space can capture more implicit correlations between features, thereby improving the training efficiency of deep learning models.

[0154] The embodiments of the present application can reduce the training cost of deep learning models.

[0155] To better implement the above method, an embodiment of the present application further provides a data processing device, which can be integrated into an electronic device, such as a terminal, a server, or the like. A terminal can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, or a personal computer (PC); a server can be a single server or a server cluster consisting of multiple servers. For example, in this embodiment, the device of the embodiment of the present application will be described in detail using the example of a data processing device being specifically integrated into a server.

[0156] For example, Figure 3 As shown, the device is applied to an electronic device, and the device includes:

[0157] A linear transformation unit 301 is configured to perform a linear transformation on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence; wherein the number of features in each of the query feature sequence, the key feature sequence, and the value feature sequence is n;

[0158] a high-dimensional space unit 302, configured to calculate, for each key feature in the key feature sequence, a similarity score of the key feature with a target query feature in a high-dimensional space; wherein the target query feature is any one of the query feature sequence;

[0159] a self-attention unit 303, configured to calculate a self-attention representation corresponding to the target query feature based on the n similarity scores and n value features in the value feature sequence, so as to substitute the self-attention representation into the deep learning model.

[0160] In an implementation, the high-dimensional space unit 302 comprises:

[0161] an inner product calculation subunit, configured to calculate an inner product of a mapping value of the key feature in the high-dimensional space and a mapping value of the target query feature in the high-dimensional space;

[0162] a score calculation subunit, configured to calculate the similarity score based on the inner product and a dimension value of the query feature sequence.

[0163] In an implementation, the inner product calculation subunit is specifically configured to substitute the key feature and the target query feature into a kernel function to obtain a function value of the kernel function, wherein the function value is the inner product of the mapping value of the key feature in the high-dimensional space and the mapping value of the target query feature in the high-dimensional space.

[0164] In an implementation, the inner product calculation subunit comprises:

[0165] a first intermediate subunit, configured to calculate a product of a transpose of the target query feature and the key feature to obtain a first intermediate result;

[0166] a first function subunit, configured to calculate the function value based on the first intermediate result and at least one hyperparameter.

[0167] In an implementation, the inner product calculation subunit comprises:

[0168] a second intermediate subunit, configured to calculate an absolute value of a difference between the target query feature and the key feature to obtain a second intermediate result;

[0169] a second function subunit, configured to calculate the function value based on the second intermediate result and at least one hyperparameter.

[0170] In an implementation, the kernel function is any one of a linear kernel function, a polynomial kernel function, a Laplacian kernel function, a sigmoid kernel function, and a Gaussian kernel function.

[0171] In an implementation, the score calculation subunit comprises:

[0172] a square root subunit configured to perform square root operation on the dimension value to obtain a square root result;

[0173] a ratio subunit configured to calculate a ratio of the inner product and the square root result, the ratio being the similarity score.

[0174] In an implementation, the self-attention unit 303 comprises:

[0175] a normalization subunit configured to perform normalization processing on n similarity scores to obtain n weight values; the n weight values correspond to n value features in the value feature sequence one by one;

[0176] a product subunit configured to calculate, for each value feature, a product of the value feature and a corresponding weight value;

[0177] a sum subunit configured to calculate a sum of n products, the sum being a self-attention representation corresponding to the target query feature.

[0178] In an implementation, the normalization subunit comprises:

[0179] an exponentiation subunit configured to calculate, for each similarity score, an exponentiation result of the similarity score;

[0180] an exponentiation sum subunit configured to calculate a sum of n exponentiation results of the similarity scores;

[0181] a weight subunit configured to calculate a ratio of each exponentiation result and the sum of the exponentiation results to obtain a corresponding weight value.

[0182] In an implementation, the linear transformation unit 301 comprises:

[0183] a query weight matrix subunit configured to transform the input sequence by using a query weight matrix to obtain the query feature sequence;

[0184] a key weight matrix subunit configured to transform the input sequence by using a key weight matrix to obtain the key feature sequence;

[0185] a value weight matrix subunit configured to transform the input sequence by using a value weight matrix to obtain the value feature sequence.

[0186] In implementation, each of the above units can be implemented as an independent entity, or can be combined as the same or several entities. The specific implementation of each of the above units can be referred to the method embodiments above, and will not be described here.

[0187] In the data processing method provided in the embodiment of the present application, the input sequence can be linearly transformed to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence. Among them, the query feature sequence includes n query features, the key feature sequence includes n key features, and the value feature sequence includes n value features. The n query features, n key features, and n value features all correspond one-to-one to the n features of the input sequence. For each of the n key features, the similarity score between the key feature and the target query feature in the high-dimensional space is calculated, and a total of n similarity scores are obtained. Among them, the target query feature is any query feature among the n query features. Subsequently, based on the n similarity scores and the n value features, the self-attention representation corresponding to the target query feature is calculated, and the self-attention representation is substituted into the deep learning model to reduce the learning difficulty of the deep learning model. In the embodiment of the present application, the similarity score between the key feature and the target query feature can be calculated in the high-dimensional space. Compared with the existing technology of directly calculating the inner product of key features and query features as the similarity score, calculation in high-dimensional space can capture more implicit correlations between features, thereby improving the training efficiency of deep learning models.

[0188] The embodiments of the present application can reduce the training cost of deep learning models.

[0189] The present application also provides an electronic device, which may be a terminal, a server, or the like. The terminal may be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a personal computer, or the like; the server may be a single server or a server cluster consisting of multiple servers, or the like.

[0190] In some embodiments, the data processing device may also be integrated into multiple electronic devices. For example, the data processing device may be integrated into multiple servers, and the data processing method of the present application may be implemented by multiple servers.

[0191] In this embodiment, the electronic device of this embodiment is an electronic device as an example for detailed description, for example, Figure 4 , which shows a schematic diagram of the structure of the electronic device involved in the embodiment of the present application, specifically:

[0192] The electronic device may include one or more processing core processors 401, one or more computer-readable storage media memories 402, a power supply 403, an input module 404, and a communication module 405. Those skilled in the art will appreciate that Figure 4The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0193] Processor 401 is the control center of the electronic device, connecting all components of the electronic device using various interfaces and circuits. It executes software programs and / or modules stored in memory 402 and accesses data stored in memory 402 to perform various functions and process data. In some embodiments, processor 401 may include one or more processing cores. In some embodiments, processor 401 may integrate an application processor and a modem processor, with the application processor primarily processing the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 401.

[0194] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0195] The electronic device also includes a power supply 403 for supplying power to various components. In some embodiments, the power supply 403 can be logically connected to the processor 401 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 403 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0196] The electronic device may further include an input module 404, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0197] The electronic device can also include a communication module 405, which in some embodiments can include a wireless module, through which the electronic device can perform short-range wireless transmission, thereby providing the user with wireless broadband Internet access. For example, the communication module 405 can be used to help the user send and receive emails, browse web pages, and access streaming media, etc.

[0198] Although not shown, the electronic device can also include a display unit, etc., which will not be described here. In particular, in the present embodiment, the processor 401 in the electronic device will load the executable file corresponding to the process of one or more application programs into the memory 402 according to the following instructions, and run the application program stored in the memory 402 by the processor 401, thereby realizing various functions, such as:

[0199] performing linear transformation on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence; wherein the number of features of the query feature sequence, the key feature sequence and the value feature sequence is n; for each key feature in the key feature sequence, calculate a similarity score between the key feature and a target query feature in a high-dimensional space; wherein the target query feature is any one of the query features in the query feature sequence; based on n similarity scores and n value features in the value feature sequence, calculate a self-attention representation corresponding to the target query feature, so as to substitute the self-attention representation into a deep learning model.

[0200] The specific implementation of each of the above operations can refer to the previous embodiments, which will not be described here.

[0201] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by instructions controlling relevant hardware, which can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0202] To this end, an embodiment of the present application provides a computer-readable storage medium, which stores a plurality of instructions. The instructions can be loaded by a processor to execute the steps in any data processing method provided by the embodiments of the present application. For example, the instructions can perform the following steps:

[0203] performing linear transformation on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence; wherein the number of features of the query feature sequence, the key feature sequence, and the value feature sequence is n; for each key feature in the key feature sequence, calculating a similarity score of the key feature and a target query feature in a high-dimensional space; wherein the target query feature is any one of the query features in the query feature sequence; and based on the n similarity scores and n value features in the value feature sequence, calculating a self-attention representation corresponding to the target query feature, so as to substitute the self-attention representation into a deep learning model.

[0204] The storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or the like.

[0205] According to an aspect of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method provided in any of the various optional implementations provided in the above embodiments.

[0206] Due to the instructions stored in the storage medium, the steps of any of the data processing methods provided in the embodiments of the present application can be executed, and thus the beneficial effects of any of the data processing methods provided in the embodiments of the present application can be achieved. Details are described in the above embodiments, and thus will not be described here.

[0207] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.

[0208] The above describes in detail the data processing method, device, electronic device and computer readable storage medium provided by the embodiments of the present application. The principles and implementation manners of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. Meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed. In conclusion, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A data processing method, characterized in that: Used to perform data processing on an input sequence, where the input sequence includes n features, where n is a positive integer; The method comprises: Performing a linear transformation on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence; wherein the number of features in the query feature sequence, the key feature sequence, and the value feature sequence is n; For each key feature in the key feature sequence, calculating a similarity score between the key feature and a target query feature in a high-dimensional space; wherein the target query feature is any query feature in the query feature sequence; Based on the n similarity scores and the n value features in the value feature sequence, a self-attention representation corresponding to the target query feature is calculated so as to substitute the self-attention representation into a deep learning model; the self-attention representation is used to reflect the degree of association between the feature corresponding to the target query feature in the input sequence and the n features of the input sequence.

2. The method according to claim 1, wherein The calculation of the similarity score between the key feature and the target query feature in the high-dimensional space includes: Calculating the inner product of the mapping value of the key feature in the high-dimensional space and the mapping value of the target query feature in the high-dimensional space; The similarity score is calculated based on the inner product and the dimension value of the query feature sequence.

3. The method according to claim 2, wherein The calculating the inner product of the mapping value of the key feature in the high-dimensional space and the mapping value of the target query feature in the high-dimensional space includes: Substituting the key feature and the target query feature into a kernel function to obtain a function value of the kernel function, wherein the function value is the inner product of a mapping value of the key feature in the high-dimensional space and a mapping value of the target query feature in the high-dimensional space.

4. The method according to claim 3, wherein Substituting the key feature and the target query feature into a kernel function to obtain a function value of the kernel function includes: Calculating the product of the transpose of the target query feature and the key feature to obtain a first intermediate result; The function value is calculated based on the first intermediate result and at least one hyperparameter.

5. The method according to claim 3, wherein Substituting the key feature and the target query feature into a kernel function to obtain a function value of the kernel function includes: Calculating the absolute value of the difference between the target query feature and the key feature to obtain the second intermediate result; The function value is calculated based on the second intermediate result and at least one hyperparameter.

6. The method according to claim 3, wherein The kernel function is any one of a linear kernel function, a polynomial kernel function, a Laplace kernel function, a sigmoid kernel function, and a Gaussian kernel function.

7. The method according to claim 2, wherein The calculating the similarity score based on the inner product and the dimension value of the query feature sequence includes: Performing a square root operation on the dimension value to obtain a square root result; A ratio of the inner product to the square root result is calculated, where the ratio is the similarity score.

8. The method according to claim 1, wherein The calculating, based on the n similarity scores and the n value features in the value feature sequence, a self-attention representation corresponding to the target query feature, includes: Normalizing the n similarity scores to obtain n weight values; the n weight values ​​correspond one-to-one to the n value features in the value feature sequence; For each of the value features, calculating the product of the value feature and the corresponding weight value; Calculate the sum of n products, where the sum is the self-attention representation corresponding to the target query feature.

9. The method according to claim 8, wherein The normalizing process is performed on the n similarity scores to obtain n weight values, including: For each of the similarity scores, calculating an exponential result of the similarity score; Calculating the sum of the exponential results of n similarity scores; The ratio of each indexation result to the sum of the indexation results is calculated to obtain a corresponding weight value.

10. The method according to claim 1, wherein The linear transformation of the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence includes: Converting the input sequence using a query weight matrix to obtain the query feature sequence; Converting the input sequence using a key weight matrix to obtain the key feature sequence; The input sequence is transformed using a value weight matrix to obtain the value feature sequence.

11. A data processing device, characterized in that: The apparatus is configured to process data of an input sequence, wherein the input sequence includes n features, where n is a positive integer; the apparatus comprises: a linear transformation unit, configured to perform a linear transformation on the input sequence to obtain a query feature sequence corresponding to the input sequence, a key feature sequence corresponding to the input sequence, and a value feature sequence corresponding to the input sequence; wherein the number of features of the query feature sequence, the key feature sequence, and the value feature sequence are all n; a high-dimensional space unit, configured to calculate, for each key feature in the key feature sequence, a similarity score between the key feature and a target query feature in the high-dimensional space; wherein the target query feature is any query feature in the query feature sequence; A self-attention unit is used to calculate the self-attention representation corresponding to the target query feature based on the n similarity scores and the n value features in the value feature sequence, so as to substitute the self-attention representation into the deep learning model.

12. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads instructions from the memory to execute the steps of the data processing method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the data processing method according to any one of claims 1 to 10.

14. A computer program product, characterized in that The method comprises a computer program / instruction, wherein when the computer program / instruction is executed by a processor, the steps of the data processing method according to any one of claims 1 to 10 are implemented.