Data processing method and device, electronic equipment and storage medium
By performing feature extraction and similarity matrix processing on the to-process data and reference data, and using self-multiply operation to obtain the attention matrix, the problem of high computational complexity in high-dimensional data processing is solved, and efficient data processing and model inference are achieved.
Patent Information
- Application Number
- CN202510124989.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-26
AI Technical Summary
The Softmax operation has high computational complexity when processing high-dimensional data, resulting in time delay and computing resource consumption problems.
By performing feature extraction of the to-be-processed data and reference data, the similarity matrix between the feature map is determined, and the attention matrix is obtained through self-multiply operation, which is replaced by the Softmax operation for weighting.
It effectively reduces the consumption of computing resources, improves the inference speed and efficiency of the model, and realizes the sparse processing of high-dimensional data.
Smart Images

Figure CN119992122A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of data processing technology and artificial intelligence technology. Specifically, the present application relates to a data processing method, device, electronic device, storage medium and product. Background Art
[0002] In deep learning models, especially those involving attention mechanisms, the Softmax operation is used to calculate the weight of the input data at each position. Softmax converts the input data into a corresponding probability distribution, thereby assigning a weight to each position. These weights reflect the relative importance of each element in the input data. In this way, the Softmax operation achieves sparse processing of the input data, allowing the model to focus on the parts that are more relevant to the current input, thereby improving the expressiveness and prediction accuracy.
[0003] However, the Softmax operation needs to perform exponential operations on each element, and the computational complexity is high. In computationally intensive models such as the Stable Diffusion model, the dimension of the input data of the corresponding Softmax operation is usually large. As the input feature dimension increases, the amount of Softmax calculation increases exponentially. Therefore, there are significant time delays and computing resource consumption problems, which in turn affects the reasoning speed and efficiency of the model. Summary of the invention
[0004] The embodiments of the present application provide a data processing method, device, electronic device, computer-readable storage medium, and computer program product, which can solve the problems of computational complexity and high resource consumption of the Softmax operation when processing high-dimensional data. The technical solution is as follows: According to a first aspect of an embodiment of the present application, a data processing method is provided, which is executed by a processor, and the method includes: By performing feature extraction on the data to be processed and the reference data used to perform target processing on the data to be processed, a first feature map of the data to be processed and a second feature map of the parameter data are obtained, wherein the data to be processed and the reference data are the same or different; wherein the sizes of the first feature map and the second feature map are , h represents the number of eigenvectors included in the feature map, and m represents the number of eigenvalues included in each eigenvector in the feature map; Determine a first similarity matrix between the first feature map and the second feature map, wherein an element value of an element in an i-th row and a j-th column in the first similarity matrix represents a similarity between an i-th eigenvector in the first feature map and a j-th eigenvector in the second feature map; Obtaining a similarity threshold, subtracting the elements in the first similarity matrix from the similarity threshold to obtain a second similarity matrix, and setting the elements in the second similarity matrix that are less than 0 to 0 to obtain a third similarity matrix; Obtaining an attention matrix between the to-be-processed data and the reference data by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix; The second feature map is weighted based on the attention matrix to obtain a weighted feature map, and target data after the target operation is performed on the data to be processed is obtained according to the weighted feature map.
[0005] According to a second aspect of an embodiment of the present application, a data processing device is provided, the device comprising: The feature extraction module is used to extract features from the data to be processed and the reference data used to perform target processing on the data to be processed, respectively, to obtain a first feature map of the data to be processed and a second feature map of the parameter data, wherein the data to be processed and the reference data are the same or different; wherein the sizes of the first feature map and the second feature map are , h represents the number of eigenvectors included in the feature map, and m represents the number of eigenvalues included in each eigenvector in the feature map; A similarity determination module, configured to determine a first similarity matrix between the first feature map and the second feature map, wherein the element value of the element in the i-th row and the j-th column in the first similarity matrix represents the similarity between the i-th eigenvector in the first feature map and the j-th eigenvector in the second feature map; A similarity threshold obtaining module is used to obtain a similarity threshold, subtract the elements in the first similarity matrix from the similarity threshold to obtain a second similarity matrix, and set the elements in the second similarity matrix that are less than 0 to 0 to obtain a third similarity matrix; An attention matrix obtaining module is used to obtain an attention matrix between the to-be-processed data and the reference data by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix; A weighting module is used to weight the second feature map based on the attention matrix to obtain a weighted feature map, and obtain target data after the target operation is performed on the data to be processed according to the weighted feature map.
[0006] According to a third aspect of an embodiment of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory, and when the processor executes the program, the steps of the method provided in the first aspect are implemented.
[0007] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method provided in the first aspect are implemented.
[0008] According to the fifth aspect of the embodiment of the present application, a computer program product is provided, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. When a processor of a computer device reads the computer instructions from the computer-readable storage medium, the processor executes the computer instructions, so that the computer device executes the steps of implementing the method provided in the first aspect.
[0009] The beneficial effects of the technical solution provided by the embodiment of the present application are: By extracting features from the data to be processed and the reference data, a first feature map of the data to be processed and a second feature map of the reference data are obtained, wherein the data to be processed and the reference data may be the same or different, so that the method provided in the embodiment of the present application can be applied to the self-attention mechanism and the cross-attention mechanism. By determining the first similarity matrix between the first feature map and the second feature map, the degree of association between the data to be processed and the reference data in the feature space is obtained; by using a similarity threshold, the elements in the first similarity matrix are subtracted from the similarity threshold to obtain a second similarity matrix, and the elements less than 0 in the second similarity matrix are set to 0 to obtain a third similarity matrix, so that the obtained third similarity matrix retains feature pairs with high similarity; by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix, the sparse processing of the third similarity matrix is realized, and the attention matrix between the data to be processed and the reference data is obtained; based on the attention matrix, the second feature map is weighted to obtain a weighted feature map, so that the model can predict and infer based on the weighted feature map to obtain the target data after the target operation of the data to be processed. In an embodiment of the present application, a similarity threshold is obtained to convert the first similarity matrix into a third similarity matrix with more 0 elements. By performing a self-multiplication operation on each element in the third similarity matrix, elements with smaller values are made closer to 0, thereby achieving sparseness of the first similarity matrix. Compared with the exponential operation in the Softmax operation, the self-multiplication operation of the elements has lower computational complexity and better numerical stability, thereby effectively reducing the consumption of computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.
[0011] Figure 1 A schematic diagram of the data processing system architecture provided in an embodiment of the present application; Figure 2 A flowchart of a data processing method provided in an embodiment of the present application; Figure 3 A schematic diagram of determining a third similarity matrix provided in an embodiment of the present application; Figure 4 A schematic diagram of obtaining an attention matrix provided in an embodiment of the present application; Figure 5 A comparison diagram of data distribution after sparseness provided in an embodiment of the present application; Figure 6 A flowchart of a data processing method provided in an embodiment of the present application; Figure 7 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application; Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0012] The embodiments of the present application are described below in conjunction with the drawings in the present application. It should be understood that the implementation methods described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions of the embodiments of the present application.
[0013] It will be understood by those skilled in the art that, unless specifically stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application refer to that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation as other features, information, data, steps, operations, elements, components and / or combinations thereof supported by the technical field. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element may be directly connected or coupled to the other element, or it may refer to that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" may be implemented as "A", or as "B", or as "A and B".
[0014] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0015] First, several terms involved in this application are introduced and explained: Stable Diffusion is a type of diffusion model. It regards the data distribution as the steady state of a diffusion process and generates data using the reverse diffusion process. It is currently commonly used in image generation, text generation, audio generation, etc. The UNet network is the most commonly used basic network structure in Stable Diffusion. The UNet network can effectively capture multi-scale features and maintain high-resolution details. In the UNet network, the implementation based on the attention mechanism helps to improve the performance of the model in the feature extraction and generation process, especially in image generation tasks, which can focus on key information more accurately to generate high-quality images.
[0016] Attention Mechanism is a data processing method in machine learning, which is widely used in various types of machine learning tasks such as natural language processing, image recognition and speech recognition. The degree of attention paid to different information by the attention mechanism is reflected by the weight. The attention mechanism can be regarded as a multilayer perceptron (MLP) composed of a query matrix (Query), a key matrix (key) and a weighted average. The idea of attention is similar to addressing. Given an element Query in the target Target, the weight coefficient of each Key corresponding to the Value is obtained by calculating the similarity or correlation between the Query and each Key, and then the Value is weighted summed to obtain the final Attention value. Therefore, in essence, the Attention mechanism is a weighted sum of the Value value, and the Query and Key are used to calculate the weight coefficient of the corresponding Value.
[0017] The data processing method, device, electronic device, computer-readable storage medium, and computer program product provided in this application are intended to solve the above technical problems in the prior art.
[0018] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application. It should be noted that the following embodiments can refer to, draw on or combine with each other, and the same terms, similar features and similar implementation steps in different embodiments will not be described repeatedly.
[0019] Figure 1 A schematic diagram of the data processing system architecture provided in the embodiment of the present application is shown in FIG. Figure 1 As shown, the data processing system includes a terminal 101 and a server 102. The server 102 may be deployed with any neural network model including an attention mechanism module, and the server may run the neural network model based on a processor NPU.
[0020] The terminal 101 can send the input data required by the neural network model to the server 102, so that the processor processes the input data based on the neural network model. When the neural network model processes the input data, it can obtain the data to be processed and the reference data corresponding to the attention mechanism module according to the input data. The data to be processed and the reference data can be the same or different. When the data to be processed and the reference data are the same, the attention mechanism module is a self-attention mechanism module, and when the data to be processed and the reference data are different, the attention mechanism module is a cross-attention mechanism module.
[0021] The server 102 performs the following operations through the attention mechanism module: by extracting features from the data to be processed and the reference data, a first feature map of the data to be processed and a second feature map of the reference data are obtained; further, by similarity analysis, a first similarity matrix between the first feature map and the second feature map is obtained; based on a preset similarity threshold, the similarity thresholds of each element in the first similarity matrix are subtracted to obtain a second similarity matrix; elements less than 0 in the second similarity matrix are set to 0 to obtain a third similarity matrix; the element values of each element in the third similarity matrix are self-multiplied to obtain an attention matrix of the image and text to be processed; the second feature map is weighted based on the attention matrix to obtain a weighted feature map; and according to the weighted feature map. The attention mechanism model in the embodiment of the present application replaces the Softmax operation in the attention mechanism of the prior art by presetting a similarity threshold and performing a self-multiplication operation on the elements, and further improves the processing efficiency of the NPU on the basis of achieving the sparseness of the first similarity matrix.
[0022] The neural network model can perform subsequent early operations based on the weighted feature map obtained by the attention mechanism module to obtain the final target data.
[0023] Taking a specific application scenario as an example, the terminal 101 can be any terminal running the Wenshengtu application / software, and the terminal can be a user's mobile phone or a computer, etc. The user terminal 101 can establish a communication connection with the Wenshengtu application / software server 102 in a wired or wireless manner, and the user can use various services provided by the application based on the user interface of the application displayed on the terminal 101.
[0024] Text is input into the user interface of the application displayed on the terminal 101, and the terminal 101 sends the text to the server 102. A stable diffusion model can be deployed in the server 102. The server runs the stable diffusion model based on the processor NPU, and generates an initial noise image based on the acquired text. The initial noise image is used as the image to be processed. The image to be processed and the text are input to the cross-attention mechanism module. The above operations are performed based on the cross-attention mechanism module to obtain a weighted feature map. The stable diffusion model performs noise prediction based on the obtained weighted feature map, and generates an image corresponding to the content described by the text by performing noise prediction on the image to be processed.
[0025] The server 102 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal 102 (also referred to as a user terminal or user device) may be a smart phone, a tablet computer, a laptop computer, a desktop computer, an intelligent voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a vehicle terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in this disclosure.
[0026] Optionally, the data processing method provided in the embodiment of the present application can be implemented as an independent application or a functional module / plug-in of an application. For example, the application can be an application with data processing function.
[0027] The present application provides a data processing method, such as Figure 2 As shown, the method includes S201-S205.
[0028] S201 , extracting features from the data to be processed and reference data used for performing target processing on the data to be processed, respectively, to obtain a first feature map of the data to be processed and a second feature map of the parameter data.
[0029] The data to be processed and the reference data are the same or different; the sizes of the first feature map and the second feature map are , h represents the number of eigenvectors included in the feature map, and m represents the number of eigenvalues included in each eigenvector in the feature map.
[0030] In the embodiment of the present application, the data to be processed and the reference data may be data such as images or texts, or feature maps obtained by processing in the intermediate process of the neural network model. The target processing may be determined according to actual needs. For example, in a text map, the target processing may be denoising of a noisy image. In text translation, the target processing may be semantic understanding of text data to obtain a translation result. In image classification, the target processing may be determining the category to which the image belongs.
[0031] When the data to be processed is the same as the reference data, the embodiment of the present application determines the attention weights of each feature of the data to be processed relative to its own features based on the self-attention mechanism by executing steps S201-S205, and obtains the target data after the target operation is performed on the data to be processed. When the data to be processed is different from the reference data, the embodiment of the present application determines the attention weights of each feature of the data to be processed relative to each feature of the reference data based on the cross-attention mechanism by executing steps S201-S205, and obtains the target data after the target operation is performed on the data to be processed.
[0032] In the embodiment of the present application, feature extraction can be performed on the processed data and the reference data respectively by the feature extraction module. Optionally, feature extraction can be performed on the processed data by the first feature extraction module to obtain a first feature map of the processed data, and feature extraction can be performed on the reference data by the second feature extraction module to obtain a second feature map of the reference data. When the processed data and the reference data are feature maps, the feature map extraction module can be a linear transformation module.
[0033] Among them, the size of the first feature map and the second feature is , where h is used to represent the number of feature vectors included in the feature map, and m is used to represent the number of eigenvalues included in each feature vector, that is, the dimension of the feature vector.
[0034] As an optional embodiment, feature extraction can be performed on the processed data and the reference data respectively to obtain feature maps corresponding to the processed data and the reference data, wherein the feature maps corresponding to the processed data and the reference data are then matrix transformed so that the feature maps after the matrix transformation have the same size, and the feature map of the processed data after the matrix transformation is used as the first feature map, and the feature map of the reference data after the matrix transformation is used as the second feature map.
[0035] In the embodiment of the present application, the first feature map can be used as a query matrix, which can be called Qurey or simply Q, and the second feature map can be used as a key matrix, which can be called Key or simply K. In the embodiment of the present application, by performing feature extraction on the data to be processed and the reference data, the feature vectors corresponding to each feature of the data to be processed and the feature vectors corresponding to each feature of the reference data are obtained, so as to determine the attention matrix of the data to be processed and the reference data according to each feature vector of the data to be processed and each feature vector of the reference data.
[0036] S202: Determine a first similarity matrix between the first feature map and the second feature map.
[0037] The element value of the element in the i-th row and j-th column in the first similarity matrix represents the similarity between the i-th eigenvector in the first feature map and the j-th eigenvector in the second feature map.
[0038] In the embodiment of the present application, the first similarity matrix between the first feature map and the second feature map can be obtained by performing an inner product operation, that is, a matrix multiplication operation, on the transposed first feature map and the second feature map. The first matrix of To obtain the first similarity matrix between the first feature map and the second feature map, the second feature map is transposed to obtain The third matrix of the first matrix can be obtained, so that the inner product operation can be performed on the first matrix and the third matrix, and the result of the inner product operation on the first matrix and the third matrix is used as the first similarity matrix.
[0039] It should be noted that the element value of the element in the i-th row and j-th column of the first similarity matrix is obtained by multiplying and adding the element values of each element in the i-th row of the first matrix and the original values of each element in the j-th column of the third matrix, that is, the result of the inner product operation of the vector represented by the i-th row of the first matrix and the vector represented by the j-th column of the third matrix. The inner product operation of the two vectors is used to calculate the similarity of the two vectors, and the vector represented by the i-th row of the first matrix is the i-th eigenvector of the first feature map, and the vector represented by the j-th column of the third matrix is the j-th eigenvector of the second feature map. Therefore, the first similarity matrix of the first feature map and the second feature map can be obtained by performing the inner product operation on the transpose of the first feature map and the second feature map.
[0040] As an optional embodiment, the cosine similarity, Euclidean distance or Manhattan distance between the vector represented by the i-th row of the first matrix and the vector represented by the j-th column of the third matrix can be determined as the element value of the element in the i-th row and j-th column of the first similarity matrix.
[0041] As an optional embodiment, the fourth matrix can be obtained by performing an inner product operation on the transpose of the first feature map and the second feature map, and each element in the fourth matrix can be divided by the scaling factor to obtain a matrix after scaling the fourth matrix as the first similarity matrix. Considering that when the dimension of each feature in the first feature map and the second feature map is large, their inner product may be very large, and the effect of subsequent sparseness is poor, the dimension of each feature can be As the scaling factor, the scaled matrix is used as the first similarity matrix.
[0042] In an embodiment of the present application, a first similarity matrix is determined to obtain the similarity between different features in the first feature map and the second feature map, so as to obtain an attention matrix of the first feature map and the second feature map through subsequent sparse processing.
[0043] S203 , obtaining a similarity threshold, subtracting the elements in the first similarity matrix from the similarity threshold to obtain a second similarity matrix, and setting the elements in the second similarity matrix that are less than 0 to 0 to obtain a third similarity matrix.
[0044] In an embodiment of the present application, the similarity threshold may be a preset value, and the phase velocity threshold may be used as an adjustable hyperparameter in the neural network model. The similarity threshold may be obtained by obtaining the value preset by the technician for the hyperparameter.
[0045] As an optional embodiment, the similarity threshold may also be obtained according to element values of elements of the first similarity matrix, and the average value of each element in the first similarity matrix may be used as the similarity threshold.
[0046] In an embodiment of the present application, the element values of the elements in the first similarity matrix are subtracted from the similarity threshold to obtain a second similarity matrix, and the element values of the elements less than 0 in the second similarity matrix are set to 0 to obtain a third similarity matrix. The similarity threshold is used to divide the elements in the first similarity matrix into elements with larger values and elements with smaller values. The elements with smaller values are less than 0 after subtracting the similarity threshold. By setting the elements less than 0 in the second similarity matrix to 0, the elements with smaller values in the first similarity matrix are set to 0. The amount of calculation can be reduced by reducing the number of 0 elements in the similarity matrix, and preliminary sparse processing is achieved. The amount of calculation is reduced by limiting the focus to only a few positions with higher similarity, and calculation of unimportant elements is avoided.
[0047] As an optional embodiment of the present application, the step of subtracting the elements in the first similarity matrix from the similarity threshold to obtain the second similarity matrix includes: Obtaining a similarity threshold corresponding to each row in the first similarity matrix; For each row of the first similarity matrix, subtract the element value of each element in the row from the similarity threshold corresponding to the row to obtain a second similarity matrix; Wherein, for each row of the first similarity matrix, the similarity threshold corresponding to the row is determined in the following manner: Determine the maximum element value in the row; A preset bias coefficient is obtained, and the maximum element value in the row is adjusted by using the bias coefficient to obtain a similarity threshold corresponding to the row, wherein the bias coefficient is a positive number greater than or equal to 0 and less than 1.
[0048] In an embodiment of the present application, for a feature vector in the first feature map, the numerical range of the similarity between the feature vector and each feature vector in the second feature map may be different from the numerical range of the similarity between another feature vector in the first feature map and each feature vector in the second feature map. Therefore, in order to accurately distinguish the numerical size of the similarity of each feature vector in the first feature map with respect to each feature vector in the second feature map, a corresponding similarity threshold is set for each row of the first similarity matrix.
[0049] In the embodiment of the present application, for each row of the first pixel matrix, the maximum element value max in the row is obtained, the preset bias coefficient α is obtained, and As the similarity threshold of the row, the value range of α is [0,1), that is, α is a positive number greater than or equal to 0 and less than 1. Figure 3 As shown, Figure 3 A schematic diagram for determining a third similarity matrix is provided for an embodiment of the present application, wherein the first similarity matrix is a 4×4 matrix, for each row in the matrix, the maximum value of the row is determined, and the bias coefficient is 0.5 to obtain the similarity threshold of each row, and each element in the first similarity matrix is subtracted from the similarity threshold of the corresponding row to obtain the second similarity matrix.
[0050] As an optional embodiment, the element values less than 0 in the second similarity matrix can be set to 0 through the ReLU activation function to obtain a third similarity matrix. It can be seen from the third similarity matrix that the embodiment of the present application achieves the sparseness of the first similarity matrix.
[0051] As an optional embodiment, the specific value of α can be determined according to the operation of the target processing. For example, the data processing method provided in the embodiment of the present application can be applied to a stable diffusion model. In this case, the target processing is denoising processing, and the value range of α can be (0.9, 0.95). The value range is obtained based on specific experiments, and the target data obtained based on the bias coefficient within this value range has better effect.
[0052] By As the similarity threshold of the row, the other element values in the row are subtracted from the similarity threshold. If the subtraction result is a negative number, it means that the similarity represented by the element value is small. The similarity can be set to 0 to achieve preliminary sparsification of the first similarity matrix. If the subtraction result is a positive number, it means that the similarity represented by the element value is large. The subtraction result corresponding to the element value can be included to obtain a third similarity matrix. There may still be some smaller element values in the third similarity matrix, and the third similarity matrix can be further sparsed through subsequent operations.
[0053] S204. Obtain an attention matrix between the data to be processed and the reference data by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix.
[0054] In an embodiment of the present application, a self-multiplication operation is performed on the element value of each element in the third similarity matrix element, wherein one self-multiplication operation is a square operation, and two self-multiplication operations are cubic operations. The exponential operation in the Softmax operation is replaced by at least one self-multiplication operation, which reduces the complexity of the algorithm, improves the efficiency of the attention mechanism processing, and obtains the attention matrix between the data to be processed and the reference data.
[0055] In the embodiment of the present application, by performing a self-multiplication operation on each element in the third similarity matrix, elements less than 1 become smaller, thereby further sparseening the third similarity matrix.
[0056] As an optional embodiment, at least one self-multiplication operation is a square operation. The square operation is more user-friendly and efficient for hardware such as the end-side device NPU and TPU that are optimized for multiplication and addition operations. That is, the end-side device NPU and TPU are more efficient when processing square operations than processing exponential operations.
[0057] As an optional embodiment of the present invention, an attention matrix between the to-be-processed data and the reference data is obtained by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix, including: Obtaining a fourth similarity matrix by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix; A normalization coefficient is obtained, and a normalization operation is performed on the fourth similarity matrix according to the normalization coefficient to obtain the attention matrix.
[0058] In an embodiment of the present application, after each element in the third similarity matrix is self-multiplied at least once, elements less than 1 will become smaller, and elements greater than 1 will become larger. In order to scale the result of the element after self-multiplication at least once to between 0 and 1, the result of self-multiplication of each element in the third similarity matrix at least once is used as the fourth similarity matrix, and the elements in the fourth similarity matrix are normalized.
[0059] It should be understood that the embodiment of the present application processes the first similarity matrix to obtain an attention matrix, which is used to replace the Softmax operation in the attention mechanism. The value range of the elements in the matrix after the Softmax operation is 0-1. In order not to change the value range of the input data corresponding to the subsequent operations after the Softmax operation in the attention mechanism, the elements in the fourth similarity matrix are scaled to between 0-1, that is, the fourth similarity matrix is normalized, and the normalized matrix is used as the attention matrix.
[0060] Among them, the value range of the normalization coefficient β can be (0,1), which is used to reduce the element values of elements greater than 1 in the fourth similarity matrix to between 0 and 1. The normalization coefficient can be used as a hyperparameter in the neural network model. The technician can set the hyperparameter according to actual needs. When running the neural network model, the value set to β will be read and normalized based on the value. Optionally, the value of the normalization coefficient β can be 0.01. Experiments have shown that the target data generated based on the normalization coefficient with this value has better effect.
[0061] As an optional embodiment, when performing a normalization operation on the fourth similarity matrix, each element in the fourth similarity matrix may be multiplied by the normalization coefficient to obtain a normalized result of each element, thereby obtaining an attention matrix.
[0062] As an optional embodiment of the present application, performing a normalization operation on the fourth similarity matrix includes: For each row of the fourth similarity matrix, determine a normalization coefficient corresponding to the row according to an element value of at least one element in the row; For each element of each row of the fourth similarity matrix, the element is adjusted to a range of 0-1 according to a normalization coefficient corresponding to the row.
[0063] In an embodiment of the present application, the numerical ranges of the element values of the elements in each row in the fourth similarity matrix may be different. By setting a normalization coefficient for each row, the normalized matrix can reflect the difference in similarity of a certain eigenvector in the first feature map with respect to each eigenvector in the second feature map.
[0064] For each row in the fourth similarity matrix, the normalization coefficient of the row is determined according to the element value of at least one element in the row. Specifically, the maximum element value (Square max) of the row can be obtained, and the preset adjustment coefficient γ can be obtained, and the ratio of γ to Square max is used as the normalization coefficient of the row, where the value of γ can be 0.1. For each element in each row of the fourth similarity matrix, the element is multiplied by the normalization coefficient of the row so that the element can be adjusted to a range of 0-1.
[0065] like Figure 4 As shown, Figure 4 A schematic diagram of obtaining an attention matrix is provided for an embodiment of the present application. For a third similarity matrix, the element value of each element in the third similarity matrix is squared to obtain a fourth similarity matrix. For each row in the fourth similarity matrix, the maximum value of the row is determined, and the adjustment coefficient is taken as 0.1 to obtain the normalization coefficient of each row. For each element of the fourth similarity matrix, the element value of the element is multiplied by the normalization coefficient of the corresponding row (four significant digits are retained for convenience of representation) to obtain the attention matrix.
[0066] In the embodiment of the present application, a Softmax approximation operation is implemented through steps S203 and S204, and the approximation operation realizes the sparseness of the first similarity matrix, and the processor uses the approximation operation for sparse processing, which is more efficient than the sparse processing through Softmax. Furthermore, for processors such as end-side NPU or TPU based on optimized multiplication and addition, the approximation operation can greatly improve hardware utilization and increase computing speed. In addition, the approximation scheme supports processors with only fixed-point operations, such as Zhouyi Z series and X1 NPU processors, where fixed-point operations refer to that the representation of numbers is fixed during operations, that is, the precision and range of all numerical values are predetermined.
[0067] For an M × N Feature Map A1, the calculation steps of the Softmax operation are: S1. For each row of matrix A1, find the maximum value of N data; S2. For each data in matrix A1, subtract the maximum value of the corresponding row to obtain matrix A2; S3, for each data of matrix A2 , find , and get the matrix A3; S4. For each row of matrix A3, find the sum of all the data in the row; S5. For each data in the matrix A3, divide the data by the sum of all the data in this row.
[0068] For an M×N Feature Map B1, the calculation steps of the approximate operation of the Softmax operation in the embodiment of the present application are: S1. For each row of matrix B1, find the maximum value of N data; S2, obtain the bias coefficient α∈[0,1); S3, for each data in the matrix B1, subtract the maximum value of the row from the data and multiply by α to obtain the matrix B2; S4, for each data in the matrix B2, perform the ReLU activation function on the data to obtain the matrix B3; S5, perform a square operation on each data in B3 to obtain a matrix B4; S6. For each row of matrix B4, find the maximum value of N data; S7. For each data in B4, multiply the data by a normalization coefficient β, β=0.1 / (maximum value of the row).
[0069] The core of this approximate operation is to replace the sparse process Exp operation with a square operation, and the square multiplication is more friendly and efficient for end-side NPU, TPU and other hardware optimized for multiplication and addition operations. For fine-tuned diffusion models such as InstaFlow, images can be generated in one step. This type of model avoids the problem of large differences in data distribution for the same operator caused by multiple iterations. When processing, steps S1, S2, S3, S6, and S7 can be integrated into the offline quantization process and saved in advance to save time traversing each row of data during inference. Therefore, the process of approximate operation is simplified to four steps: S1. For matrix B1, get the value of each row in B1. , for each data in matrix B1, subtract the maximum value of the row from the data and multiply it by α to obtain matrix B2; S2, perform the ReLU activation function on each data in the matrix B2 to obtain the matrix B3; S3, perform self-multiplication, that is, square operation, on each data in matrix B3 to obtain matrix B4; S4. For the matrix B4, obtain β for each row in the matrix B4, β=0.1 / (maximum value of the row), and for each data in the matrix B4, multiply the data by the corresponding β.
[0070] The principle of the Softmax approximation operation running faster on the end-side platform in the embodiment of the present application is mainly to reduce the number of unit calculations. Taking Zhouyi NPU as an example, the floating-point operation process of the Softmax operation is more complicated than the additional operation process of the approximate operation.
[0071] In the S3 step of the Softmax operation, each element of the matrix needs to be calculate , exponential calculations are very complex, especially in hardware implementation, and require complex computational processes. For example: 1) Traverse element by element to limit the size extreme value: avoid input that is too large or too small to ensure calculation stability; 2) Look up the table and calculate the approximate value: Implemented by looking up the table , and perform two element-by-element subtractions and corrections; 3) Multiple element-by-element multiplications and additions: To further approximate the exponent value, at least two multiplications and one addition are required.
[0072] 4) Element-by-element multiplication: used to adjust the final result.
[0073] It can be seen that exponential operations require multiple table lookups, multiplications, and additions, involving complex unit calculations in hardware.
[0074] In the approximate operation, S4 and S5 use ReLU activation function and square operation instead of traditional exponential operation. Specific calculation categories include: 1) Calculate activation values by one table lookup: Implement ReLU operation (set negative values to zero), which is computationally simple and hardware-friendly.
[0075] 2) Calculate the square by multiplying each element at a time: The square operation only requires one multiplication unit to complete, which is much simpler than the exponential operation.
[0076] In contrast, the computational complexity of the square approximation algorithm is significantly reduced because it avoids complex exponential calculations and is more efficient in hardware implementation, occupying fewer resources.
[0077] The Softmax approximation operation of the embodiment of the present application can be completed using fewer unit calculations in the Zhouyi NPU. Similarly, in other end-side computing architectures, if the exponential calculation requires more complex unit calculations than the square calculation, then the Softmax operation can be replaced based on the approximate operation in the end-side computing architecture to effectively improve the computing performance of the end-side computing architecture.
[0078] To compare the difference between the Softmax approximation operation and the Softmax operation in the embodiment of the present application in the effect of sparse data, as shown in FIG. Figure 5 As shown, Figure 5A data distribution comparison diagram after sparseness is provided in an embodiment of the present application, wherein the original data data distribution diagram is used to represent the data distribution of a normally distributed random number sequence with a data volume of 2048. The random number sequence is used as the first similarity matrix and is processed based on the Softmax operation and the approximate operation in the embodiment of the present application, respectively. The data distribution obtained after the Softmax operation and the data distribution obtained after the approximate operation are as follows: Figure 5 As shown. Through analysis, the cosine similarity of the data distribution obtained by the Softmax operation and the approximate operation is measured to be 75%, and the mean square error MSE is 2.4e-4. The approximate operation effectively sparses the data and highlights the features near the maximum value. In terms of performance, the approximate operation in the Zhouyi NPU simulator has about 40% fewer instructions than the Softmax implementation.
[0079] This approximation has the following advantages: (1) Fast inference speed: The approximate algorithm reduces the computational complexity of the Softmax operator, effectively improving the running speed while ensuring the model effect. (2) Wide scope of application: Not limited to a specific end-side chip architecture, the approximate algorithm that reduces complexity is universal (3) Easy to implement: Unlike other Softmax alternatives that require re-fine-tuning and model training, this approximate algorithm can be directly tuned and replaced on the inference side.
[0080] S205. Weight the second feature map based on the attention matrix to obtain a weighted feature map, and obtain target data after the target operation is performed on the data to be processed according to the weighted feature map.
[0081] In an embodiment of the present application, the second feature map can also be used as a value matrix in the attention mechanism, called Value or V for short. By performing an inner product operation, that is, a matrix multiplication operation, on the second feature map as a value matrix and the weight matrix in the attention mechanism, a feature map required for further processing can be generated.
[0082] As an optional embodiment of the present application, the second feature map includes a third feature map obtained by extracting features from the reference data through the first module, and a fourth feature map obtained by extracting features from the reference data through the second module; Determining a first similarity matrix between the first feature map and the second feature map includes: Perform an inner product operation on the transpose of the first feature map and the second feature map to obtain a first similarity matrix; The second feature map is weighted based on the attention matrix to obtain a weighted feature map, including: The fourth feature map is weighted by the attention matrix to obtain a weighted feature map.
[0083] In the embodiment of the present application, when extracting features from the reference data, two different feature extraction modules can be used to obtain a third feature map and a fourth feature map. The first module can be a feature extraction module for obtaining a key matrix, and the second module can be a feature extraction module for obtaining a value matrix. For the data to be processed, the third module can be used to extract features from the data to be processed to obtain a query matrix.
[0084] In an embodiment of the present application, an inner product operation is performed on the transpose of the first feature map and the second feature map. The inner product operation, that is, matrix multiplication operation, is performed on the transpose of the first feature map and the third feature map to obtain a first similarity matrix to obtain the correlation or similarity between the query matrix and the key matrix.
[0085] In an embodiment of the present application, the fourth feature map can be weighted by the attention matrix as a weighted feature map. The weighted sum operation is performed using the attention matrix and the value matrix to obtain a weighted feature map, which can better focus on important areas or features in the reference data, thereby improving the effectiveness of the target operation.
[0086] As an optional embodiment of the present application, the data to be processed is an image to be processed, the reference data is a content description text corresponding to the image to be processed, the target processing is denoising processing, and obtaining target data after the target operation is performed on the data to be processed according to the weighted feature map includes: Predicting a first noise in the image to be processed according to the weighted feature map; According to the first noise, the image to be processed is denoised to obtain a denoised target image.
[0087] In the embodiment of the present application, the data to be processed may be a noise image generated in a stable diffusion model, the reference data may be a content description text, and the target operation is a denoising operation. In the stable diffusion model, the noise image is gradually denoised by guiding the content description text to generate a target image, the content of which is the content described by the content description text.
[0088] Specifically, the stable diffusion model includes a Unet network, which denoises the noisy image through a downsampling structure and an upsampling structure. The Unet network includes a Transformer module, and the most important part of the Transformer module is a cross-attention module. The input of the cross-attention mechanism module is a noisy image and text. The cross-attention mechanism module uses the noisy image and text as the image to be processed and the content description text, respectively. Through steps S201-S204 in the embodiment of the present application, the attention matrix of the image to be processed and the content description text is obtained, and the attention matrix and the fourth feature map of the content description text are weighted to obtain a weighted feature map. The Unet network will predict the first noise in the current image to be processed based on the weighted feature map to obtain a prediction result of the first noise.
[0089] After obtaining the prediction result in the current image to be processed, the noise is removed by using the reverse diffusion process according to the prediction result of the first noise to generate a clearer image. The Unet network can repeatedly perform the denoising process until a target image with a clarity greater than a preset value is obtained.
[0090] In the embodiment of the present application, after denoising is completed, a denoised target image is obtained, and its content and quality are significantly affected by the weighted feature map, that is, the weighted feature map helps the model focus on important image areas related to the text description.
[0091] In the embodiment of the present application, the approximate operation of Softmax is applied to the stable diffusion model, which improves the efficiency of the stable diffusion model in generating the target image. The stable diffusion model Stable Diffusion 1.5 is used as a model for verification, and the positions of the five Softmax operations with the largest amount of data to be processed in the model are determined. The approximate operation is used to replace any one of the five Softmax operations to obtain the generated target image. The approximate operation is used to replace any two of the five Softmax operations to obtain the generated target image. The experimental results show that the effect of the target image is similar to that of the Softmax operation, and using the approximate operation to replace Softmax can greatly improve the inference speed of the model on the end-side computing platform.
[0092] In order to more clearly illustrate the process of the data processing method provided in the embodiment of the present application, a complete embodiment of the data processing method of the present application is now given, please refer to Figure 6 , Figure 6 A flowchart of a data processing method provided in an embodiment of the present application includes steps S601-S610.
[0093] S601, extracting features from the data to be processed and the reference data respectively, to obtain a first feature map of the data to be processed and a third feature map and a fourth feature map of the parameter data; S602, performing an inner product operation on the first feature map and the third feature map to obtain a first similarity matrix; S603, obtaining a similarity threshold corresponding to each row in the first similarity matrix according to the maximum element value and the bias coefficient of each row in the first similarity matrix; S604, subtract each element value in the first similarity matrix from the similarity threshold of the corresponding row to obtain a second similarity matrix; S605, setting the elements in the second similarity matrix that are less than 0 to 0 to obtain a third similarity matrix; S606, performing a square operation on each element in the third similarity matrix to obtain a fourth similarity matrix; S607: Obtain a normalization coefficient corresponding to each row of the fourth similarity matrix according to the maximum element value and the adjustment coefficient of each row of the fourth similarity matrix; S608, multiplying each element in the fourth similarity matrix by the normalization coefficient of the corresponding row to obtain an attention matrix; S609, weighting the fourth feature map by using the attention matrix to obtain a weighted feature map; S610: Obtain target data after a target operation is performed on the data to be processed according to the weighted feature map.
[0094] The present application embodiment provides a data processing device, such as Figure 7 As shown, the data processing device may include: a feature extraction module 701, a similarity determination module 702, a similarity threshold acquisition module 703, an attention matrix acquisition module 704 and a weighting module 705, wherein: The feature extraction module 701 is used to extract features from the data to be processed and the reference data used to perform target processing on the data to be processed, respectively, to obtain a first feature map of the data to be processed and a second feature map of the parameter data, wherein the data to be processed and the reference data are the same or different; wherein the sizes of the first feature map and the second feature map are , h represents the number of eigenvectors included in the feature map, and m represents the number of eigenvalues included in each eigenvector in the feature map; A similarity determination module 702 is used to determine a first similarity matrix between the first feature map and the second feature map, wherein the element value of the element in the i-th row and the j-th column in the first similarity matrix represents the similarity between the i-th eigenvector in the first feature map and the j-th eigenvector in the second feature map; A similarity threshold obtaining module 703 is used to obtain a similarity threshold, subtract the elements in the first similarity matrix from the similarity threshold to obtain a second similarity matrix, and set the elements in the second similarity matrix that are less than 0 to 0 to obtain a third similarity matrix; An attention matrix obtaining module 704 is used to obtain an attention matrix between the to-be-processed data and the reference data by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix; The weighting module 705 is used to weight the second feature map based on the attention matrix to obtain a weighted feature map, and obtain the target data after the target operation is performed on the data to be processed according to the weighted feature map.
[0095] The device of the embodiments of the present application can execute the method provided by the embodiments of the present application, and the implementation principles are similar. The actions performed by each module in the device of each embodiment of the present application correspond to the steps in the method of each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, which will not be repeated here.
[0096] As an optional embodiment of the present application, when the module for obtaining a similarity threshold value is used to obtain a second similarity matrix by subtracting the elements in the first similarity matrix from the similarity threshold value, the module is specifically used to: Obtaining a similarity threshold corresponding to each row in the first similarity matrix; For each row of the first similarity matrix, subtract the element value of each element in the row from the similarity threshold corresponding to the row to obtain a second similarity matrix; Wherein, for each row of the first similarity matrix, the similarity threshold corresponding to the row is determined in the following manner: Determine the maximum element value in the row; A preset bias coefficient is obtained, and the maximum element value in the row is adjusted by using the bias coefficient to obtain a similarity threshold corresponding to the row, wherein the bias coefficient is a positive number greater than or equal to 0 and less than 1.
[0097] As an optional embodiment of the present application, when the attention matrix obtaining module is used to obtain the attention matrix between the to-be-processed data and the reference data by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix, it is specifically used to: Obtaining a fourth similarity matrix by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix; A normalization coefficient is obtained, and a normalization operation is performed on the fourth similarity matrix according to the normalization coefficient to obtain the attention matrix.
[0098] As an optional embodiment of the present application, when the attention matrix obtaining module is used to perform a normalization operation on the fourth similarity matrix, it is specifically used to: For each row of the fourth similarity matrix, determine a normalization coefficient corresponding to the row according to an element value of at least one element in the row; For each element of each row of the fourth similarity matrix, the element is adjusted to a range of 0-1 according to a normalization coefficient corresponding to the row.
[0099] As an optional embodiment of the present application, the second feature map includes a third feature map obtained by extracting features from the reference data through the first module, and a fourth feature map obtained by extracting features from the reference data through the second module; When the similarity determination module is used to determine the first similarity matrix between the first feature map and the second feature map, it is specifically used to: Performing an inner product operation on the transpose of the first feature map and the third feature map to obtain the first similarity matrix; The weighting module, when used to weight the second feature map based on the attention matrix to obtain the weighted feature map, is specifically used to: The fourth feature map is weighted by the attention matrix to obtain the weighted feature map.
[0100] As an optional embodiment of the present application, the data to be processed is an image to be processed, the reference data is a content description text corresponding to the image to be processed, and the target processing is a denoising process; when the weighting module is used to obtain the target data after the target operation is performed on the data to be processed according to the weighted feature map, it is specifically used to: Predicting a first noise in the image to be processed according to the weighted feature map; According to the first noise, the image to be processed is denoised to obtain a denoised target image.
[0101] In an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory. The processor executes the above-mentioned computer program to implement the steps of the data processing method. Compared with the related art, it can achieve: by extracting features of the data to be processed and the reference data, a first feature map of the data to be processed and a second feature map of the reference data are obtained, wherein the data to be processed and the reference data can be the same or different, so that the method provided in the embodiment of the present application can be applied to the self-attention mechanism and the cross-attention mechanism. By determining a first similarity matrix between a first feature map and a second feature map, the correlation degree between the data to be processed and the reference data in the feature space is obtained; by using a similarity threshold, the elements in the first similarity matrix are subtracted from the similarity threshold to obtain a second similarity matrix, and the elements in the second similarity matrix that are less than 0 are set to 0 to obtain a third similarity matrix, so that the obtained third similarity matrix retains feature pairs with high similarity; by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix, the third similarity matrix is sparsely processed to obtain an attention matrix between the data to be processed and the reference data; based on the attention matrix, the second feature map is weighted to obtain a weighted feature map, so that the model can perform prediction and inference based on the weighted feature map to obtain target data after the target operation is performed on the data to be processed. In an embodiment of the present application, a similarity threshold is obtained to convert the first similarity matrix into a third similarity matrix with more 0 elements. By performing a self-multiplication operation on each element in the third similarity matrix, elements with smaller values are made closer to 0, thereby achieving sparseness of the first similarity matrix. Compared with the exponential operation in the Softmax operation, the self-multiplication operation of the elements has lower computational complexity and better numerical stability, thereby effectively reducing the consumption of computing resources.
[0102] In an alternative embodiment, an electronic device is provided, such as Figure 8 As shown, Figure 8 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may also include a transceiver 4004, which may be used for data interaction between the electronic device and other electronic devices, such as data transmission and / or data reception. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present application.
[0103] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0104] The bus 4002 may include a path to transmit information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0105] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, without limitation herein.
[0106] The memory 4003 is used to store the computer program for executing the embodiment of the present application, and the execution is controlled by the processor 4001. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the above method embodiment.
[0107] Among them, electronic devices may include but are not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0108] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps and corresponding contents of the aforementioned method embodiment can be implemented. Compared with the prior art, it can be achieved that: by extracting features from the data to be processed and the reference data, a first feature map of the data to be processed and a second feature map of the reference data are obtained, wherein the data to be processed and the reference data can be the same or different, so that the method provided in the embodiment of the present application can be applied to the self-attention mechanism and the cross-attention mechanism. By determining a first similarity matrix between a first feature map and a second feature map, the correlation degree between the data to be processed and the reference data in the feature space is obtained; by using a similarity threshold, the elements in the first similarity matrix are subtracted from the similarity threshold to obtain a second similarity matrix, and the elements in the second similarity matrix that are less than 0 are set to 0 to obtain a third similarity matrix, so that the obtained third similarity matrix retains feature pairs with high similarity; by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix, the third similarity matrix is sparsely processed to obtain an attention matrix between the data to be processed and the reference data; based on the attention matrix, the second feature map is weighted to obtain a weighted feature map, so that the model can perform prediction and inference based on the weighted feature map to obtain target data after the target operation is performed on the data to be processed. In an embodiment of the present application, a similarity threshold is obtained to convert the first similarity matrix into a third similarity matrix with more 0 elements. By performing a self-multiplication operation on each element in the third similarity matrix, elements with smaller values are made closer to 0, thereby achieving sparseness of the first similarity matrix. Compared with the exponential operation in the Softmax operation, the self-multiplication operation of the elements has lower computational complexity and better numerical stability, thereby effectively reducing the consumption of computing resources.
[0109] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer readable signal media may also be any computer readable medium other than computer readable storage media, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0110] The embodiment of the present application also provides a computer program product, including a computer program, which can implement the steps and corresponding contents of the aforementioned method embodiment when executed by a processor. Compared with the prior art, it can be achieved that: by extracting features from the data to be processed and the reference data, a first feature map of the data to be processed and a second feature map of the reference data are obtained, wherein the data to be processed and the reference data can be the same or different, so that the method provided in the embodiment of the present application can be applied to the self-attention mechanism and the cross-attention mechanism. By determining a first similarity matrix between a first feature map and a second feature map, the correlation degree between the data to be processed and the reference data in the feature space is obtained; by using a similarity threshold, the elements in the first similarity matrix are subtracted from the similarity threshold to obtain a second similarity matrix, and the elements in the second similarity matrix that are less than 0 are set to 0 to obtain a third similarity matrix, so that the obtained third similarity matrix retains feature pairs with high similarity; by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix, the third similarity matrix is sparsely processed to obtain an attention matrix between the data to be processed and the reference data; based on the attention matrix, the second feature map is weighted to obtain a weighted feature map, so that the model can perform prediction and inference based on the weighted feature map to obtain target data after the target operation is performed on the data to be processed. In an embodiment of the present application, a similarity threshold is obtained to convert the first similarity matrix into a third similarity matrix with more 0 elements. By performing a self-multiplication operation on each element in the third similarity matrix, elements with smaller values are made closer to 0, thereby achieving sparseness of the first similarity matrix. Compared with the exponential operation in the Softmax operation, the self-multiplication operation of the elements has lower computational complexity and better numerical stability, thereby effectively reducing the consumption of computing resources.
[0111] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that shown or described in the drawings.
[0112] It should be understood that, although each operation step is indicated by arrows in the flowchart of the embodiment of the present application, the implementation order of these steps is not limited to the order indicated by the arrows. Unless clearly stated herein, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be performed in other orders according to demand. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on actual implementation scenarios. Some or all of these sub-steps or stages may be executed at the same time, and each sub-step or stage in these sub-steps or stages may also be executed at different times respectively. In different scenarios of execution time, the execution order of these sub-steps or stages may be flexibly configured according to demand, and the embodiment of the present application does not limit this.
[0113] The above is only an optional implementation method for some implementation scenarios of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the technical concept of the solution of the present application, other similar implementation methods based on the technical ideas of the present application are also within the protection scope of the embodiments of the present application.
Claims
1. A data processing method, characterized in that: Executed by a processor, the method includes: By performing feature extraction on the data to be processed and the reference data used to perform target processing on the data to be processed, a first feature map of the data to be processed and a second feature map of the parameter data are obtained, wherein the data to be processed and the reference data are the same or different; wherein the sizes of the first feature map and the second feature map are , Indicates the number of feature vectors included in the feature map, Indicates the number of eigenvalues included in each eigenvector in the feature map; Determine a first similarity matrix between the first feature map and the second feature map, wherein an element value of an element in an i-th row and a j-th column in the first similarity matrix represents a similarity between an i-th eigenvector in the first feature map and a j-th eigenvector in the second feature map; Obtaining a similarity threshold, subtracting the elements in the first similarity matrix from the similarity threshold to obtain a second similarity matrix, and setting the elements in the second similarity matrix that are less than 0 to 0 to obtain a third similarity matrix; Obtaining an attention matrix between the to-be-processed data and the reference data by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix; The second feature map is weighted based on the attention matrix to obtain a weighted feature map, and target data after the target operation is performed on the data to be processed is obtained according to the weighted feature map.
2. The method according to claim 1, characterized in that Subtracting the elements in the first similarity matrix from the similarity threshold to obtain a second similarity matrix includes: Obtaining a similarity threshold corresponding to each row in the first similarity matrix; For each row of the first similarity matrix, subtract the element value of each element in the row from the similarity threshold corresponding to the row to obtain a second similarity matrix; Wherein, for each row of the first similarity matrix, the similarity threshold corresponding to the row is determined in the following manner: Determine the maximum element value in the row; A preset bias coefficient is obtained, and the maximum element value in the row is adjusted by using the bias coefficient to obtain a similarity threshold corresponding to the row, wherein the bias coefficient is a positive number greater than or equal to 0 and less than 1.
3. The method according to claim 1, characterized in that The step of performing at least one self-multiplication operation on the element value of each element in the third similarity matrix to obtain an attention matrix between the data to be processed and the reference data comprises: Obtaining a fourth similarity matrix by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix; Obtain a normalization coefficient, and perform a normalization operation on the fourth similarity matrix according to the normalization coefficient to obtain the attention matrix.
4. The method according to claim 3, characterized in that The normalizing operation on the fourth similarity matrix includes: For each row of the fourth similarity matrix, determine a normalization coefficient corresponding to the row according to an element value of at least one element in the row; For each element of each row of the fourth similarity matrix, the element is adjusted to a range of 0-1 according to a normalization coefficient corresponding to the row.
5. The method according to claim 1, characterized in that The second feature map includes a third feature map obtained by extracting features from the reference data through the first module, and a fourth feature map obtained by extracting features from the reference data through the second module; The determining a first similarity matrix between the first feature map and the second feature map comprises: Performing an inner product operation on the transpose of the first feature map and the third feature map to obtain the first similarity matrix; The step of weighting the second feature map based on the attention matrix to obtain a weighted feature map includes: The fourth feature map is weighted by the attention matrix to obtain the weighted feature map.
6. The method according to any one of claims 1 to 5, characterized in that: The data to be processed is an image to be processed, the reference data is a content description text corresponding to the image to be processed, the target processing is a denoising process, and obtaining target data after the target operation is performed on the data to be processed according to the weighted feature map includes: Predicting a first noise in the image to be processed according to the weighted feature map; According to the first noise, the image to be processed is denoised to obtain a denoised target image.
7. A data processing device, characterized in that: The device comprises: The feature extraction module is used to extract features from the data to be processed and the reference data used to perform target processing on the data to be processed, respectively, to obtain a first feature map of the data to be processed and a second feature map of the parameter data, wherein the data to be processed and the reference data are the same or different; wherein the sizes of the first feature map and the second feature map are , h represents the number of eigenvectors included in the feature map, and m represents the number of eigenvalues included in each eigenvector in the feature map; A similarity determination module, configured to determine a first similarity matrix between the first feature map and the second feature map, wherein the element value of the element in the i-th row and the j-th column in the first similarity matrix represents the similarity between the i-th eigenvector in the first feature map and the j-th eigenvector in the second feature map; A similarity threshold obtaining module is used to obtain a similarity threshold, subtract the elements in the first similarity matrix from the similarity threshold to obtain a second similarity matrix, and set the elements in the second similarity matrix that are less than 0 to 0 to obtain a third similarity matrix; An attention matrix obtaining module is used to obtain an attention matrix between the to-be-processed data and the reference data by performing at least one self-multiplication operation on the element value of each element in the third similarity matrix; A weighting module is used to weight the second feature map based on the attention matrix to obtain a weighted feature map, and obtain target data after the target operation is performed on the data to be processed according to the weighted feature map.
8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Text intention recognition method and device, equipment and storage medium
CN111221944A
Voice recognition method and device, computer readable storage medium and computer equipment
CN113823264A
Data processing method, processor and computer equipment
CN118965021A