Feature extraction method, electronic equipment and computer readable storage medium

By splicing key-value vectors in the buffer, the problem of memory occupancy and memory access delay of the self-attention layer on the terminal device is solved, and efficient memory management and computing performance is achieved.

CN120372238APending Publication Date: 2025-07-25HONOR DEVICE CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202410459035.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2024-04-16
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art has high computational cost and occupies a lot of memory when processing long sequences of the self-attention layer. Especially when deploying the Transformer architecture on terminal devices with limited memory resources, there are memory pressure and memory access delay problems.

Method used

By writing the key vector and value vector of the elements to be processed in advance in the buffer, the key cache vector and value cache vector in the buffer are spliced to avoid the use of temporary storage space and reduce memory redundancy and memory access delay.

Benefits of technology

It reduces the redundancy of memory data, reduces the peak memory usage and memory access delay of the device, and improves computing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372238A_ABST
    Figure CN120372238A_ABST
Patent Text Reader

Abstract

The invention provides a feature extraction method, electronic equipment and a computer readable storage medium, and relates to the field of artificial intelligence. According to the method, a challenge vector, a first key vector and a first value vector of a to-be-processed element can be obtained, then the first key vector and the first value vector are written into a buffer area, and a self-attention representation of the to-be-processed element is obtained based on the challenge vector, a second key vector and a second value vector. Vector splicing is achieved in the mode that the first key vector and the first value vector of the to-be-processed element are written into the buffer area in advance on the terminal device with limited memory resources, the device does not need to additionally use temporary storage space to store the splicing result, redundancy of memory data can be reduced, peak memory occupation of the device is reduced, and the splicing efficiency is improved. And additional time delay overhead caused by writing the splicing result into a temporary memory can be avoided, and the memory access time delay of the equipment is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202410063470.6 and the invention title "An Operator Fusion Method", which was filed with the National Intellectual Property Administration on January 16, 2024. The entire content of this application is incorporated herein by reference. Technical Field

[0002] This application relates to the field of artificial intelligence, and in particular, to a feature extraction method, an electronic device, and a computer-readable storage medium. Background Art

[0003] Large language models, such as GPT-4 and Llama, rely on the hardware accelerator (Transformer) architecture. The Transformer can be used to process tasks such as machine translation, text classification, and text generation. The Transformer architecture includes a self-attention layer. When processing each element in a sequence, the self-attention layer calculates the query (Q) vector, key (K) vector, and value (V) vector of all elements in the sequence. Based on the K vector and the Q vector, it calculates the similarity between the element to be processed and other elements, normalizes these similarities into attention weights, and then performs a weighted sum of the V vector of each element and the corresponding attention weight to obtain the output of the self-attention layer. Among them, the longer the length of the sequence, the higher the computational cost of the self-attention layer.

[0004] In related technologies, in order to reduce the computational cost of the self-attention layer when processing long sequences, the key-value cache (KV cache) technology is introduced. This technology proposes that the K vector and the V vector of the processed elements can be cached for repeated use when processing different elements, achieving the effect of reducing redundant calculations and lowering the computational cost. However, this method will additionally occupy memory space, bringing memory pressure to the device. Moreover, as the task progresses and the scale of the model used increases, the size of the KV cache will increase significantly, which is very disadvantageous for deploying the Transformer architecture on edge devices with extremely limited memory resources. Summary of the Invention

[0005] Embodiments of this application provide a feature extraction method, an electronic device, and a computer-readable storage medium, which can reduce the memory space occupied and the memory access latency during the feature extraction process.

[0006] To achieve the above object, the embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides a feature extraction method, which can be applied to an electronic device, such as a mobile phone, a tablet, or other terminal devices. The method can obtain a query vector, a first key vector, and a first value vector of an element to be processed, then write the first key vector and the first value vector into a buffer, and obtain a self-attention representation of the element to be processed based on the query vector, a second key vector, and a second value vector. Since the buffer originally stores a key vector of an element associated with the element to be processed (i.e., a key cache vector) and a value vector of an element associated with the element to be processed (i.e., a value cache vector), after writing the first key vector and the first value vector into the buffer and then accessing the buffer, the second key vector and the second value vector can be obtained. The second key vector includes the first key vector and the key cache vector, and the second value vector includes the second value vector and the value cache vector, which is equivalent to realizing the concatenation of the first key vector and the key cache vector, and the concatenation of the first value vector and the value cache vector.

[0008] It can be understood that the present application takes into account the limited memory resources of terminal devices such as mobile phones and tablets. Therefore, the first key vector and the first value vector of the element to be processed are written into the buffer in advance, and the concatenation result can be obtained without using a concatenation operator, and the device does not need to use additional temporary storage space to store the concatenation result. This can not only reduce the redundancy of memory data and the peak memory occupancy of the device, but also avoid the additional latency overhead caused by writing the concatenation result into temporary memory, and reduce the latency (i.e., memory access latency) caused by the device accessing memory during the calculation process.

[0009] In an implementation manner provided in the first aspect, the electronic device includes a self-attention module. The step of obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector includes: obtaining the self-attention representation of the element to be processed through the self-attention module based on the query vector, the second key vector, and the second value vector. That is to say, the present application can obtain the self-attention representation of the element to be processed through the self-attention module.

[0010] In an implementation manner provided in the first aspect, the self-attention module includes a first matrix multiplication operator, an addition operator, a normalization operator, and a second matrix multiplication operator. The step of obtaining the self-attention representation of the element to be processed through the self-attention module based on the query vector, the second key vector, and the second value vector specifically includes: performing matrix multiplication on the query vector and the transposed vector of the second key vector through the first matrix multiplication operator to obtain a matrix multiplication result; superimposing a preset mask and the scaled result of the matrix multiplication result through the addition operator to obtain a superimposed result; performing normalization processing on the superimposed result through the normalization operator to obtain a normalized processing result; performing matrix multiplication on the normalized processing result and the second value vector through the second matrix multiplication operator to obtain the self-attention representation of the element to be processed.

[0011] In an implementation provided in the first aspect, the buffer includes a key buffer, and the key buffer includes a key cache vector. Before the step of obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector, the electronic device can also access the key buffer through a first matrix multiplication operator in a first order to obtain the transposed vector of the second key vector. The first order is opposite to the data storage order of the key cache vector. For example, if the electronic device stores the data of the key cache vector in a row-first and column-second order, then the first order is column-first and row-second; another example is that if the electronic device stores the data of the key cache vector in a column-first and row-second order, then the first order is row-first and column-second. In short, by reading the data in an order opposite to the data storage order of the key cache vector, the electronic device can read the data in the order opposite to the data storage order of the key cache vector. Through this way of reading data, the self-attention module can transpose the second key vector without using a transpose operator, which can not only simplify the structure of the self-attention layer but also reduce the number of times the self-attention module accesses the memory, thereby reducing the latency caused by accessing the memory.

[0012] In an implementation provided in the first aspect, the buffer includes a key buffer, and the key buffer includes a key cache vector. The self-attention module further includes a transpose operator. Before the step of obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector, the electronic device can also access the key buffer through the transpose operator in a second order to obtain the second key vector. The second order is the same as the data storage order of the key cache vector; the transpose operator is used to swap the rows and columns of the second key vector to obtain the transposed vector of the second key vector. That is to say, when the self-attention module includes a transpose operator, the electronic device reads the data in the same order as the data storage order of the key cache vector so as to transpose the second key vector through the transpose operator.

[0013] In an implementation provided in the first aspect, the buffer further includes a value buffer, and the value buffer includes a value cache vector. Before the step of obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector, the electronic device can also access the value buffer through a second matrix multiplication operator in a second order to obtain the second value vector. The second order is the same as the data storage order of the key cache vector.

[0014] In an implementation provided in the first aspect, the self-attention module further includes a multiplication operator. The step of obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector through the self-attention module further includes: scaling the matrix multiplication result through the multiplication operator.

[0015] In an implementation provided in the first aspect, in order to further reduce the memory space occupied by the KV cache, the electronic device can also perform quantization operations on the data when writing the data. For this purpose, the buffer is also used to store the quantization parameter matrix corresponding to the key cache vector and the quantization parameter matrix corresponding to the value cache vector. Thus, the step of writing the first key vector and the first value vector into the buffer includes: performing a quantization operation on the first key vector to obtain a third key vector and a first quantization parameter matrix, and performing a quantization operation on the first value vector to obtain a third value vector and a second quantization parameter matrix; wherein, the first quantization parameter matrix is used to indicate the mapping relationship between the first key vector and the third key vector, the precisions of the first key vector and the third key vector are different, the second quantization parameter matrix is used to indicate the mapping relationship between the first value vector and the third value vector, and the precisions of the first value vector and the third value vector are different. After performing the quantization operation, the electronic device can write the third key vector, the first quantization parameter matrix, the third value vector, and the second quantization parameter matrix into the buffer. Correspondingly, the electronic device can also perform an inverse quantization operation on the data when accessing the data. Specifically: the electronic device can obtain a second key vector and a second value vector based on a fourth key vector, a third quantization parameter matrix, a fourth value vector, and a fourth quantization parameter matrix; wherein, the fourth key vector includes the third key vector and the key cache vector, the third quantization parameter matrix includes the first quantization parameter matrix and the quantization parameter matrix corresponding to the key cache vector, the fourth value vector includes the third value vector and the value cache vector, and the fourth quantization parameter matrix includes the second quantization parameter matrix and the quantization parameter matrix corresponding to the value cache vector. In this way, through the quantization operation and the inverse quantization operation, both the memory space occupied by the KV cache can be reduced, and the calculation accuracy of the device when calculating the self-attention representation of the elements to be processed can be ensured.

[0016] In an implementation provided in the first aspect, the step of obtaining the second key vector and the second value vector based on the fourth key vector, the third quantization parameter matrix, the fourth value vector, and the fourth quantization parameter matrix specifically includes: performing an inverse quantization operation on the fourth key vector through the third quantization parameter matrix to obtain the second key vector; performing an inverse quantization operation on the fourth value vector through the fourth quantization parameter matrix to obtain the second value vector.

[0017] In a second aspect, an embodiment of the present application provides an electronic device, which includes: a memory and one or more processors; the memory is coupled to the processor; the memory is used to store computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the electronic device, the electronic device is caused to execute the method according to the first aspect and any one of its implementations.

[0018] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing computer instructions, which, when running on an electronic device, cause the electronic device to execute the method according to the first aspect and any one of its implementation manners.

[0019] In a fourth aspect, an embodiment of the present application provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the method provided by the first aspect and any one of its implementation manners.

[0020] Among them, for the technical effects brought by any one of the design manners in the second to fourth aspects, reference may be made to the technical effects brought by different implementation manners in the first aspect, which will not be elaborated herein. Description of the Drawings

[0021] Figure 1 It is a schematic diagram of a network structure of a self-attention layer;

[0022] Figure 2 It is a schematic diagram of a network structure of a self-attention layer based on KV cache provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic diagram of the structure of an electronic device 100 provided by an embodiment of the present application;

[0024] Figure 4 It is a flowchart of a feature extraction method provided by an embodiment of the present application Figure 1 ;

[0025] Figure 5 It is a storage schematic diagram of a key buffer before and after an electronic device writes a K1 vector provided by an embodiment of the present application;

[0026] Figure 6 It is a schematic diagram of the structure of a self-attention module provided by an embodiment of the present application;

[0027] Figure 7 It is a schematic diagram of the structure of another self-attention module provided by an embodiment of the present application;

[0028] Figure 8 It is a flowchart of a feature extraction method provided by an embodiment of the present application Figure 2 ;

[0029] Figure 9 It is a schematic diagram of a quantization process provided by an embodiment of the present application;

[0030] Figure 10 It is a storage schematic diagram of another key buffer before and after an electronic device writes a K1 vector provided by an embodiment of the present application;

[0031] Figure 11 Flow schematic of a feature extraction method provided by an embodiment of this application Figure 3 。 Specific implementation manners

[0032] The following describes the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Among them, in the description of the embodiments of this application, the terms used in the following embodiments are only for the purpose of describing specific embodiments, and are not intended to limit this application. As used in the specification and claims of this application, the singular forms "a", "the", "above", "this" and "this one" are also intended to include expressions such as "one or more", unless there is a clear indication to the contrary in the context. It should also be understood that in the following embodiments of this application, "at least one" and "one or more" mean one or more than two (including two). The term "and / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist; for example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally means that the associated objects before and after are an "or" relationship.

[0033] The reference to "an embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure or characteristic described in combination with this embodiment is included in one or more embodiments of this application. Thus, the statements "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "include", "comprise", "have" and their variants all mean "include but not limited to", unless otherwise specifically emphasized in other ways. The term "connection" includes direct connection and indirect connection, unless otherwise stated. "First" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features.

[0034] In the embodiments of this application, words such as "exemplarily" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0035] Since the embodiments of this application involve a large number of applications of self-attention layers, for the sake of easy understanding, the related technologies involved in the embodiments of this application will be introduced below.

[0036] When processing sequential data, each element (token) in the sequence can be associated with other elements in the sequence, rather than relying only on adjacent elements. It can adaptively capture long-range dependencies between elements by calculating the relative importance between elements.

[0037] Specifically, for each element in the sequence, the self-attention layer can calculate its similarity with other elements and normalize these similarities into attention weights. Then, by performing a weighted sum of each element with the corresponding attention weight, the output of the self-attention layer can be obtained.

[0038] Refer to Figure 1 , Figure 1 is a schematic diagram of a network structure of the self-attention layer. As Figure 1 shown, the self-attention layer includes a query (Q) feature transformation module (i.e., the Q module in Figure 1 ), a key (K) feature transformation module (i.e., the K module in Figure 1 ), and a value (V) feature transformation module (i.e., the V module in Figure 1 ). The input x of the self-attention layer is respectively subjected to feature transformation by the Q module, the K module, and the V module to become Q (vector), K (vector), and V (vector); weight coefficients are generated based on K and Q, and then V is subjected to a weighted sum according to the weight coefficients, and finally the weighted feature out is output.

[0039] Please continue to refer to Figure 1 ,and the processing process of the self-attention layer will be introduced below. The input of the self-attention layer is x. First, three feature transformations are respectively performed on x to obtain Q, K, and V:

[0040] Q = xW Q ,K = xW k ,V = xW V Equation (1)

[0041] In Equation (1), x is the input of the self-attention layer, W Q ,W k ,W V are feature transformation matrices, W Q ,W k ,W V are fully connected layers in the convolutional neural network, and their weights are learnable parameters. Multiply x with W Q ,W k ,W VMultiply them to obtain Q, K, and V. The purpose of Q, K, and V is to map the x features to another dimension so that the mapped features can meet the requirements for Q, K, and V in the Attention operation. W Q , W k , W V The functions of W

[0042] will gradually emerge as the deep neural network training process progresses. Among them, after the transformation by W k , the obtained K can describe the content features of the input, representing what the input features are like; after the transformation by W Q , the obtained Q can contain the guiding features regarding the input, representing what information is needed for the model to process; and after the transformation by W V , the obtained V is a vector expressing the input features.

[0043] After obtaining Q, K, and V, the self-attention layer can perform matrix multiplication (MatMul) on Q and K, and then use a scaling layer for numerical scaling processing; optionally, in the field of text processing, the self-attention layer can also use a masking layer to process the numerically scaled data; finally, after a normalization (softmax) operation, an Attention Map describing the self-correlation between certain dimensions of x is obtained, that is, the Attention Map is:

[0044]

[0045] In formula (2), d is a constant, which is a custom value and is specifically set according to the different needs of the model. Performing matrix multiplication on Q and K T is equivalent to searching for guiding features in the content features.

[0046] Finally, weight V with the Attention Map, that is, perform matrix multiplication on V and the Attention Map, and then use it as the output of the self-attention layer after feature transformation. Thus, the output out of the self-attention layer is:

[0047] out = AVW out Formula (3)

[0048] The W in formula (3) out is a feature transformation matrix.

[0049] To establish associations between each element and other elements in the sequence, when the self-attention layer processes any element in the sequence, it will calculate the K and Q of all elements in the sequence. This will increase the computational cost of the device, and the longer the sequence, the higher the computational cost.

[0050] In related technologies, to reduce the computational cost when the self-attention layer processes long sequences, KV cache is introduced. KV cache can reduce redundant calculations and lower the computational cost by caching the keys and values of the processed data for subsequent reuse.

[0051] Reference Figure 2 , Figure 2 is a schematic diagram of the network structure of a self-attention layer based on KV cache. Briefly comparing Figure 1 and Figure 2 shows that the processing flows of the two are similar, and the differences are as follows:

[0052] (1) In the structure shown in Figure 2 , after the self-attention layer obtains the K vector, it will concatenate (concat) the K vector and the key cache vector to get Kout, and then perform a transpose operation on Kout and then do a matrix multiplication (MatMul) with the Q vector. Among them, Figure 2 's Kout is equivalent to Figure 1 's K. In addition, after concatenating to get Kout, the self-attention layer will store Kout in temporary memory until it completes the matrix multiplication with Q and updates the key cache vector before releasing it.

[0053] (2) In the structure shown in Figure 2 , after the self-attention layer obtains the V vector, it will concatenate (concat) the V vector and the value cache vector to get Vout, and then do a matrix multiplication of Vout and the result of softmax to get the output out. Among them, Figure 2 's Vout is equivalent to Figure 1 's V. In addition, after concatenating to get Vout, the self-attention layer will store Vout in temporary memory until it completes the matrix multiplication with the result of softmax and updates the value cache vector before releasing it.

[0054] It can be seen that in the network structure shown in Figure 2 , before releasing Kout and Vout, the self-attention layer not only needs to occupy memory to store the key cache vector and the value cache vector, but also needs to occupy temporary memory to store Kout and Vout. Taking the length of the key cache vector and the value cache vector as 2048, there are 32 self-attention layers in the electronic device, and each cache vector includes 32 * 128 data as an example, then Kout and Vout need to occupy 2048 * 32 * 2 * 32 * 128 * 2 / 1024 / 1024 / 1024 = 1GB of memory space.

[0055] But in fact, there is a large amount of duplicate data between Kout and the key cache vector (i.e., the key cache vector), and there is a large amount of duplicate data between Vout and the value cache vector (i.e., the value cache vector). This will lead to redundancy of memory data, increase the peak memory usage of the self-attention layer during operation and the memory pressure of the device, and may further cause the process to be killed by the system or affect the use of other processes on the device.

[0056] In particular, as the task progresses and the scale of the model used increases, the size of the KV cache will increase significantly, which is not conducive to deploying the Transformer architecture on end-side devices with extremely tight memory resources.

[0057] In order to at least solve the above problems, an embodiment of the present application provides a feature extraction method. Taking into account the limited memory resources of terminal devices such as mobile phones and tablets, the K vector and V vector of the input data are written into the key buffer and the value buffer in advance respectively to realize the splicing of the K vector and the key cache vector, and the splicing of the V vector and the V cache vector. There is no need for the device to use additional temporary storage space to store the splicing results. This can not only reduce the redundancy of memory data and the peak memory occupancy of the self-attention layer, but also avoid the additional delay overhead caused by writing the splicing results into the temporary memory, and reduce the delay caused by the self-attention layer accessing the memory during the calculation process (i.e., memory access delay).

[0058] The feature extraction method provided in the embodiment of the present application can be applied to electronic devices. The electronic device can be a mobile phone (including a straight-screen mobile phone, a folding screen mobile phone, etc.), a tablet computer, a personal communication service (PCS) phone, a virtual reality (VR) electronic device, an augmented reality (AR) device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical surgery, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc., without specific limitation.

[0059] Figure 3 FIG. 1 is a schematic diagram of the structure of an electronic device 100 provided in an embodiment of the present application. Figure 3As shown in the figure, the electronic device 100 may include: a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a key 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc.

[0060] Among them, the processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors. The processor 210 may be the nerve center and command center of the electronic device 100. The processor 210 may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0061] A memory may also be provided in the processor 210 for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory may save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can be directly called from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0062] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0063] The external memory interface 220 may be used to connect to an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device 100. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.

[0064] The internal memory 221 may be used to store computer-executable program code, and the executable program code includes instructions. The processor 210 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 221. For example, in the embodiments of the present application, the processor 210 may execute the instructions stored in the internal memory 221. The internal memory 221 may include a storage program area and a storage data area.

[0065] Among them, the storage program area may store an operating system, application programs required for at least one function (such as a shopping function, etc.). The storage data area may store data created during the use of the electronic device 100 (such as a downgrading policy sent from the cloud side, user characteristics, a feature set of a certain type of user, a function list, etc.). In addition, the internal memory 221 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0066] It can be understood that the interface connection relationships between the modules illustrated in this embodiment are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments, the electronic device 100 may also include more or fewer modules than those provided in the above embodiments, and different interface connection manners or combinations of multiple interface connection manners may also be adopted between the modules.

[0067] The feature extraction method provided by the embodiments of the present application will be specifically described below with reference to the accompanying drawings.

[0068] Refer to Figure 4 , which is a flowchart illustration of a feature extraction method provided by an embodiment of the present application Figure 1 . As Figure 4 shown, the feature extraction method provided by the embodiments of the present application at least includes S401 to S404.

[0069] S401, the electronic device obtains the Q1 vector, K1 vector, and V1 vector of the element to be processed.

[0070] Among them, the element to be processed may be words, phrases, words, phrases, etc. that can be understood by users, and no specific limitation is made here. For example, the element to be processed may be "yes", "you", "is", etc.

[0071] Among them, the Q1 vector is the query vector of the element to be processed, the K1 vector is the key vector of the element to be processed, and the V1 vector is the value vector of the element to be processed. Among them, the query vector can be understood as the query target, which can be used to match other elements and calculate the association or relationship between the element to be processed and other elements; the key vector can be understood as the key information of the element and can be used to match the query vector; the value vector can be used to represent the element itself. In a possible design, the Q1 vector may also be referred to as the first query (Q) vector, the K1 vector may also be referred to as the first key (K) vector, and the V1 vector may also be referred to as the first value (V) vector.

[0072] In the embodiments of the present application, the electronic device may perform an embedding operation on the element to be processed to obtain the embedding vector of the element to be processed, and then perform feature transformation on the embedding vector of the element to be processed to obtain the Q1 vector, K1 vector, and V1 vector. Among them, the embedding vector may also be referred to as the word vector.

[0073] Among them, the embedding operation can map elements such as words, phrases, etc. into a new multi-dimensional space to obtain the multi-dimensional vector representation of the element. Among them, the more similar the elements are, the more similar their corresponding embedding vectors are.

[0074] In a possible design, an electronic device can perform an embedding operation on an element to be processed using models such as Word2Vec and Bidirectional Encoder Representations from Transformers (BERT) to obtain an embedding vector of the element to be processed.

[0075] In an embodiment of the present application, the electronic device can utilize three feature transformation matrices W Q , W K , W V to perform feature transformations on the embedding vector of the element to be processed respectively, to obtain a Q1 vector, a K1 vector, and a V1 vector. Thus, the Q1 vector, the K1 vector, and the V1 vector satisfy:

[0076] Q1 = xW Q , K1 = xW K , V1 = xW V Equation (4)

[0077] where x is the embedding vector of the element to be processed, W Q , W K , W V are feature transformation matrices respectively, Q1 is the Q1 vector, K1 is the K1 vector, and V1 is the V1 vector.

[0078] As Figure 4 shown, the electronic device can include a feature transformation module, and the electronic device can utilize the feature transformation module to execute S401, that is, the feature transformation module can obtain the Q1 vector, the K1 vector, and the V1 vector of the element to be processed.

[0079] S402, the electronic device writes the K1 vector into the key buffer, and writes the V1 vector into the value buffer.

[0080] Among them, the key buffer can be used to store K cache vectors, and the K cache vectors are K vectors of historical elements; the value buffer can be used to store V cache vectors, and the V cache vectors are V vectors of historical elements. Historical elements refer to elements that the electronic device has already processed. Optionally, the K cache vector and the V cache vector can also be the K vector and the V vector of an element associated with the element to be processed respectively. For example, an element associated with the element to be processed is an element involved in a single question-and-answer process or multiple related question-and-answer processes.

[0081] In a possible design, the electronic device can write the K1 vector to the end of the K cache vector sequence, and write the V1 vector to the end of the V cache vector sequence.

[0082] Refer to Figure 5 , which is a storage schematic diagram of the key buffer before and after writing the K1 vector into the electronic device. As Figure 5 shown, before the electronic device writes the K1 vector into the key buffer, the key buffer includes P K cache vectors, such as K cache vector 1 to K cache vector P. Among them, the dimensions of the K1 vector, K cache vector 1 to K cache vector P are all M×N. After the electronic device writes the K1 vector into the key buffer, the key buffer includes P + 1 K cache vectors, specifically including the original P K cache vectors and the newly written K1 vector, where the K1 vector is located at the end of the data sequence stored in the key buffer, that is, after the P K cache vectors.

[0083] Similarly, before the electronic device writes the V1 vector into the value buffer, the value buffer includes P V cache vectors, such as V cache vector 1 to V cache vector P. Among them, the dimensions of the V1 vector, V cache vector 1 to V cache vector P are also M×N. After the electronic device writes the V1 vector into the value buffer, the value buffer includes P + 1 V cache vectors, specifically including the original P V cache vectors and the newly written V1 vector, where the V1 vector is located at the end of the data sequence stored in the value buffer, that is, after the P V cache vectors.

[0084] It can be understood that after writing the K1 vector and V1 vector into the key buffer and value buffer respectively, the key buffer stores the original K cache vectors and the newly written K1 vector, and the value buffer stores the original V cache vectors and the newly written V1 vector. It is equivalent to the electronic device achieving the splicing of the K1 vector and the K cache vectors, and the splicing of the V1 vector and the V cache vectors without performing the concat operation.

[0085] S403, the electronic device accesses the key buffer and the value buffer to obtain the K2 vector and the V2 vector.

[0086] Among them, the K2 vector includes the K cache vector and the K1 vector, and the V2 vector includes the V cache vector and the V1 vector. It should be noted that the K2 vector can also be called the second key vector, and the V2 vector can also be called the second value vector.

[0087] In a possible design, still as Figure 4 shown, the electronic device further includes a self-attention module. Among them, the self-attention module of the electronic device can access the key buffer and the value buffer to obtain the K2 vector and the V2 vector.

[0088] S404, the electronic device obtains the attention representation of the element to be processed based on the K2 vector, V2 vector and Q1 vector.

[0089] As Figure 4As shown in the figure, the electronic device can use the self-attention module to obtain the attention representation of the element to be processed based on the K2 vector, the V2 vector, and the Q1 vector. It can be understood that the attention representation of the element to be processed can be used to reflect the association between the element to be processed and the processed elements.

[0090] The following will specifically describe the process by which the self-attention module obtains the attention representation of the element to be processed based on the K2 vector, the V2 vector, and the Q1 vector in combination with the structure of the self-attention module.

[0091] Please refer to Figure 6 , which is a schematic structural diagram of a self-attention module provided by an embodiment of the present application. As Figure 6 shown, the self-attention module includes a transpose operator, a matrix multiplication (MatMul) operator 1, a multiplication (Mul) operator, an addition (add) operator, a normalization (softmax) operator, and a matrix multiplication operator 2. Among them, the transpose operator can be used to exchange the rows and columns of a vector, the matrix multiplication operator can be used to perform a dot product operation on vectors, the multiplication operator can be used to perform an operation of magnifying / shrinking vectors, the addition operator can be used to stack vectors, and the normalization operator can be used to perform a normalization process on vectors.

[0092] Figure 6 The operation process of the self-attention module shown in the figure is as follows: First, the transpose operator reads the key buffer in the order of first row and then column to obtain the K2 vector, and then the transpose operator performs a transpose operation on the K2 vector to obtain the K2 T vector and stores the K2 T vector in the temporary storage space. The K2 T vector is the transposed K2 vector. Secondly, the matrix multiplication operator 1 accesses the temporary storage space to obtain the K2 T vector, performs matrix multiplication on the Q1 vector and the K2 T vector, and stores the result of the matrix multiplication in the temporary storage space. Then, the multiplication operator accesses the temporary storage space to obtain the result of the matrix multiplication, scales the result of the matrix multiplication, and writes it into the temporary storage space. Then, the addition operator accesses the temporary storage space to obtain the scaled result, stacks the scaled result and the mask, and writes the stacked result into the temporary storage space. Next, the normalization operator accesses the temporary storage space to obtain the stacked result, performs a normalization process on the stacked result, and writes the normalized result into the temporary storage space. Finally, the matrix multiplication operator 2 accesses the value buffer to obtain the V2 vector, accesses the temporary storage space to obtain the normalized result, and performs matrix multiplication on the V2 vector and the normalized result to obtain the attention representation of the element to be processed.

[0093] According to the above process, it can be known that the K2 vector, the V2 vector, the Q1 vector, and the attention representation of the element to be processed satisfy:

[0094]

[0095] Among them, out is the attention representation of the element to be processed, Q1 is the Q1 vector, K2 is the K2 vector, V2 is the V2 vector, and K2 T is the transposed K2 vector, and d is a preset constant.

[0096] In a possible design, the matrix multiplication operator 1 can be fused with the transpose operator. Please refer to Figure 7 , which is a schematic structural diagram of another self-attention module provided by the embodiments of the present application. As Figure 7 shown, the self-attention module includes a matrix multiplication operator 1, a multiplication operator, an addition operator, a normalization operator, and a matrix multiplication operator 2.

[0097] Figure 7 The operation process of the self-attention module shown is as follows: First, the matrix multiplication operator 1 reads the key buffer in the order of columns first and then rows to obtain the K2 T vector, performs matrix multiplication on the Q1 vector and the K2 T vector, and writes the matrix multiplication result into the temporary storage space. Then, the multiplication operator accesses the temporary storage space to obtain the result of the matrix multiplication, scales the result of the matrix multiplication, and writes it into the temporary storage space. Next, the addition operator accesses the temporary storage space to obtain the scaled result, superimposes the scaled result and the mask, and writes the superimposed result into the temporary storage space. Then, the normalization operator accesses the temporary storage space to obtain the superimposed result, normalizes the superimposed result, and writes the normalized result into the temporary storage space. Finally, the matrix multiplication operator 2 accesses the value buffer to obtain the V2 vector, accesses the temporary storage space to obtain the normalized result, and performs matrix multiplication on the V2 vector and the normalized result to obtain the attention representation of the element to be processed.

[0098] It should be noted that in the embodiments of the present application, the electronic device can write data into the key buffer and the value buffer in the order of rows first and then columns. For example, if the key vector is then the electronic device can write the values of the key vector into the key buffer in the order of "1", "3", "2", "4". In this way, when the matrix multiplication operator 1 reads the key buffer in the order of columns first and then rows, the transposed vector of the key vector can be obtained

[0099] Of course, in other embodiments, the electronic device can also write data into the key buffer and the value buffer in the order of columns first and then rows; correspondingly, the matrix multiplication operator 1 accesses the key buffer in the order of rows after columns to obtain the K2 T vector.

[0100] In short, the matrix multiplication operator 1 accesses the key buffer in the order opposite to the data storage order of the key cache vector (which can be called the first order), and the transpose of the vector can be achieved to obtain K2T Vector. Additionally, the matrix multiplication operator 2 accesses the value buffer in the same order as the data storage order of the key buffer vector (which can be referred to as the second order), and thus the V2 vector can be obtained.

[0101] It can be seen that in Figure 7 the structure shown, by adjusting the order in which the matrix multiplication operator 1 reads data from the key buffer, the matrix multiplication operator 1 realizes the transpose of the vector during the data reading process, which can not only realize the fusion of the transpose operator and the matrix multiplication operator and simplify the structure of the self-attention module, but also reduce the number of I / O operations and the memory access latency during the calculation process of the self-attention module.

[0102] In a possible design, the matrix multiplication operator 1 can also be fused with the multiplication operator, that is, the self-attention module may not include the multiplication operator. In this case, before storing the matrix multiplication result in the temporary storage space, the matrix multiplication operator 1 can first scale the matrix multiplication result using a preset coefficient, and then write the scaled result into the temporary storage space for the addition operator to use. This can further simplify the structure of the self-attention module and reduce the memory access latency during the calculation process of the self-attention module.

[0103] In summary, in the embodiments of the present application, the electronic device can write the K1 vector and the V1 vector of the elements to be processed into the key buffer and the value buffer respectively in advance before the self-attention module calculates, and then the self-attention module accesses the key buffer and the value buffer to obtain the K2 vector and the V2 vector, and finally the self-attention module obtains the attention representation of the elements to be processed. By reusing the key buffer and the value buffer, the self-attention module does not need to perform a splicing operation, further avoiding the self-attention module from additionally occupying the temporary storage space to store the splicing result, which can not only reduce the redundancy of the memory data, but also reduce the memory occupancy.

[0104] Furthermore, the fact that no splicing operation is required enables the self-attention module to remove the splicing operator and simplify the structure of the self-attention module. And the fact that the self-attention module does not need to perform a splicing operation also makes the self-attention module not need to store / read the splicing result, which can achieve the effect of reducing the memory access latency during the calculation process of the self-attention module.

[0105] In the related art, the precision of the data used during the model operation is mostly FP16, that is, the data is 16-bit floating-point data. Correspondingly, the precision of the data stored in the key buffer and the value buffer is also FP16. Of course, the precision may also be FP32, etc., which is not specifically limited here. This will occupy more memory space.

[0106] In the embodiments of the present application, the electronic device can perform quantization processing on data, converting high-precision floating-point data into integer representations with lower bit widths. This can further save the resident memory consumption of the self-attention model during the inference process without significantly sacrificing precision. Among them, the integer with a lower bit width is, for example, an 8-bit integer, a 4-bit integer, etc., denoted as Int8 and Int4 respectively.

[0107] Refer to Figure 8 , which is a schematic flowchart of a feature extraction method provided by an embodiment of the present application. Figure 2 . As Figure 8 shown, S402 includes S4021 to S4022.

[0108] S4021, perform a quantization operation on the K1 vector to obtain the K3 vector and the quantization parameter matrix s1, and perform a quantization operation on the V1 vector to obtain the V3 vector and the quantization parameter matrix s2.

[0109] Among them, the K3 vector is the quantized K1 vector, and the quantization parameter matrix s1 is used to convert the K1 vector into the K3 vector; the V3 vector is the quantized V1 vector, and the quantization parameter matrix s2 is used to convert the V1 vector into the V3 vector. The precision of the K1 vector is higher than that of the K3 vector, and the precision of the V1 vector is higher than that of the V3 vector. For example, the precision of both the K1 vector and the V1 vector is FP16, and the precision of both the K3 vector and the V3 vector is Int8. In a possible design, the K3 vector can also be referred to as the third key vector, the V3 vector can also be referred to as the third value vector, the quantization parameter matrix s1 can also be referred to as the first quantization parameter matrix, and the quantization parameter matrix s2 can also be referred to as the second quantization parameter matrix.

[0110] Among them, the quantization parameter matrix includes quantization parameters, and the quantization parameters include a quantization step and a zero point.

[0111] In the embodiments of the present application, the electronic device can perform quantization operations on the K1 vector and the V1 vector in an asymmetric quantization or symmetric quantization manner. The basic processes of asymmetric quantization and symmetric quantization will be introduced separately below.

[0112] (1) Asymmetric quantization means mapping the data to be quantized within the range [Rmin, Rmax] to quantization values within the range [0, 2 b -1] through the quantization step and the zero point, where b is the number of bits of the quantization value. For example, if the precision of the quantization value is Int8, then b = 8, so asymmetric quantization can map the data to be quantized within the range [Rmin, Rmax] to the range [0, 255].

[0113] For asymmetric quantization, the quantization step and the zero point satisfy the formula:

[0114]

[0115] Among them, s is the quantization step size, z is the zero point, Rmax is the maximum value in the data to be quantized, Rmin is the minimum value in the data to be quantized, and round() is the rounding function.

[0116] After obtaining the quantization step size and zero point, for any data to be quantized, its corresponding quantization value is:

[0117]

[0118] Among them, r is the data to be quantized, and q is the quantization value.

[0119] (2) Symmetric quantization means mapping the data to be quantized in the range [Rmin, Rmax] to the quantization value in the range [-2 b-1 , 2 b-1 -1] through the quantization step size and zero point, and b is the number of bits of the quantization value. For example, if the precision of the quantization value is Int8, then b = 4, so symmetric quantization can map the data to be quantized in the range [Rmin, Rmax] to the range [-128, 127].

[0120] For symmetric quantization, the zero point z = 0, and the quantization step size satisfies the formula:

[0121]

[0122] Among them, s is the quantization step size, and Rmax is the maximum value in the data to be quantized.

[0123] After obtaining the quantization step size and zero point, for any data to be quantized r, its corresponding quantization value q also satisfies the above formula (7), the difference being that z = 0.

[0124] In the embodiments of the present application, both the K1 vector and the V1 vector include multiple data to be quantized. Taking the quantization operation of the K1 vector as an example, the electronic device can divide the K1 vector into multiple data groups, each data group includes at least one data to be quantized, and then the electronic device calculates the corresponding quantization step size and zero point (i.e., quantization parameters) for each data group respectively, and finally calculates the quantization value of each data to be quantized in the data group through the corresponding quantization step size and zero point. Among them, the quantization process of the V1 vector is the same and will not be elaborated here.

[0125] It can be understood that the quantization parameter matrix s1 includes as many quantization parameters as the number of data groups into which the electronic device divides the K1 vector.

[0126] Comparing the above two quantization methods, it can be seen that the zero point z of symmetric quantization remains unchanged and is 0, while the zero point z of asymmetric quantization can change with the change of the range of data to be quantized. Therefore, there are fewer operations involved in the calculation process of symmetric quantization, and it is not necessary to occupy memory space to record the zero point z. Therefore, in order to reduce the running cost and memory occupancy, the embodiments of the present application may adopt the method of symmetric quantization to quantize data, and the following will also take symmetric quantization as an example for illustration. Of course, in other embodiments, the electronic device may still adopt the method of asymmetric quantization, which is not specifically limited herein.

[0127] That is to say, in the embodiments of the present application, the quantization parameter may only include the quantization step size.

[0128] Exemplarily, referring to Figure 9 , taking the K1 vector / V1 vector including M×N data to be quantized, and each data group including m1×n1 data to be quantized as an example, it exemplarily illustrates the process of the electronic device quantizing the K1 vector / V1 vector; where M, N, m1, and n1 are all 2 n , and n is an integer. As Figure 9 shown, for any data group in the K1 vector / V1 vector, the electronic device can calculate the quantization step size corresponding to this data group by using formula (8); then for each data group, the electronic device can calculate the quantization value of each data to be quantized in this data group based on the corresponding quantization step size by using formula (7).

[0129] Still as Figure 9 shown, data group a includes m1×n1 data to be quantized. The electronic device can determine the maximum value among these m1×n1 data to be quantized, and combine the precision of the quantization value to determine the quantization step size a corresponding to data group a. Then the electronic device calculates the quantization value of each data to be quantized in data group a through the quantization step size a.

[0130] In this way, the electronic device can obtain a quantization parameter matrix with a size of (M / m1)×(N / n1), and a K3 vector / V3 vector with a size of M×N. In other words, the quantization parameter matrix includes (M / m1)×(N / n1) quantization step sizes, and the K3 vector / V3 vector includes M×N quantization values.

[0131] For example, the size of the K1 vector / V1 vector can be 32×128 (i.e., M = 32, N = 128), and each data group can include 1×16 (i.e., m1 = 1, n1 = 16) data to be quantized. In this way, the electronic device can obtain a quantization parameter matrix with a size of 32×8, and a K3 vector / V3 vector with a size of 32×128. It can be understood that the above M, N, m1, and n1 can also be other values, which are not specifically limited herein.

[0132] It should be noted that the number of data to be quantified included in each data group can affect the memory space occupied by the quantization parameter matrix and the quantization accuracy. Among them, the more the number of data to be quantified included in each data group, the fewer the number of quantization steps included in the quantization parameter matrix, the smaller the memory space occupied by the quantization parameter matrix, and the greater the accuracy loss during the quantization process; the fewer the number of data to be quantified included in each data group, the more the number of quantization steps included in the quantization parameter matrix, the larger the memory space occupied by the quantization parameter matrix, and the smaller the accuracy loss during the quantization process.

[0133] In the embodiments of the present application, the electronic device can set the number of data to be quantified included in each data group according to actual needs to adjust the quantization accuracy, which is not specifically limited herein.

[0134] S4022, write the K3 vector and the quantization parameter matrix s1 into the key buffer, and write the V3 vector and the quantization parameter matrix s2 into the value buffer.

[0135] In the embodiments of the present application, the key buffer stores the K cache vector and the corresponding quantization parameter matrix, and the value buffer stores the V cache vector and the corresponding quantization parameter matrix. Among them, the K cache vector and the V cache vector are also quantized vectors and have the same accuracy as the K3 vector and the V3 vector.

[0136] Exemplarily, refer to Figure 10 , which shows the storage schematic diagram of the key buffer before and after the electronic device writes the K1 vector. Among them, the accuracies before and after quantization are FP16 and Int8 respectively. As Figure 10 shown, before the electronic device writes the K1 vector into the key buffer, the key buffer includes K cache vectors 1 to K cache vector P, and quantization parameter matrices 1 to quantization parameter matrix P. During the process of writing the K1 vector, the K1 vector can be first converted into a K3 vector of Int8 and a quantization parameter matrix s1 of FP16, and then the K3 vector and the quantization parameter matrix s1 are written into the key buffer together.

[0137] Among them, the process of the electronic device writing the V1 vector into the value buffer is the same as Figure 10 the above process and will not be elaborated here.

[0138] Refer to Figure 11 , which is a flowchart of a feature extraction method provided by the embodiments of the present application Figure 3 . As Figure 11 shown, S403 includes S4031 to S4032.

[0139] S4031, the electronic device reads the key buffer to obtain the K4 vector and the quantization parameter matrix S3, and reads the value buffer to obtain the V4 vector and the quantization parameter matrix S4.

[0140] Among them, the K4 vector includes a K cache vector and a K3 vector (i.e., the quantized K1 vector), and the quantization parameter matrix S3 includes the quantization parameter matrix corresponding to the K cache vector and the quantization parameter matrix s1. The V4 vector includes a V cache vector and a V3 vector (i.e., the quantized V1 vector), and the quantization parameter matrix S4 includes the quantization parameter matrix corresponding to the V cache vector and the quantization parameter matrix s1. In a possible design, the K4 vector can also be referred to as the fourth key vector, the V4 vector can also be referred to as the fourth value vector, the quantization parameter matrix S3 can also be referred to as the third quantization parameter matrix, and the quantization parameter matrix S4 can also be referred to as the fourth quantization parameter matrix.

[0141] Exemplarily, still as Figure 10 shown, the K4 vector includes K cache vectors 1 to K cache vector P and the K3 vector, and the quantization parameter matrix S3 includes quantization parameter matrices 1 to quantization parameter matrix P and the quantization parameter matrix s1.

[0142] S4032, the electronic device performs inverse quantization on the K4 vector through the quantization parameter matrix S3 to obtain the K2 vector, and performs inverse quantization on the V4 vector through the quantization parameter matrix S4 to obtain the V2 vector.

[0143] For symmetric quantization, the quantization value and the data to be quantized satisfy the formula: r = q·s.

[0144] For non-symmetric quantization, the quantization value and the data to be quantized satisfy the formula: r = (q - z)·s.

[0145] Exemplarily, still as Figure 10 shown, during the process of the electronic device accessing the key buffer, it can first read the K4 vector in Int8 format and the quantization parameter matrix S3 in FP16 format from the key buffer. The K4 vector specifically includes K cache vectors 1 to K cache vector P and the K3 vector, and the quantization parameter matrix S3 includes quantization parameter matrices 1 to quantization parameter matrix P and the quantization parameter matrix s1. Then, the electronic device can perform inverse quantization operation on the K4 vector through the quantization parameter matrix S3 to obtain the K2 vector.

[0146] In the embodiments of the present application, through the above quantization and inverse quantization operations, the data can be stored in a data type with lower precision and used in a data type with higher precision, thereby reducing the memory space required for storing data while ensuring the calculation accuracy during the operation process.

[0147] The embodiments of the present application also provide a computer storage medium, which includes computer instructions. When the computer instructions run on the above electronic device, the electronic device is enabled to execute each function or step executed by the electronic device in the above method embodiments.

[0148] The embodiments of the present application also provide a computer program product. When the computer program product runs on a computer, the computer is enabled to execute each function or step executed by the electronic device in the above method embodiments.

[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and conciseness of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0150] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, indirect coupling or communication connection of devices or units, and can be in electrical, mechanical or other forms.

[0151] The units described as separate components may or may not be physically separated. The components displayed as units may be one physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0152] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0153] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0154] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claimed rights.

Claims

1. A feature extraction method, characterized in that, Applied to an electronic device, the method includes: Obtaining a query vector, a first key vector, and a first value vector of an element to be processed; Writing the first key vector and the first value vector into a buffer; wherein, the buffer is used to store a key cache vector and a value cache vector, the key cache vector includes key vectors of elements associated with the element to be processed, and the value cache vector includes value vectors of elements associated with the element to be processed; Obtaining a self-attention representation of the element to be processed based on the query vector, a second key vector, and a second value vector; wherein, the second key vector includes the key cache vector and the first key vector, the second value vector includes the value cache vector and the first value vector, and the second key vector and the second value vector are obtained by accessing the buffer.

2. The method according to claim 1, wherein The electronic device includes a self-attention module, and obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector includes: Obtaining the self-attention representation of the element to be processed by the self-attention module based on the query vector, the second key vector, and the second value vector.

3. The method according to claim 2, wherein The self-attention module includes a first matrix multiplication operator, an addition operator, a normalization operator, and a second matrix multiplication operator; Obtaining the self-attention representation of the element to be processed by the self-attention module based on the query vector, the second key vector, and the second value vector includes: Performing matrix multiplication on the query vector and the transposed vector of the second key vector by the first matrix multiplication operator to obtain a matrix multiplication result; Superimposing a preset mask and the result obtained by scaling the matrix multiplication result by the addition operator to obtain a superimposed result; Performing normalization processing on the superimposed result by the normalization operator to obtain a normalized processing result; Performing matrix multiplication on the normalized processing result and the second value vector by the second matrix multiplication operator to obtain the self-attention representation of the element to be processed.

4. The method according to claim 3, wherein The buffer includes a key buffer, and the key buffer includes the key cache vector. Before obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector, the method further includes: Accessing the key buffer in a first order by the first matrix multiplication operator to obtain the transposed vector of the second key vector, and the first order is opposite to the data storage order of the key cache vector.

5. The method according to claim 3, wherein The buffer includes a key buffer, and the key buffer includes the key cache vector. The self-attention module further includes a transpose operator. Before obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector, the method further includes: Accessing the key buffer in a second order by the transpose operator to obtain the second key vector, and the second order is the same as the data storage order of the key cache vector; Exchanging the rows and columns of the second key vector by the transpose operator to obtain the transposed vector of the second key vector.

6. The method according to any one of claims 2-5, characterized in that The buffer further includes a value buffer, and the value buffer includes the value cache vector. Before obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector, the method further includes: Accessing the value buffer in a second order by the second matrix multiplication operator to obtain the second value vector, where the second order is the same as the data storage order of the key cache vector.

7. The method according to any one of claims 2-6, characterized in that, The self-attention module further includes a multiplication operator. The obtaining the self-attention representation of the element to be processed by the self-attention module based on the query vector, the second key vector, and the second value vector further includes: Scaling the matrix multiplication result by the multiplication operator.

8. The method according to any one of claims 1-7, characterized in that The buffer is further configured to store the quantization parameter matrix corresponding to the key cache vector and the quantization parameter matrix corresponding to the value cache vector; The writing the first key vector and the first value vector into the buffer includes: Performing a quantization operation on the first key vector to obtain a third key vector and a first quantization parameter matrix; wherein, the first quantization parameter matrix is used to indicate the mapping relationship between the first key vector and the third key vector, and the precisions of the first key vector and the third key vector are different; Performing a quantization operation on the first value vector to obtain a third value vector and a second quantization parameter matrix; wherein, the second quantization parameter matrix is used to indicate the mapping relationship between the first value vector and the third value vector, and the precisions of the first value vector and the third value vector are different; Writing the third key vector, the first quantization parameter matrix, the third value vector, and the second quantization parameter matrix into the buffer; Before obtaining the self-attention representation of the element to be processed based on the query vector, the second key vector, and the second value vector, the method further includes: Obtaining the second key vector and the second value vector based on a fourth key vector, a third quantization parameter matrix, a fourth value vector, and a fourth quantization parameter matrix; wherein, the fourth key vector includes the third key vector and the key cache vector, the third quantization parameter matrix includes the first quantization parameter matrix and the quantization parameter matrix corresponding to the key cache vector, the fourth value vector includes the third value vector and the value cache vector, and the fourth quantization parameter matrix includes the second quantization parameter matrix and the quantization parameter matrix corresponding to the value cache vector.

9. The method according to claim 8, wherein The obtaining the second key vector and the second value vector based on the fourth key vector, the third quantization parameter matrix, the fourth value vector, and the fourth quantization parameter matrix includes: Performing an inverse quantization operation on the fourth key vector by the third quantization parameter matrix to obtain the second key vector; Performing an inverse quantization operation on the fourth value vector by the fourth quantization parameter matrix to obtain the second value vector.

10. An electronic device, characterized in that, The electronic device includes: a memory and one or more processors; the memory is coupled to the processor; the memory is used to store computer program code, and the computer program code includes computer instructions. When the computer instructions are executed by the electronic device, the electronic device is caused to execute the method according to any one of claims 1-9.

11. A computer-readable storage medium, characterized in that, Computer instructions are stored in the computer-readable storage medium. When the computer instructions run in an electronic device, the electronic device is caused to execute the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Intention recognition method, device, readable medium and electronic equipment

    CN114090740A

  • Neural network model optimization method and related equipment

    CN115841134A

  • Neural network quantification method and device, chip, board card and equipment

    CN115841136A

  • Selective batching of inference system for converter-based task generation

    CN116245181A

  • Calculation method and device of neural network model, electronic equipment and storage medium

    CN117273084A