User interaction method, device, electronic device, medium and product

The weight of the multi-head attention mechanism is quantified by a gradient-based post-training quantization method and stacking multiple decoder blocks in the text inference large model, solving the problems of limited accuracy and slow inference speed when converting MHA to GQA, and achieving improvements in accuracy and speed without pre-training or fine-tuning.

CN119808719BActive Publication Date: 2025-05-23INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510295109.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-23
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In the prior art, when the multi-head attention mechanism (MHA) is converted to grouped query attention (GQA), the accuracy is limited, and pre-training or fine-tuning is required to solve the accuracy loss, which in turn affects the inference speed.

Method used

By using a gradient-based post-training quantization method, weights in the multi-head attention mechanism are quantified and multiple decoder blocks are stacked in the text inference large model to achieve feature extraction and semantic associations, accelerating model inference without pre-training or fine-tuning.

Benefits of technology

While maintaining a certain accuracy, by quantizing the architecture of multi-head attention weights and multiple decoder blocks superposition, model inference can be accelerated, solving the problems of limited accuracy and slow inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808719B_ABST
    Figure CN119808719B_ABST
Patent Text Reader

Abstract

The present invention discloses a user interaction method, device, electronic device, medium and product, which relate to the technical field of electronic digital data processing. The invention can utilize a quantized multi-head attention mechanism and a plurality of decoder blocks superimposed architecture to perform feature extraction and semantic association on the question text input by the user in successive cycles to obtain more accurate reasoning data, and update the multi-head attention weights in a quantized manner, which can accelerate the model reasoning while maintaining a certain accuracy without pre-training or fine-tuning. Therefore, the technical problem in the related technology that the accuracy is limited and further pre-training or fine-tuning is required to solve the accuracy loss, thereby affecting the reasoning speed, can be solved, and the technical effect of simultaneously ensuring the model reasoning speed and the model accuracy can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a user interaction method, device, electronic equipment, medium and product. Background Art

[0002] In related technologies, large-scale language models can maintain accuracy by pooling several adjacent heads and then continuing pre-training. They can also use singular value decomposition to retain some singular values ​​and fine-tune by adding low-rank matrices within the GQA (Grouped Query Attention) group to solve the problem of accuracy loss. They can also fine-tune by judging the similarity between different attention heads and grouping attention heads with high similarity. Although the above three methods can all convert MHA to GQA to take into account both model performance and inference speed, they are all aimed at processing weights, and the accuracy they can achieve is limited, and they need to continue pre-training or fine-tuning to solve the accuracy loss. Summary of the invention

[0003] The present invention provides a user interaction method, device, electronic device, medium and product to at least solve the technical problem in the related art that the accuracy is limited and further pre-training or fine-tuning is required to solve the accuracy loss, thereby affecting the inference speed.

[0004] The present invention provides a user interaction method, comprising: receiving a user's question text; inputting the question text into a text reasoning model, so as to output reasoning data corresponding to the question text after the question text gradually passes through a plurality of decoder blocks stacked in the text reasoning model, wherein the multi-head attention weights in the decoder blocks are obtained by gradient-based post-training quantization; and pushing the reasoning data to the user.

[0005] The present invention also provides a user interaction device, comprising: a receiving module for receiving a user's question text; a processing module for inputting the question text into a text reasoning model, so as to output reasoning data corresponding to the question text after the question text gradually passes through multiple decoder blocks stacked in the text reasoning model, wherein the multi-head attention weights in the decoder blocks are obtained by gradient-based post-training quantization; and a reply module for pushing the reasoning data to the user.

[0006] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any one of the above-mentioned user interaction methods when executing the computer program.

[0007] The present invention also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned user interaction methods are implemented.

[0008] The present invention also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned user interaction methods when executed by a processor.

[0009] Through the present invention, a quantized multi-head attention mechanism and a plurality of decoder block superposition architecture can be used to perform feature extraction and semantic association on the question text input by the user in successive cycles to obtain more accurate reasoning data, and the multi-head attention weights are updated in a quantized manner, which can accelerate the model reasoning while maintaining a certain accuracy without pre-training or fine-tuning. Therefore, the technical problem of limited accuracy in related technologies, which requires continued pre-training or fine-tuning to solve the loss of accuracy and thus affects the reasoning speed, can be solved, thereby achieving the technical effect of simultaneously ensuring the model reasoning speed and model accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0011] Figure 1 A flowchart of a user interaction method provided by an embodiment of the present invention;

[0012] Figure 2 Schematic diagram of the principles of three multi-head attention mechanisms in related technologies;

[0013] Figure 3 A schematic diagram of the principle of a large text reasoning model of a user interaction method provided according to an embodiment of the present invention;

[0014] Figure 4 A schematic diagram of the principle of a large text reasoning model of a user interaction method provided according to another embodiment of the present invention;

[0015] Figure 5 A schematic diagram of the principle of average pooling according to an embodiment of the present invention;

[0016] Figure 6 A schematic diagram of the principle of a feedforward neural network provided according to an embodiment of the present invention;

[0017] Figure 7A flowchart of building and launching a large text reasoning model according to an embodiment of the present invention;

[0018] Figure 8 A schematic diagram of the structure of a user interaction device provided by an embodiment of the present invention.

[0019] Among them, 10 is a user interaction device, 100 is a receiving module, 200 is a processing module, and 300 is a reply module. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] It should be noted that, in the description of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0022] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0023] In conjunction with the specific application environment architecture or the specific hardware architecture on which the execution of the user interaction device method depends, the specific application environment architecture or the specific hardware architecture is described herein.

[0024] An embodiment of the present invention provides a user interaction method.

[0025] like Figure 1 FIG. 1 is a flowchart of a user interaction method according to an embodiment of the present invention, wherein the user interaction method comprises the following steps:

[0026] In step S101, a question text from a user is received.

[0027] It is understandable that online customer service is a role that provides services to customers through online channels. For larger companies, there are many customers, and the demands for customer service are increasing. However, there is a great waste of manpower if every question raised by a user is answered manually. Therefore, the primary way to reduce manpower waste is to use automatic replies to infer and answer questions raised by users.

[0028] The embodiment of the present invention can pre-build a large text reasoning model to reason about the question text after receiving the user's question text, and reply according to the reasoning result, thereby realizing automatic reply, thereby saving labor costs, reducing the technical knowledge requirements for manual customer service, and speeding up the resolution of user problems.

[0029] In the actual implementation process, the user can enter the relevant customer service page of the enterprise and enter his own questions. At this time, the embodiment of the present invention can receive the user's question text and perform reasoning based on the question text.

[0030] In step S102, the question text is input into the text reasoning model to output the reasoning data corresponding to the question text after the question text gradually passes through multiple decoder blocks stacked in the text reasoning model, wherein the multi-head attention weights in the decoder blocks are obtained by gradient-based post-training quantization.

[0031] It is understandable that with the continuous development of deep learning, large models have great performance in fields such as natural language processing and computer vision. The development and application of large models mainly include three process steps: pre-training, fine-tuning, and reasoning.

[0032] Reasoning of large models is the practical application and inference work carried out based on pre-trained models.

[0033] The two stages of reasoning include the prefill stage and the decoding stage.

[0034] The pre-filling phase can calculate all user inputs and then generate the corresponding key-value cache. It is like a preparation for the reasoning process and can provide the required context information for the subsequent decoding phase.

[0035] In the decoding phase, the model can gradually generate word units in the output sequence based on the key-value cache generated in the pre-filling phase. At the same time, the key values ​​corresponding to the newly generated word units will be placed in the key-value cache, and the internal state of the model will be updated for the next decoding.

[0036] In related technologies, methods for accelerating the inference performance of large models may include:

[0037] The first type of method compresses the size of the model and the user's input sequence, thereby reducing redundant or irrelevant information to achieve the purpose of acceleration.

[0038] The second method is to compress the space size of the key-value cache to achieve the acceleration effect.

[0039] Among them, the process of compressing the space size of the key-value cache can include methods such as intra-layer sharing, inter-layer sharing, and inter-layer fusion. Among them, GQA is a common method for intra-layer sharing. It divides the query (query parameter) header into G groups, each group shares a key and value header. In this way, GQA is between MQA (Multi-Query Attention) and MHA (Multi-Head Attention) in terms of performance and speed, which can ensure a relatively fast inference speed and maintain a high model performance.

[0040] Among them, the relationship between MHA, GQA and MQA can be as follows Figure 2 As shown in the figure, using the GQA method for reasoning is usually strongly related to training. Usually, only when the trained model implements GQA can GQA be used for reasoning acceleration during reasoning, and the number of GQA groups is fixed. This method makes the acceleration effect of the model and the range of model selection affected by business requirements and thus has certain limitations. Currently, there are some technologies that can convert MHA to GQA, but it still requires continued pre-training or fine-tuning to solve the problem of accuracy loss.

[0041] Based on the above problems, the embodiments of the present invention can consider the activation value and use the quantization idea of ​​GPTQ (Gradient-based Post-training Quantization) to realize the conversion of MHA into GQA. It can accelerate model reasoning while maintaining a certain accuracy without pre-training or fine-tuning.

[0042] That is to say, when constructing the text reasoning big model of the embodiment of the present invention, it is possible to convert MHA into GQA by drawing on the idea of ​​GPTQ quantization, so that after the MHA in the decoder blocks stacked in the text reasoning big model is converted into GQA, the question text can gradually pass through multiple new decoder blocks, realize the gradual extraction of features and semantics, complete the reasoning, and output the reasoning data corresponding to the question text.

[0043] Optionally, in one embodiment of the present invention, the question text is input into a text reasoning model to output reasoning data corresponding to the question text after the question text gradually passes through multiple decoder blocks stacked in the text reasoning model, including: splitting the question text to obtain multiple word elements that meet preset processing conditions; mapping the multiple word elements to the vector space of the first target dimension respectively to obtain multiple word vectors; inputting the multiple word vectors into a loop processing layer formed by stacking multiple decoder blocks to output processed data corresponding to the question text; decoding the processed data to obtain reasoning data.

[0044] like Figure 3 and Figure 4 As shown, the structure of the text reasoning large model in the embodiment of the present invention can be an online LLM (Large Language Model) reasoning architecture.

[0045] Taking Llama2 13B as an example, Llama2 13B is pre-trained on publicly available online data sources, and then optimized through supervised fine-tuning (SFT) and reinforcement learning with human feedback (RLHF) to improve its performance on specific tasks. Context window: The context window length of Llama2 13B is 4096, which can handle longer text content and provide more comprehensive contextual understanding.

[0046] Llama2 13B is a large language model based on the Transformer architecture. It adopts a decoder-only structure (a model structure with only a decoder part but no encoder part), that is, only the decoder part of the Transformer is used.

[0047] Input layer: The input layer segments the input text and converts the segmented results into corresponding word vector representations, that is, multiple word units are obtained. Multiple word units will be used as input to the model.

[0048] Attention mechanism, Llama2 13B adopts a grouped query attention mechanism. This mechanism shares key and value projections in multi-head attention models, thereby reducing the memory costs associated with caching. By using GQA, larger models can maintain performance while optimizing memory usage.

[0049] Feed-Forward Neural Network, after the attention mechanism, the model uses FFN (Feed-Forward Neural Network) to further transform the representation of each position. FFN consists of two linear transformations and an activation function, and the activation function usually uses SwiGLU (Sigmoid-Weighted Linear Unit, adaptive activation function).

[0050] The output layer converts the final representation of the model into a probability distribution of the predicted next word. These representations are usually converted into probabilities using a softmax function so that the most likely next word can be selected.

[0051] like Figure 2 As shown, the text reasoning model of the embodiment of the present invention includes a stack of multiple decoder blocks in addition to the input layer and the output layer, wherein the structure (attention mechanism, feedforward neural network, etc.) in each decoder block can be as follows Figure 3 shown.

[0052] Generally speaking, the structure of the decoder block can include:

[0053] Masked Multi-Head Self-Attention: It is necessary to mask "future information" because future information cannot be seen during prediction, so the current word and subsequent words must all be marked.

[0054] Encoder-Decoder Attention: This layer allows the decoder to pay attention to the output of the encoder, so that the decoder can generate the output sequence based on the information of the input sequence. Here, the query (Q) comes from the decoder output of the previous layer, while the key (K) and value (V) come from the output of the encoder. This setup allows the decoder to "query" different parts of the input sequence in order to more accurately generate the next token.

[0055] Feed-Forward Network: The feed-forward fully connected layer here is exactly the same as that in the Encoder module, consisting of two linear transformations with a Relu activation function in the middle.

[0056] Workflow

[0057] Input processing: The input received by the bottom layer of blocks is the data after being processed by MASK. The input patterns received by other layers of blocks are consistent, which are all the outputs of the previous layer.

[0058] Multi-head self-attention layer calculation: In this layer, each word will be paid attention to the words before it in the sentence to obtain a weighted representation. This process will be performed in parallel on multiple heads, and the results are finally concatenated.

[0059] Encoder-decoder attention layer calculation: The Q of this layer comes from the output of the previous step, and K and V come from the output of the Encoder. By calculating the attention weights, the Decoder can focus on the part of the input sequence that is most relevant to the currently generated word.

[0060] Feedforward fully connected layer calculation: The output of the previous step is sent to the feedforward fully connected layer, and the features are further transformed after two linear transformations and a Relu activation function.

[0061] Output: After the above steps, the output of the Decoder Block will be used as the input of the next Block, or in the last block, as the output of the entire Decoder to generate the final prediction result, such as the target language sentence in the translation task.

[0062] In an embodiment of the present invention, the input data can be first split and processed, then embedded (embedding) to form vectorized data, and finally normalized (RMSNorm) -> multi-head attention mechanism (MHA) -> normalized (RMSNorm) -> feedforward neural network (FeedForward SwishGLU) and other series of processing processes, and finally after completing the multi-layer data processing, the characters are decoded through normalization (RMSNorm) -> embedding (embedding) -> activation function (Softmax).

[0063] Among them, when performing text segmentation, the embodiment of the present invention can be implemented using a tokenizer (Tokenizier), which can be used to convert the text into a data form that can be processed by the model, that is, to segment the text and build a vocabulary, and at the same time map the text to an integer sequence.

[0064] The tokenizer can decompose the input question text data (such as a paragraph of text) into small units, which are called tokens. For example, the sentence "I love the green leaves in spring" may be decomposed into tokens such as "I", "love", "spring", and "green leaves". The purpose of this is to convert the original data into a format that the model can process, because large models usually cannot directly process continuous text or other complex data structures, and need to be split into basic units first.

[0065] After being processed by the tokenizer, each word in the data is mapped to a low-dimensional vector space. This process is called embedding. Through embedding, the originally discrete word units are converted into vector form, and these vectors can capture the semantic relationship between word units. For example, words with similar semantics will be closer in the vector space.

[0066] Normalization helps stabilize the training process and improve the convergence speed and performance of the model. Before the word vector or processed data enters the multi-head attention layer or feedforward neural network or after the processed data is output, the data can be normalized through normalization so that the inputs of different neurons have similar distributions, thereby avoiding the output of some neurons being too large or too small, affecting the training of the entire model.

[0067] Among them, the formula for normalization processing can be:

[0068] ,

[0069] ,

[0070] in, d is the feature dimension, , is a normalization hyperparameter, which is determined after the model is trained. , The value of , The dimensions and d same. Represents the element-level multiplication operation, that is, element-by-element multiplication. RMSNorm(x) represents the root mean square normalization formula, represents the normalized input vector, RMS(x) represents the root mean square formula, Represents the input vector of the model.

[0071] The multi-head attention mechanism allows the model to focus on different parts of the input data at the same time and obtain information from multiple representation subspaces. For example, when processing text, it can focus on different words in a sentence at the same time to better understand the semantic relationship between words. The multi-head attention mechanism captures semantic information at different levels through multiple "heads" (i.e., multiple parallel attention calculations), and then combines this information to provide richer feature representations for subsequent processing.

[0072] Feedforward neural network is a basic structure in neural network, which processes input data through a series of linear transformations and nonlinear activation functions.

[0073] SwishGLU is a specific combination of activation functions. The role of the activation function is to introduce nonlinear characteristics to the neural network, so that the model can learn complex functional relationships. In this process, the feedforward neural network structure combines the SwishGLU activation function to further transform the features and extract information from the previously processed data.

[0074] When processing data, each layer processes the output of the previous layer and gradually extracts the features of the data, making the data representation more and more advanced and better adapted to the task requirements of the model, such as accurately predicting the next character in a language model.

[0075] Optionally, in one embodiment of the present invention, before the question text gradually passes through multiple decoder blocks stacked in the text reasoning model, it also includes: quantizing the multi-head attention mechanism in the initial decoder block to obtain a quantized multi-head attention module; constructing a decoder block using the multi-head attention module and a feedforward neural network structure to use multiple decoder blocks to perform step-by-step reasoning on the question text.

[0076] like Figure 4 As shown, each decoder block in the embodiment of the present invention includes a multi-head attention module and a feedforward neural network including a multi-head attention mechanism. In the embodiment of the present invention, the multi-head attention mechanism in the multi-head attention module can convert MHA into GQA by drawing on the idea of ​​GPTQ quantization, thereby obtaining a new multi-head attention module to optimize the initial decoder block.

[0077] Optionally, in one embodiment of the present invention, the multi-head attention mechanism in the initial decoder block is quantized to obtain a quantized multi-head attention module, including: grouping the multiple attention heads in the initial decoder block to obtain multiple attention groups, wherein each attention group shares the same set of key vector weights and value vector weights, and the query vector weights of the attention heads in each attention group are consistent; constructing an error constraint on the weight of the key vector when the weight of the query vector in each attention head does not change; constructing a corresponding key vector cost function based on the error constraint to update the weight of the key vector of each attention group; updating the weight of the value vector of each attention group based on the updated key vector weight; updating the initial decoder block based on the updated key vector weight and the updated value vector weight to obtain a decoder block.

[0078] Among them, constructing the error constraint of the weight of the key vector includes: obtaining a calibration sample in the same application scenario as the question text; using the calibration sample to construct a calibration sample set; obtaining the calibration weight of the key vector based on the unquantized multi-head attention mechanism and the calibration sample set in the initial decoder block; and constructing the error constraint using the calibration weight and the weight of the key vector in the quantized multi-head attention module.

[0079] The formula of MHA in the multi-head attention module before conversion can be:

[0080] ,

[0081] in,

[0082] ,

[0083] In an embodiment of the present invention, a calibration sample in the same application scenario as the question text can be input into the MHA to obtain data for calibration. For example, if the model of the embodiment of the present invention is applied to an e-commerce platform, the calibration sample in the same application scenario should include the reply text of the question text related to the items sold on the e-commerce platform.

[0084] therefore, , , are the calibration weights for query vector, key vector and value vector respectively, are the weights of the linear layer.

[0085] Among them, the embodiment of the present invention can use the Softmax activation function to calculate the attention weight. For the attention calculation of each head, the original score is first obtained by calculating the dot product of the query Q and the key K, and then these scores are divided by a scaling factor (usually the square root of the key vector dimension) to stabilize the output of the Softmax function, and finally the scaled scores are converted into probability distributions, i.e., attention weights, through the Softmax function. These attention weights represent the importance of each tag to the current query tag, thereby realizing dynamic adjustment of the relationship and dependency between different tags in the input sequence.

[0086] Among them, the formula of Softmax can be:

[0087] Softmax( )= .

[0088] Among them, Softmax( ) represents the result of the Softmax function, which represents the input vector z Middle i The output probability value corresponding to the element. In multi-classification problems, this value can be interpreted as the input belonging to the i The probability of the class. Represents the input vector z No. i In a neural network, this is usually the output value of a node (or neuron), which is calculated by taking the weighted sum of the previous layer and adding a bias term. K Represents the input vector z Length.

[0089] After obtaining data that can be used for calibration, the embodiment of the present invention can gradually replace MHA with GQA.

[0090] When GQA is used for calculation, the weights within the group should be consistent. Assuming that there are 8 groups, and each group has five attention heads, then , Etc. The situation in other groups is similar.

[0091] In related technologies, the weights in a group can be averaged by pooling, but the error caused by the average pooling operation is large. Because the weights in each group are less correlated, weights with strong correlation can be grouped by weight clustering. This can reduce the error after the pooling operation. SVD (Singular Value Decomposition) can also be used to decompose and retain the weights with high importance within or between groups. However, the GQA errors obtained by these methods are relatively large. In the end, fine-tuning or re-pre-training is required.

[0092] In the present invention, from the perspective of the attention formula, If there is no change, if you keep If unchanged, it should be minimized The error is that we need to minimize the mean square error, that is, the key vector cost function is:

[0093] ,

[0094] in, K is the key vector, is the calibration weight of the key vector, is the weight of the key vector, that is, the weight within the GQA group, and the characteristic of the weight within the group is that each The weights of the heads are the same, so the same weight only needs to be calculated using one weight to get the weights of other heads in the group, thereby saving calculation amount and reducing cache space.

[0095] This cost function means that after GQA grouping Should be consistent with the original MHA The calculation results are consistent, and finally Be consistent.

[0096] Furthermore, the embodiment of the present invention can use the parameters trained by the feedforward neural network to perform Taylor expansion on the cost function, and the obtained results are as follows:

[0097] ,

[0098] in, is the error term, The compensation matrix of weights, is the gradient, is a second-order Hessian matrix, which is a positive definite symmetric matrix. The cost function That is .

[0099] ,

[0100] in, is the weight matrix obtained after average pooling within the group, that is, the key vector weight.

[0101] The change of loss function is:

[0102] .

[0103] Due to the training is 0, and can be ignored, so

[0104] ,

[0105] The embodiments of the present invention can e m The weight of the position is assigned to e q , e q The matrix is ​​similar to the following , when its transpose is The result of multiplication Indicates selection The error of the qth column. e m Indicates location m The weight assignment of e q Indicates location q The weight assignment of .

[0106] So ,in, is the difference between the weight of the mth column and the weight of the qth column. This formula means that The difference compensation is given , so when the weight is updated, the qth column position will become the weight of the mth column position. At this time, the problem is transformed into Minimize under the constraints of The Lagrange multiplication operator is introduced to transform the problem of constraining each set of parameters to be the same into an unconstrained problem:

[0107] .

[0108] Where H represents the Hessian matrix, represents the Lagrange multiplier.

[0109] right Taking the derivative we get: .

[0110] get , this result means that if the loss function is minimized, the compensation required for the weight is .

[0111] The embodiment of the present invention can be recorded is the value of the inverse matrix of the Hessian matrix at the qth row and qth column position. , substitute We can get:

[0112] .

[0113] The value added to the loss function is:

[0114] .

[0115] For the cost function, the Hessian matrix of each row is the same, that is, It is only related to the input data and is a positive definite symmetric matrix. If a matrix is ​​a positive definite symmetric matrix, it can be decomposed by Cholesky matrix, and there is a lower triangular matrix with positive diagonal elements. Make Established.

[0116] Assumptions , , .

[0117] Can get

[0118] ,

[0119] in, , , ,

[0120] Define a series of phalanx: ,in ,

[0121] .

[0122] There is also a series of lower triangular matrices :

[0123] ,

[0124] It can be found that:

[0125] ,

[0126] and

[0127] ,

[0128] Finally decomposed into:

[0129] ,

[0130] In this algorithm, the first element is deleted each time, and the inverse matrix of the Hessian matrix is ​​recorded as .

[0131] .

[0132] At the same time, each time the first element is removed,

[0133] .

[0134] therefore,

[0135] .

[0136] For The inverse matrix Perform Cholesky decomposition to obtain After that, the weight values ​​can be updated without explicit update every time.

[0137] Therefore, the pseudo code for updating weights can be obtained as follows:

[0138] Algorithm 1 Average W given inversr Hessian , . Calculate the inverse Hessian matrix

[0139] Create Q matrix, store

[0140] Create E matrix, store

[0141] The inverse Hessian matrix break down

[0142] For k = 0,1,2, …,M do

[0143] for i = 0, 1*B, 2*B, …,G*B do

[0144] for j = i, … , i+B-1 do

[0145] deal with

[0146] calculate

[0147] calculate And update the attention head weight

[0148] end for

[0149] Update the weights of other attention heads in the group

[0150] end for

[0151] end for

[0152] in, is the inverse matrix of the Hessian matrix, and For each group Matrix, Llama2 13B original The size is , the original attention head size is 40, assuming it is divided into 8 groups, then the weight in each group is , , ,at the same time , , .

[0153] The average pooling operation diagram can be shown as follows Figure 5 shown.

[0154] The embodiment of the present invention may assume that , and To form a group, the average pooling operation is to average the corresponding elements of each attention head and assign them to each , will soon , and The first row of the element at position 1, the element at position 5, and the element at position 9 are averaged and assigned to 1, 5, 9, 2, 6, 10, and other positions in the same way, and then processed row by row. This process can ensure that each The elements are the same.

[0155] So far, the embodiment of the present invention can complete the data update of the weight of the key vector.

[0156] Optionally, in one embodiment of the present invention, the weight of the value vector of each attention group is updated based on the weight of the updated key vector, including: constructing a value vector cost function of the value vector using the weight of the updated key vector; and obtaining the weight of the updated value vector based on the value vector cost function.

[0157] Among them, the value vector cost function is:

[0158] ,

[0159] Among them, Softmax is the activation function, Q is the query vector, V is the value vector, is the weight of the query vector, is the weight of the value vector, is a vector of calibration weights.

[0160] Further, Updates and The update is similar, and the above process is also used The update is different. The update requires the use of After the data is updated, the cost function will also change. The update cost function is:

[0161] ,

[0162] That is Update if known , it should also be noted that when updating When the inverse matrix of the Hessian matrix is ​​not , , but The value becomes .

[0163] Optionally, in one embodiment of the present invention, multiple word vectors are input into a recurrent processing layer formed by stacking multiple decoder blocks to output processed data corresponding to the question text, including: using the attention mechanism of the current decoder block to respectively calculate the attention scores of the multiple word vectors; obtaining corresponding attention output data based on the attention scores; performing linear transformation and nonlinear activation function on the attention output data to perform feature transformation and information extraction on the attention output data, and outputting intermediate processed data; inputting the intermediate processed data into the next encoder block until all decoder blocks are processed to obtain processed data.

[0164] The feedforward neural network and SwishGLU architecture can be Figure 6 shown.

[0165] like Figure 6 As shown, the data output after normalization processing is used as the input of the feedforward neural network and SwishGLU architecture, and passes through two routes, gate linear transformation (Gate_Linear) -> Swish function and dimension increase linear transformation (Up_Linear), and then performs Hadamard product, and finally passes through dimension reduction linear transformation (Down_Linear) as output and is transmitted to the next unit.

[0166] Among them, the Swish function is as follows:

[0167] ,

[0168] in, For input, are the parameters that have been learned by the feedforward neural network. σ ( βx ) is the Sigmoid function σ application.

[0169] pass The function can still generate non-zero gradients when the input is negative, avoiding the "neuron death" problem of traditional activation functions such as ReLU when the input is negative. At the same time, Gate_Linear, Up_Linear and Down_Linear are all MLP networks.

[0170] Optionally, in one embodiment of the present invention, the processed data is decoded to obtain inference data, including: normalizing the processed data to obtain normalized processed data; mapping the normalized processed data to a representation space that satisfies preset decoding conditions to obtain data to be decoded; and decoding the data to be decoded to obtain inference data.

[0171] After layer-by-layer processing, the embodiment of the present invention may perform a normalization operation again to make a final adjustment to the data so as to better perform subsequent decoding operations.

[0172] The embodiment of the present invention may perform some form of embedding operation again on the previously processed data, possibly to convert the data into a representation space that is more suitable for decoding.

[0173] The embodiment of the present invention may also utilize a Softmax activation function for decoding to obtain inference data.

[0174] In step S103, the inference data is pushed to the user.

[0175] Furthermore, the embodiment of the present invention can push the inference data to the corresponding user to complete the answer to the question.

[0176] Optionally, in one embodiment of the present invention, after obtaining the inference data, it also includes: mapping the inference data into vocabulary probabilities; matching corresponding sub-words based on the vocabulary probabilities, and merging multiple sub-words into a complete text based on the sub-word sequence; converting the complete text into a display format, and pushing the complete text in the display format to the user.

[0177] The embodiment of the present invention can convert the output of the model into a probability distribution, so that each possible character has a corresponding probability. By selecting the character with the highest probability, the predicted character can be decoded from the model. For example, if the model is a language model, after decoding, the next most likely character can be determined based on the probability distribution.

[0178] The matching subwords are merged to obtain the complete text, and the corresponding display format is determined according to the user's needs, that is, it is converted into a text format that can be sent online, and the complete text is pushed to the user to complete the user's inquiry.

[0179] In summary, it can be seen that the embodiment of the present invention can use the idea of ​​GPTQ quantization to build a large text reasoning model based on the conversion of MHA to GQA. Figure 7 As shown, the construction and launch process of the text reasoning large model in an embodiment of the present invention is illustrated by example.

[0180] Step S701: define an online LLM reasoning architecture.

[0181] like Figure 3 and Figure 4 As shown, the embodiment of the present invention can first perform word segmentation on the input data, then embed it to form vectorized data, and finally go through a series of processing processes such as normalization processing -> multi-head attention module -> normalization processing -> feedforward neural network and SwishGLU function, and finally after completing 40 layers of data processing, normalization processing -> embedding -> Softmax function is used to decode the characters.

[0182] Step S702: PTQ (Post-Training Quantization Calibration) calibration.

[0183] The calibration dataset selects C4-validation as the calibration dataset. The key to the calibration dataset is to select samples that can represent the model application scenarios, and use these samples to calibrate the model parameters in the GQA weight generation process to reduce performance loss. The calibration dataset is used as the input of the network, and the weights of each layer are updated in turn from Layer 0 to Layer 40 through the steps of Algorithm 1, and the updated weights are saved.

[0184] Step S703: Modify the configuration.

[0185] Modify the model configuration file.

[0186] Step S704: The model is online.

[0187] In the embodiment of the present invention, the model may be loaded into a vllm (large model deployment tool) and deployed. For example, the deployment command may be:

[0188] python3 -m vllm.entrypoints.openai.api_server --model {llama2 13Bmodel path} --swap-space 16 --disable-log-requests --port 8081

[0189] Then access it through port 8081 to get the llama2 13B model with MHA changed to GQA.

[0190] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.

[0191] An embodiment of the present invention further provides a user interaction device.

[0192] like Figure 8 As shown, the user interaction device 10 includes: a receiving module 100 , a processing module 200 and a reply module 300 .

[0193] Specifically, the receiving module 100 is used to receive a question text from a user.

[0194] The processing module 200 is used to input the question text into the text reasoning model, so as to output the reasoning data corresponding to the question text after the question text gradually passes through multiple decoder blocks stacked in the text reasoning model, wherein the multi-head attention weights in the decoder blocks are obtained by gradient-based post-training quantization.

[0195] The reply module 300 is used to push the inference data to the user.

[0196] Optionally, in one embodiment of the present invention, the processing module 200 includes: a splitting unit, an embedding unit, a processing unit and a decoding unit.

[0197] The splitting unit is used to split the question text to obtain multiple word units that meet preset processing conditions.

[0198] The embedding unit is used to map the multiple word units to the vector space of the first target dimension respectively to obtain multiple word vectors.

[0199] A processing unit is used to input multiple word vectors into a recurrent processing layer formed by stacking multiple decoder blocks to output processed data corresponding to the question text.

[0200] The decoding unit is used to decode the processed data to obtain inference data.

[0201] Optionally, in one embodiment of the present invention, the user interaction device 10 further includes: a quantification module and a construction module.

[0202] Among them, the quantization module is used to quantize the multi-head attention mechanism in the initial decoder block to obtain a quantized multi-head attention module.

[0203] A building module for building decoder blocks using a multi-head attention module and a feed-forward neural network structure to perform step-by-step reasoning on question text using multiple decoder blocks.

[0204] Optionally, in one embodiment of the present invention, the quantization module includes: a grouping unit, a construction unit, a first updating unit, a second updating unit and a third updating unit.

[0205] Among them, the grouping unit is used to group multiple attention heads in the initial decoder block to obtain multiple attention groups, wherein each attention group shares the same set of key vector weights and value vector weights, and the query vector weights of the attention heads in each attention group are consistent.

[0206] A construction unit for constructing error constraints on the weights of the key vectors while keeping the weights of the query vectors unchanged in each attention head.

[0207] The first updating unit is used to construct a corresponding key vector cost function based on the error constraint to update the weight of the key vector of each attention group.

[0208] The second updating unit is used to update the weight of the value vector of each attention group based on the weight of the updated key vector.

[0209] A third updating unit is used to update the initial decoder block based on the updated weight of the key vector and the updated weight of the value vector to obtain a decoder block.

[0210] Optionally, in one embodiment of the present invention, the construction unit includes: an acquisition subunit, a first construction subunit, a calculation subunit and a second construction subunit.

[0211] The acquisition subunit is used to acquire calibration samples in the same application scenario as the question text.

[0212] The first construction subunit is used to construct a calibration sample set using the calibration samples.

[0213] The first computing subunit is used to obtain a calibration weight of the key vector based on an unquantized multi-head attention mechanism and a calibration sample set in an initial decoder block.

[0214] The second construction subunit is used to construct error constraints using the calibration weights and the weights of the key vectors in the quantized multi-head attention module.

[0215] Optionally, in one embodiment of the present invention, the key vector cost function is:

[0216] ,

[0217] in, K is the key vector, is the calibration weight of the key vector, is the weight of the key vector.

[0218] Optionally, in one embodiment of the present invention, the second updating unit includes: a third constructing subunit and a second calculating subunit.

[0219] The third construction subunit is used to construct a value vector cost function of the value vector using the updated weight of the key vector.

[0220] The second calculation subunit is used to obtain the weight of the updated value vector based on the value vector cost function.

[0221] Optionally, in one embodiment of the present invention, the value vector cost function is:

[0222] ,

[0223] Among them, Softmax is the activation function, Q is the query vector, V is the value vector, is the weight of the query vector, is the weight of the value vector, is a vector of calibration weights.

[0224] Optionally, in one embodiment of the present invention, the processing module 200 includes: a first calculation unit, a second calculation unit, a third calculation unit and a circulation unit.

[0225] Among them, the first calculation unit is used to use the attention mechanism of the current decoder block to respectively calculate the attention scores of multiple word vectors.

[0226] The second computing unit is used to obtain corresponding attention output data based on the attention score.

[0227] The third computing unit is used to perform linear transformation and nonlinear activation function on the attention output data to perform feature transformation and information extraction on the attention output data, and output intermediate processed data.

[0228] The loop unit is used to input the intermediate processed data into the next encoder block until all decoder blocks have finished processing and obtained the processed data.

[0229] Optionally, in one embodiment of the present invention, the decoding unit includes: a first processing subunit, a second processing subunit and a decoding subunit.

[0230] The first processing subunit is used to normalize the processed data to obtain normalized processed data.

[0231] The second processing subunit is used to map the normalized processed data to a representation space that meets a preset decoding condition to obtain data to be decoded.

[0232] The decoding subunit is used to decode the data to be decoded to obtain inference data.

[0233] Optionally, in one embodiment of the present invention, the user interaction device 10 further includes: a mapping module, a matching module and a conversion module.

[0234] Among them, the mapping module is used to map the inference data into vocabulary probabilities.

[0235] The matching module is used to match the corresponding subwords based on the vocabulary probability and merge multiple subwords into a complete text based on the subword sequence.

[0236] The conversion module is used to convert the complete text into a display format and push the complete text in the display format to the user.

[0237] For the description of the features in the embodiments corresponding to the user interaction device, reference may be made to the relevant description of the embodiments corresponding to the user interaction device method, which will not be described one by one here.

[0238] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned method embodiments of the user interaction device.

[0239] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-mentioned method embodiments of the user interaction device when running.

[0240] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs.

[0241] An embodiment of the present invention further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned method embodiments of the user interaction device.

[0242] An embodiment of the present invention further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-mentioned method embodiments of the user interaction device.

[0243] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0244] The above has introduced in detail a user interaction device provided by the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A user interaction method, characterized in that: include: Receive user's question text; Input the question text into the text reasoning model, group the multiple attention heads in the initial decoder block to obtain multiple attention groups, wherein each attention group shares the same set of key vector weights and value vector weights, and the query vector weights of the attention heads in each attention group are consistent; construct an error constraint on the weight of the key vector when the weight of the query vector in each attention head does not change; construct a corresponding key vector cost function based on the error constraint to update the weight of the key vector of each attention group; update the weight of the value vector of each attention group based on the updated key vector weight; update the initial decoder block based on the updated key vector weight and the updated value vector weight to obtain a decoder block; After the question text passes through a plurality of decoder blocks stacked in the text reasoning model step by step, the reasoning data corresponding to the question text is output, wherein the multi-head attention weights in the decoder blocks are obtained by gradient-based post-training quantization; The inference data is pushed to the user.

2. The user interaction method according to claim 1, characterized in that: The step of inputting the question text into the text reasoning model to output reasoning data corresponding to the question text after the question text gradually passes through a plurality of decoder blocks stacked in the text reasoning model comprises: Splitting the question text to obtain multiple word units that meet preset processing conditions; Mapping the multiple word units to a vector space of a first target dimension respectively to obtain multiple word vectors; Inputting the plurality of word vectors into a recurrent processing layer formed by stacking the plurality of decoder blocks to output processed data corresponding to the question text; The processed data is decoded to obtain the inference data.

3. The user interaction method according to claim 2, characterized in that: Before the question text is gradually passed through a plurality of decoder blocks stacked in the text reasoning model, it also includes: Quantize the multi-head attention mechanism in the initial decoder block to obtain a quantized multi-head attention module; The decoder block is constructed using the multi-head attention module and the feed-forward neural network structure to perform step-by-step reasoning of the question text using the multiple decoder blocks.

4. The user interaction method according to claim 1, characterized in that: The error constraint of constructing the weight of the key vector includes: Obtaining a calibration sample in the same application scenario as the problem text; constructing a calibration sample set using the calibration samples; Obtaining calibration weights of key vectors based on an unquantized multi-head attention mechanism in the initial decoder block and the calibration sample set; The error constraint is constructed using the calibration weights and the weights of the key vectors in the quantized multi-head attention module.

5. The user interaction method according to claim 4, characterized in that: The key vector cost function is: , in, K is the key vector, is the calibration weight of the bond vector, is the weight of the key vector.

6. The user interaction method according to claim 4, characterized in that: The updating the weight of the value vector of each attention group based on the weight of the updated key vector includes: constructing a value vector cost function of the value vector using the weight of the updated key vector; A weight of the updated value vector is obtained based on the value vector cost function.

7. The user interaction method according to claim 6, characterized in that: The value vector cost function is: , in, K is the key vector, is the calibration weight of the bond vector, is the weight of the key vector, Softmax is the activation function, Q is the query vector, V is the value vector, is the weight of the query vector, is the weight of the value vector, is the calibration weight of the value vector.

8. The user interaction method according to claim 2, characterized in that: The step of inputting the plurality of word vectors into a recurrent processing layer formed by stacking the plurality of decoder blocks to output processed data corresponding to the question text includes: Calculate the attention scores of the multiple word vectors respectively using the attention mechanism of the current decoder block; Obtaining corresponding attention output data based on the attention score; Performing a linear transformation and a nonlinear activation function on the attention output data to perform feature transformation and information extraction on the attention output data, and outputting intermediate processed data; The intermediate processed data is input into the next encoder block until all decoder blocks complete the processing to obtain the processed data.

9. The user interaction method according to claim 2, characterized in that: The decoding of the processed data to obtain the inference data includes: Normalizing the processed data to obtain normalized processed data; Mapping the normalized processed data to a representation space that meets a preset decoding condition to obtain data to be decoded; The data to be decoded is decoded to obtain the inference data.

10. The user interaction method according to claim 9, characterized in that: After obtaining the inference data, the method further includes: Mapping the inference data into vocabulary probabilities; Matching corresponding subwords based on the vocabulary probability, and merging multiple subwords into a complete text based on the subword sequence; The complete text is converted into a display format, and the complete text in the display format is pushed to the user.

11. A user interaction device, characterized in that: include: A receiving module, used for receiving a user's question text; A processing module is used to input the question text into the text reasoning large model, group multiple attention heads in the initial decoder block to obtain multiple attention groups, wherein each attention group shares the same set of key vector weights and value vector weights, and the query vector weights of the attention heads in each attention group are consistent; when the query vector weight in each attention head does not change, construct an error constraint on the key vector weight; construct a corresponding key vector cost function based on the error constraint to update the key vector weight of each attention group; update the value vector weight of each attention group based on the updated key vector weight; update the initial decoder block based on the updated key vector weight and the updated value vector weight to obtain a decoder block; after the question text gradually passes through multiple decoder blocks stacked in the text reasoning large model, output the reasoning data corresponding to the question text, wherein the multi-head attention weights in the decoder block are obtained by gradient-based post-training quantization; A reply module is used to push the inference data to the user.

12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the user interaction method according to any one of claims 1 to 10 when executing the computer program.

13. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the user interaction method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the user interaction method according to any one of claims 1 to 10 are implemented.

Citation Information

Patent Citations

  • Text generation method and device and computing equipment

    CN118364877A

  • Recognition training method and recognition method based on large language model

    CN118520904A