Model pre-training method and device, equipment and storage medium

By performing specific training strategies on the word embedding layer and attention module of the causal language model, the training method of the model is optimized, the resource consumption problem of existing models is solved, and more efficient training and deployment is achieved.

CN120087501APending Publication Date: 2025-06-03GUANGZHOU CIVIL AVIATION INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510473426.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Existing language models require massive computing resources during training and deployment, limiting researchers' ability to continue their efforts on their success.

Method used

A model pre-training method is proposed, which can fix the encoder parameters of the causal language model and train the word embedding layer. Then set the adapter weight in the attention module, and jointly train the word embedding layer, model head and adapter weights to optimize the model training method.

Benefits of technology

It significantly reduces the memory cost and computing resources requirements, improves the performance and training efficiency of the model, and realizes the possibility of pre-training large-scale data sets in resource-limited environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087501A_ABST
    Figure CN120087501A_ABST
Patent Text Reader

Abstract

The invention discloses a model pre-training method and device, equipment and a storage medium. The method comprises the following steps: for a to-be-trained causal language model, fixing encoder parameters of the causal language model, and training a word embedding layer of the causal language model; in response to completion of training of the word embedding layer, setting an adapter weight in an attention module of the causal language model, and performing joint training on the word embedding layer, a head of the causal language model and the adapter weight to obtain a trained causal language model; wherein the training of the word embedding layer and the joint training respectively adopt a Chinese data set. According to the method provided by the invention, the video memory cost required by model training is remarkably reduced, so that the method can be implemented on a single civil video card, and the convenience and efficiency of training are ensured while the performance of the Chinese language model is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and more particularly, to a method, apparatus, device, and storage medium for model pre-training. Background Art

[0002] With the emergence of large language models, the field of natural language processing has undergone earth-shaking changes. Currently, large language models are trained with a fairly large-scale, comprehensive, and pure training data, and have shown excellent capabilities in understanding and generating text. Compared with traditional models, large language models have shown significant advantages in text generation, context understanding and interaction, and multi-turn question answering, especially in logical reasoning ability.

[0003] However, despite the satisfactory performance of large language models, their training and deployment have limitations. Whether it is to train a large language model from scratch based on extensive corpora or to perform pre-training on the basis of open-source models, it requires a large amount of computing resources, thus limiting the ability of researchers to build on their success.

[0004] Therefore, the present application provides a method, apparatus, device, and storage medium for model pre-training to solve one of the above technical problems. Summary of the Invention

[0005] The purpose of the present application is to provide a method, apparatus, device, and storage medium for model pre-training, which can solve at least one of the above-mentioned technical problems. The specific solutions are as follows: According to the specific embodiments of the present application, in a first aspect, the present application provides a method for model pre-training, including: For a causal language model to be trained, fix the encoder parameters of the causal language model, and train the word embedding layer of the causal language model; in response to completing the training of the word embedding layer, set adapter weights in the attention module of the causal language model, and jointly train the word embedding layer, the head of the causal language model, and the adapter weights to obtain a trained causal language model; wherein, the training of the word embedding layer and the joint training respectively use Chinese datasets.

[0006] In one implementation, the attention module is a multi-head attention module, and uses variant modules of multiple different attention mechanisms; wherein, the multiple different attention mechanisms include dilated attention mechanism and enhanced local attention mechanism, and the attention results of the multi-head attention module use fusion calculation of distributed attention.

[0007] In one implementation, sparse calculation is adopted between different heads in the multi-head attention module.

[0008] In one implementation, the query vectors between different heads in the multi-head attention module are calculated independently, and the key vectors and value vectors are shared between different heads.

[0009] In one implementation, the decoder of the causal language model includes a self-attention layer, a cross-attention layer, and a feed-forward network layer, and the self-attention layer, the cross-attention layer, and the feed-forward network layer integrate representations based on a local attention structure.

[0010] In one implementation, the attention weight parameter of the integrated representation is used to replace the weight parameter of the feed-forward network layer.

[0011] According to the specific implementation manners of the present application, in a second aspect, the present application provides a model pre-training device, including: An input processing unit, configured to input a Chinese dataset as training data into a causal language model to be trained; a training control unit, configured to fix the encoder parameters of the causal language model and train the word embedding layer of the causal language model; and in response to completing the training of the word embedding layer, set adapter weights in the attention module of the causal language model, and jointly train the word embedding layer, the head of the causal language model, and the adapter weights to obtain a trained causal language model; wherein, the Chinese dataset is used for training the word embedding layer and the joint training respectively.

[0012] In one implementation, the attention module is a multi-head attention module, and variant modules with multiple different attention mechanisms are adopted; wherein, the multiple different attention mechanisms include a dilated attention mechanism and an enhanced local attention mechanism, the attention result of the multi-head attention module adopts a fusion calculation of distributed attention, a sparsification calculation is adopted between different heads in the multi-head attention module, the query vectors between different heads in the multi-head attention module are calculated independently, and the key vectors and value vectors are shared between different heads; the decoder of the causal language model includes a self-attention layer, a cross-attention layer, and a feed-forward network layer, the self-attention layer, the cross-attention layer, and the feed-forward network layer integrate representations based on a local attention structure, and the attention weight parameter of the integrated representation is used to replace the weight parameter of the feed-forward network layer.

[0013] According to the specific implementation manners of the present application, in a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the method according to any one of the first aspect.

[0014] According to a specific embodiment of the present application, in a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program / instructions are stored, characterized in that when the computer program / instructions are executed by a processor, the method described in any one of the first aspects is implemented.

[0015] Compared with the prior art, the above solution of the embodiments of the present application has at least the following beneficial effects: The present application provides a model pre-training method, aiming to optimize and upgrade the model training method by transforming the attention module and the decoder module of the existing transformer model architecture. For optimizing the attention module, the present application integrates various variants of attention mechanisms, including dilated attention and enhanced local attention. This strategy adopts distributed attention calculation to reduce the computational complexity of attention operations. The present application realizes computational differences between different heads through sparse calculations in different parts, and constructs a set of attention results by expanding the query while keeping the keys and values shared. For the decoder module, the present application integrates the self-attention layer, the cross-attention layer and the feed-forward network in the decoder into a unified local attention structure. In this optimization, the present application uses the combined attention weights to replace the weight parameters in the feed-forward neural network, so as to focus the model's attention more on local information. On this basis, the performance and training efficiency of the model are significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 Shows a flowchart of a model pre-training method; Figure 2 Shows a training principle diagram of a causal language model; Figure 3 Shows a model processing flowchart of an attention mechanism; Figure 4 Shows a schematic diagram of a causal language model architecture and its working process; Figure 5 Shows a training experimental data display diagram of the model; Figure 6 Shows a unit block diagram of a model pre-training device according to an embodiment of the present application; Figure 7 Is a block diagram of an electronic device 700 for model pre-training shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following will further describe this application in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of them. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0018] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the", and "said" used in the embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plural" generally includes at least two.

[0019] It should be understood that the term "and / or" used herein is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0020] It should be understood that although terms such as first, second, and third may be used in the embodiments of this application for description, these descriptions should not be limited to these terms. These terms are only used to distinguish the descriptions. For example, without departing from the scope of the embodiments of this application, the first can also be called the second, and similarly, the second can also be called the first.

[0021] Depending on the context, the words "if", "when" as used herein can be interpreted as "when...", "when...", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined", "in response to determining", "when detected (stated condition or event)", or "in response to detecting (stated condition or event)".

[0022] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a commodity or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or device including the said element.

[0023] It should be particularly noted that symbols and / or numbers in the specification, if not marked in the accompanying drawing explanations, are not drawing reference numerals.

[0024] In this application, the proprietary nouns involved are directly translated from Chinese to English in the following way. For any unclear Chinese descriptions in the following embodiments, reference can be directly made here: Causal Language Modeling (CLM); Language Model heads (LM heads); Transformer model; Encoder; Decoder; Positional Encoding; Embeddings; Attention Module; Multi-Head Attention; Masked Multi-Head Attention; Dilated Attention; Cross-Attention layer; Self-Attention layer; Local-Attention; Enhanced Local Attention (EL Attention); Feed Forward Network (FFN); Layer Normalization; Residual Connection; Attention heads; Query; Key; Value.

[0025] The optional embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0026] The embodiments provided in this application are embodiments of a model pre-training method.

[0027] The following will be combined with Figure 1 The embodiments of this application will be described in detail.

[0028] Figure 1 The flowchart of a model pre-training method is shown. As Figure 1 shown, it includes the following steps.

[0029] Step S101: For the causal language model to be trained, fix the encoder parameters of the causal language model and train the word embedding layer of the causal language model.

[0030] Step S102: In response to the completion of the training of the word embedding layer, set the adapter weights in the attention module of the causal language model, and jointly train the word embedding layer, the head of the causal language model, and the adapter weights to obtain the trained causal language model.

[0031] Among them, the Chinese dataset is used for training the word embedding layer and joint training respectively.

[0032] The present invention proposes a model pre-training method. By first fixing the encoder parameters of the causal language model and separately training the word embedding layer, new added words can be effectively learned without damaging the performance of the original model. Then, adapter weights are introduced in the attention module and the word embedding layer, the model head and the adapter weights are jointly trained, which not only promotes the effective integration of new added words, but also improves the adaptability of the model to specific tasks or domains. The argument is that this two-stage training strategy significantly reduces the demand for video memory cost and computing resources, while ensuring the efficiency and accuracy during the pre-training process, thus demonstrating the possibility of implementing pre-training on large-scale datasets in resource-limited environments (such as consumer-grade graphics cards).

[0033] Figure 2 Shows a training principle diagram of a causal language model.

[0034] In the embodiments of the present application, as Figure 2 shown, a causal language model is adopted, and this model can be, for example, a Chinese LLaMA model. The input sequence The prefix words or characters in the input sequence, and these words or characters are known context information. MASK represents a special marker indicating the unknown part or the part to be predicted. During the training process, x i is the next word or character predicted by the model, and is predicted according to the given prefix sequence and the true label. The causal language model receives these inputs and predicts the next word or character based on the known prefix sequence and the MASK marker. In the first stage of pre-training, the present application fixes all the parameters of the transformer model encoder in the LLaMA large model and only trains the word embedding to adapt to the extended vocabulary while minimizing the interference to the original LLaMA model. In the second stage, the present application adds adapter weights in the attention module and jointly trains the word embedding, the language model head (language model head) and the newly added adapter weights.

[0035] In the present application, the causal language model consists of two main components: an encoder and a decoder. The following is a detailed introduction to the model structure, including the specific composition of the encoder and the decoder and the data processing process of each layer.

[0036] In the embodiments of the present application, the encoder is mainly responsible for performing preliminary processing and feature extraction on the input data, providing necessary context information for the subsequent decoding process. For example, it may include an input layer, positional encoding, multi-head attention, a feed-forward neural network, layer normalization, and residual connections. Exemplarily, the input layer receives the original Chinese text data. After preprocessing steps such as word segmentation and word embedding, the text is converted into a vector form that the model can understand. Positional encoding is used to distinguish words at different positions. Positional encoding is usually added to the word embedding as the input to the encoder. Multi-head attention is one of the core components of the Transformer model architecture, which allows the model to simultaneously focus on different parts of the input sequence in different representation subspaces. In the encoder, the multi-head attention mechanism generates an output vector containing context information through operations such as linear transformation of the input vector, dot-product attention calculation, and concatenation. The feed-forward neural network is located after the multi-head attention and is used to further process the output of the multi-head attention. It usually consists of two linear transformations and an activation function, which can perform non-linear transformation on the input and extract deeper features. In addition, to ensure the stability and training efficiency of the model, layer normalization and residual connections are performed on each layer in the encoder. Layer normalization reduces the internal covariate shift by normalizing the input, while the residual connection allows the gradient to be directly propagated to earlier layers, helping to alleviate the vanishing gradient problem.

[0037] In the embodiments of the present application, the decoder part is responsible for predicting the next word according to the output of the encoder and the previous generation result. For example, it may include an input layer, masked multi-head attention, cross-attention layer, a feed-forward neural network, layer normalization, and residual connections. Exemplarily, the input layer of the decoder also receives the Chinese text data that has undergone preprocessing steps such as word segmentation and word embedding. However, different from the encoder, the decoder also needs to receive the previously generated words as input in an autoregressive manner. During the decoding process, to ensure that the model can only predict the next word based on the previous generation result, the multi-head attention mechanism needs to be masked. The masking operation usually sets the attention weights of future positions to 0 or negative infinity, thus preventing the model from seeing future information. The cross-attention layer mechanism allows the decoder to focus on the output of the encoder, thereby obtaining global context information. In the cross-attention layer, the query vector (exemplarily from the previous layer of the decoder) and the key vector and value vector (exemplarily from the output of the encoder) perform dot-product attention calculation to generate an output vector containing global context information. The feed-forward neural network in the decoder is similar to that in the encoder and is used to further process the output of the cross-attention layer. In addition, layer normalization and residual connections are also performed on each layer in the decoder to ensure the stability and training efficiency of the model.

[0038] In some embodiments, the attention module is a multi - head attention module, which adopts variant modules of multiple different attention mechanisms. Among them, the multiple different attention mechanisms include dilated attention mechanism and enhanced local attention mechanism, and the attention results of the multi - head attention module adopt the fusion calculation of distributed attention.

[0039] In the embodiments of the present application, the model is allowed to more flexibly capture the relationships between different parts of the input sequence, especially long - range dependencies. By combining different types of attention mechanisms, the computational complexity can be effectively reduced, and the expressiveness of the model can be maintained or even improved.

[0040] In some embodiments, sparse computing is adopted between different heads in the multi - head attention module.

[0041] In some embodiments, the query vectors between different heads in the multi - head attention module are calculated independently, and the key vectors and value vectors are shared between different heads.

[0042] In some embodiments, the decoder of the causal language model includes a self - attention layer, a cross - attention layer, and a feed - forward network layer. The self - attention layer, the cross - attention layer, and the feed - forward network layer integrate representations based on a local attention structure.

[0043] In some embodiments, the attention weight parameters for integrating representations are used to replace the weight parameters of the feed - forward network layer.

[0044] Figure 3 A model processing flow chart of an attention mechanism is shown.

[0045] Exemplarily, as Figure 3 shown, the input data X undergoes a segmentation process to generate multiple subsequences X 1 ,X 2 ,… These subsequences are mapped from a high - dimensional space to a low - dimensional space through a projection operation to obtain corresponding query vectors, key vectors, and value vectors. For example, the query vector Q 1 corresponding to the subsequence X 1 , the key vector K 1 , and the value vector V 1 . Next, the sparse data structure summarizes these vectors to form the final items. Then, the attention distribution of each item is calculated through the attention mechanism, and the data length is calculated based on these distributions. Finally, the final output result is obtained through weighted summation. This process demonstrates how to optimize the data processing process through the attention mechanism and the sparse data structure to improve the training efficiency and performance of the model. Among them, the optimized attention mechanism avoids mapping and to the subspace, but maps Directly project it into a low-rank space while keeping its dimension unchanged. This improvement makes the present application no longer rely on and , thus effectively releasing the caches of the encoder and decoder attention sub-layers. In addition, the present application only needs to save the output of the encoder and share it among all decoders. At the same time, only requires a simple linear transformation, and the hidden state is directly used as and .

[0046] In some embodiments, by sparsifying different parts of the query vector, key vector, and value vector , differential calculations between each attention head are achieved. In specific operations, for the j-th attention head, the present application introduces an offset during the selection of . Then, the present application evenly divides the input into multiple segments , each segment having a length of W, and N represents the total length of the input sequence. On this basis, the present application sparsifies these segments along the sequence dimension by selecting rows at an interval of r. As shown in the following formula:

[0047] where Q represents the query matrix, which is a matrix transformed from the input sequence and used to query relevant information. K represents the key matrix, which is another matrix transformed from the input sequence and used to match the query. V represents the value matrix, which is a matrix containing actual information. When the query matches the key, the value will be extracted as the final result. i represents the row index, indicating which row of data is being processed currently. w represents the length of each segment, that is, the number of consecutive elements selected each time. r represents the sparsity, that is, how many elements to skip to select a row. s_j represents the offset, which is used to achieve differential calculations between different attention heads.

[0048] In the embodiments of the present application, by sparsely selecting different parts of Q, K, and V, differential calculations can be performed between different attention heads, which can reduce the computational complexity while maintaining the effectiveness of the model.

[0049] In the present application, for the sparsified segments , the present application constructs the attention result by only expanding , while allowing all attention heads (attention heads) to share and . The calculation process is divided into three steps: First, It is extended to the multi - head attention format and multiplied by the hidden state, i.e., the encoder output or the output of the previous layer, to calculate the attention probabilities. Then, the intermediate attention result is obtained by multiplying these attention probabilities by the same hidden state again. Finally, the present application generates the final output through a linear transformation.

[0050] In the present application, in order to further optimize the calculation process, two feed - forward networks are introduced, which are specifically used to calculate the attention result.

[0051] For example, through denoted as. Among them, , , and , is the normalization factor of the softmax equation, is the model dimension, Att(Q, K, V) represents the attention function, which is used to accept three inputs, namely the query matrix Q, the key matrix K, and the value matrix V, and returns an attention result matrix. h represents the number of attention heads, that is, how many independent attention calculations are there.

[0052] In the present application, the feed - forward neural network receives an input matrix and maps it to a new space. The feed - forward neural network i represents the i - th feed - forward neural network. \(W^k\) represents the weight matrix of the input layer. \(W^o\) represents the weight matrix of the output layer. d represents the normalization factor of the softmax equation, usually set to sqrt(d_model) to ensure calculation stability and convergence. \(d_m\) represents the model dimension, that is, d_model, which usually refers to the hidden layer size in the transformer model architecture.

[0053] In the above - mentioned embodiment, the first formula defines a composite operation. First, the attention function is applied to obtain the preliminary attention result, and then these results are passed to multiple feed - forward neural networks. Each feed - forward neural network will first perform a linear transformation on the attention result and then apply the second linear transformation. Finally, the outputs of all feed - forward neural networks are added together to obtain the final attention result.

[0054] On this basis, the attention results are integrated into the output O and input into the attention in parallel, as shown in the following formula: ; where \(O_i\) represents the i - th attention result, j represents the index of the attention head, r represents the sparsity, that is, how many elements are skipped to select one row, \(\widetilde{O}\) represents the combined attention result, and O represents the final output matrix, which is composed of the combined attention results.

[0055] In the embodiments of the present application, the formula indicates that for each attention result O_i, it will only be considered when the index j of the attention head modulo r is equal to 0. This means that only a part of the attention results will participate in the construction of the final output. This helps to reduce the computational complexity because not all attention results need to be retained. For example, assume that the present application has 5 attention heads and the sparsity r is 2, then only the attention heads with odd indices (i.e., the 1st, 3rd, and 5th attention heads) will be considered. The results of other attention heads will be ignored. The advantage of doing this is that it can reduce the computational burden while maintaining the effectiveness of the model.

[0056] Figure 4 A schematic diagram showing the causal language model architecture and its workflow is presented.

[0057] Exemplarily, as Figure 4 shown, the input first undergoes token embedding, which converts each word into a vector representation and combines positional encoding to retain the position information in the sequence. Next, these embedded vectors are processed by the multi-head attention mechanism, which allows the model to focus on different parts of the input sequence. Subsequently, the output is optimized through residual connections and layer normalization, and then passes through a per-bit perceptron feed-forward network to further enhance the feature representation. This process is repeated N times to form multiple encoder layers. Finally, the model calculates the probability distribution of each word in the output vocabulary through linear regression and softmax regression, thereby predicting the next word or character. Among them, L is set to 12 and N is set to 6. The present application realizes the goal of reducing the video memory occupancy and accelerating the model inference speed while maintaining the model accuracy by increasing the number of encoder layers and reducing the number of decoder layers.

[0058] In the embodiments of the present application, in the decoder of the Transformer model, each layer sequentially includes three sub-layers: a self-attention layer, a cross-attention layer, and a feed-forward network. The self-attention layer receives the output X of the previous sub-layer as input and produces a tensor of the same size as X. The attention distribution is calculated , and X is averaged through . The present application represents the self-attention layer as , where , t is the length of the output sentence, and d is the dimension of the hidden self-attention layer. Therefore, it can also be represented by the following formula: ; where, , Self(X) represents the output of the self-attention layer mechanism. It takes the input X and produces a tensor of the same size as X. X represents the output of the previous layer, which is a t×d matrix, where t is the length of the output sentence and d is the dimension of the hidden layer. A_x represents the attention distribution, which is a t×t matrix indicating the degree of importance of each position. Y_x represents the final output of the self-attention layer mechanism, also a t×d matrix. SoftMax represents a normalization function used to normalize the attention distribution into a probability distribution. W_q represents the query weight matrix for converting the input X into a query vector. W_k represents the key weight matrix for converting the input X into a key vector. W_v represents the value weight matrix for converting the input X into a value vector. W_o represents the output weight matrix for converting the attention distribution back to the original dimension.

[0059] Exemplarily, the input X is converted into query, key, and value vectors, the similarity scores between the query and the key are calculated, and then they are transformed into a probability distribution through the SoftMax function. The attention distribution is used to multiply the value vector to obtain a new output. Further, a linear transformation is performed on the new output to restore it to the original dimension. During this process, calculating the attention distribution reflects the correlation between each position and other positions. In this way, the self-attention layer mechanism can understand the input sequence from a global perspective to facilitate the model to perform natural language processing tasks.

[0060] In the embodiments of the present application, the cross-attention layer is similar to the self-attention layer, except that it additionally introduces the encoder output H as an input. The cross-attention layer in the present application is represented as , where , s is the length of the input sentence, which can be represented by the following formula: ; where , Cross(X, H) represents the output of the cross-attention layer mechanism. It takes the inputs X and H and produces a tensor of the same size as X. X represents the input of the decoder, which is a t×d matrix, where t is the length of the output sentence and d is the dimension of the hidden layer. H represents the output of the encoder, which is an s×d matrix, where s is the length of the input sentence. A_h represents the cross-attention layer distribution, which is a t×s matrix indicating the correlation between each position of the decoder and each position of the encoder. Y_h represents the final output of the cross-attention layer mechanism, also a t×d matrix. SoftMax represents the normalization function for normalizing the attention distribution into a probability distribution. W_q represents the query weight matrix for converting the input X into a query vector. W_k represents the key weight matrix for converting the input H into a key vector. W_v represents the value weight matrix for converting the input H into a value vector.

[0061] In the above embodiments, the input X of the decoder is converted into a query vector, and the output H of the encoder is converted into key and value vectors. The similarity score between the query and the key is calculated and then transformed into a probability distribution through the SoftMax function. Further, the attention distribution is multiplied by the value vector to obtain a new output. Furthermore, a linear transformation is performed on the new output to restore it to the original dimension. Among them, the key of the cross-attention layer mechanism lies in introducing the output of the encoder, enabling the decoder to focus on the important parts of the input sequence. This helps to improve the model's understanding ability and prediction accuracy.

[0062] In the embodiments of the present application, the feed-forward network receives the non-linearly transformed result X as input. The present application represents the feed-forward neural network as , where , FeedForwardNet(X) represents the output of the feed-forward neural network, which accepts the input X and undergoes two linear transformations and the ReLU activation function. Among them, ReLU represents the non-linear activation function used to increase the non-linear expression ability of the model, W_1 represents the weight matrix of the first layer, b_1 represents the bias term of the first layer, W_2 represents the weight matrix of the second layer, and b_2 represents the bias term of the second layer.

[0063] In the above embodiments, the formula describes the basic steps of the feed-forward neural network. Among them, the input X is multiplied by the weight matrix of the first layer, added with the bias term, and then the ReLU activation function is applied. The result of ReLU is multiplied by the weight matrix of the second layer and then added with the bias term to obtain the final output. The feed-forward neural network is usually used in deep learning models and can learn complex non-linear relationships. In this embodiment, the feed-forward neural network receives the non-linearly transformed result as input and can usually be implemented through activation functions such as ReLU or GELU. Based on this, the feed-forward neural network can simulate arbitrary functions to a certain extent, thereby improving the performance of the model.

[0064] In the causal language model task, as a variant of the attention mechanism, local attention enables the model to focus on only a part of the input sequence rather than the entire sequence when generating each output. This method helps to process long texts, improves the computational efficiency of the model, and in some cases helps to capture local context information. It is different from the global self-attention layer, and the output at each position only depends on some adjacent positions in the input sequence. This helps to alleviate the computational and storage pressure when processing long sequences. Among them, implementing local attention includes the following key steps: determining the local region at each time step, calculating the attention weights, and using these weights to perform weighted summation on the value vectors.

[0065] Determine the local region: Define a dynamically determined local region for each time step. This local region represents the input context that the model needs to focus on when generating the word at the current position.

[0066] Calculate attention weights: For each time step, calculate the similarity between the query vector and the key vectors, and then apply the softmax function to obtain the attention weights. Different from global attention, only the key vectors within the local region are considered here.

[0067] Weighted sum: Use the calculated attention weights to perform a weighted sum on the value vectors to obtain the context information at the current time step. Therefore, local attention can be expressed by the following formula:

[0068] where t is the current time step, and radius is the local region radius, which refers to the local region that the model focuses on at the current time step. Based on the above information, the combined local attention is expressed as: ; where , is the local context matrix. Local attention(q, K, V) represents the output of the local attention mechanism, which takes the query q, keys K, and values V as inputs. t represents the current time step, indicating the position processed by the model. radius represents the local region radius, specifying the range that the model focuses on. softmax represents the normalization function, which is used to normalize the attention distribution into a probability distribution. Attention X represents the output of the combined local attention mechanism. C represents the local context matrix, which is a two-dimensional array that records the context information near the current position.

[0069] In the above embodiments, the feed-forward neural network represents the output of the feed-forward network, which receives the result of local attention and performs a non-linear transformation. Based on this, the above formula describes the basic steps of the local attention mechanism. In this process, according to the current time and the local region radius, the attention distribution is calculated, the value vectors are multiplied by the attention distribution to obtain a new output, and the new output is passed to the feed-forward network for non-linear transformation. Among them, the local attention mechanism limits the attention range of the model, avoiding the computational overhead brought by global search. In this way, the model can focus more on local information, improving efficiency and effectiveness.

[0070] In the embodiment of the present application, by replacing K and V in the self-attention layer and the cross-attention layer with the local context matrix C, the present application successfully realizes the localization of attention. In this way, the network only pays attention to the information of the local area when generating each word, rather than the entire input sequence. In addition, the present application can adjust the size of the local context according to the requirements of the specific task and data set to control the scope of the model's attention. This can be achieved by changing the size of the local context matrix C.

[0071] Figure 5 The figure shows the training experimental data display diagram of the model.

[0072] like Figure 5 As shown, through ablation experiments and comparative analysis, this application determines the impact of each module on the experimental results and its optimization effect. In order to evaluate the effects of different pre-training methods, this application divides pre-training into four types: unoptimized pre-training, pre-training that optimizes only the attention module, pre-training that optimizes only the decoder module, and pre-training that optimizes both the attention and decoder modules.

[0073] This application focuses on solving the problems of excessive resource usage and slow training speed in the Chinese LLaMA pre-training model. By carefully optimizing the attention module and the decoder module, this application successfully achieved the goal of significantly reducing video memory costs and increasing training speed. The optimization method of this application can be directly applied to consumer-grade graphics cards. For example, a single RTX3090 (24G) graphics card can be used to complete pre-training tasks while ensuring the accuracy of the model. This optimization result greatly reduces the burden on video memory and brings practical benefits to the training and application of the Chinese LLaMA pre-training model. This not only means more efficient utilization of hardware resources, but also provides possibilities for a wider range of research and application scenarios. The method of this application enables the high accuracy of the model to be maintained even when resources are limited, providing more flexibility for the development of the field of Chinese natural language processing.

[0074] In addition, this memory saving achievement lays the foundation for end-to-end Chinese natural language processing deployment on multiple devices, providing solid support for the widespread application and promotion of Chinese natural language processing technology.

[0075] The present application also provides an apparatus embodiment that is consistent with the above-mentioned embodiment, which is used to implement the method steps described in the above-mentioned embodiment. The explanation based on the same name meaning is the same as that of the above-mentioned embodiment, and has the same technical effect as that of the above-mentioned embodiment, and will not be repeated here.

[0076] like Figure 6 As shown, the present application provides a model pre-training device 600, comprising: An input processing unit 601 for inputting a Chinese dataset as training data into a causal language model to be trained; A training control unit 602 for fixing the encoder parameters of the causal language model and training the word embedding layer of the causal language model; and in response to completing the training of the word embedding layer, setting adapter weights in the attention module of the causal language model, and jointly training the word embedding layer, the head of the causal language model, and the adapter weights to obtain a trained causal language model; wherein the Chinese dataset is used for training the word embedding layer and the joint training respectively.

[0077] In one implementation, the attention module is a multi-head attention module, adopting variant modules of multiple different attention mechanisms; wherein the multiple different attention mechanisms include a dilated attention mechanism and an enhanced local attention mechanism, the attention results of the multi-head attention module are calculated by fusing distributed attention, sparse calculation is adopted between different heads in the multi-head attention module, the query vectors between different heads in the multi-head attention module are calculated independently, and the key vectors and value vectors are shared between different heads; The decoder of the causal language model includes a self-attention layer, a cross-attention layer, and a feed-forward network layer. The self-attention layer, the cross-attention layer, and the feed-forward network layer integrate representations based on a local attention structure, and the attention weight parameters of the integrated representation are used to replace the weight parameters of the feed-forward network layer.

[0078] Regarding the device in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0079] Figure 7 It is a block diagram of an electronic device 700 for model pre-training shown according to an exemplary embodiment.

[0080] As Figure 7As shown, an embodiment of the present application provides an electronic device 700. Among them, the electronic device 700 includes a memory 701, a processor 702, and an input / output (I / O) interface 703. Among them, the memory 701 is used to store instructions. The processor 702 is used to call the instructions stored in the memory 701 to execute the model pre-training method of the embodiments of the present application. Among them, the processor 702 is respectively connected to the memory 701 and the I / O interface 703, and can be connected, for example, through a bus system and / or other forms of connection mechanisms (not shown). The memory 701 can be used to store programs and data, including the programs of the model pre-training method involved in the embodiments of the present application. The processor 702 executes various functional applications and data processing of the electronic device 700 by running the programs stored in the memory 701.

[0081] In the embodiments of the present application, the processor 702 can be implemented in at least one hardware form of a digital signal processor (DSP), a field programmable gate array (FPGA), and a programmable logic array (PLA). The processor 702 can be a central processing unit (CPU) or a combination of one or more of other forms of processing units with data processing capabilities and / or instruction execution capabilities.

[0082] The memory 701 in the embodiments of the present application may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0083] In the embodiments of the present application, the I / O interface 703 can be used to receive input instructions (such as numerical or character information, and key signal inputs related to user settings and function controls of the electronic device 700, etc.), and can also output various information to the outside (such as images or sounds, etc.). In the embodiments of the present application, the I / O interface 703 can include one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a mouse, a joystick, a trackball, a microphone, a speaker, and a touch panel, etc.

[0084] In some embodiments, the present application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, perform any of the methods described above.

[0085] In some embodiments, the present application provides a computer program product including a computer program that, when executed by a processor, performs any of the methods described above.

[0086] Although the operations are described in a specific order in the drawings, it should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order, or requiring all of the operations shown to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.

[0087] The methods, apparatuses, devices, and storage media of the present application can be completed using standard programming techniques, and various method steps are implemented using rule-based logic or other logics. It should also be noted that the terms "apparatus" and "module" as used herein and in the claims are intended to include implementations using one or more lines of software code and / or hardware implementations and / or devices for receiving inputs.

[0088] Any of the steps, operations, or programs described herein can be performed or implemented using one or more hardware or software modules alone or in combination with other devices. In one embodiment, the software module is implemented using a computer program product including a computer-readable medium containing computer program code that can be executed by a computer processor to perform any or all of the described steps, operations, or programs.

[0089] For purposes of illustration and description, the foregoing description of the embodiments of the present application has been given. The foregoing description is not exhaustive and is not intended to limit the present application to the exact form disclosed, and various modifications and variations may be possible according to the above teachings, or various modifications and variations may be obtained from the practice of the present application. These embodiments are selected and described to illustrate the principles of the present application and its practical applications, so that those skilled in the art can utilize the present application in various embodiments and various modifications suitable for the conceived specific purposes.

[0090] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0091] It can be further understood that, unless otherwise specified, "connection" includes direct connection between two parties without other components therebetween, and also includes indirect connection between two parties with other elements therebetween.

[0092] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of the present application, it should not be construed as requiring the operations to be performed in the specific order or serial order shown, or requiring all the operations shown to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.

[0093] Those skilled in the art will readily conceive of other implementations of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the field of the present application not disclosed herein. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0094] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

[0095] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A model pre-training method, characterized in that: The method comprises: For a causal language model to be trained, fixing encoder parameters of the causal language model, and training a word embedding layer of the causal language model; In response to completing the training of the word embedding layer, setting adapter weights in the attention module of the causal language model, and jointly training the word embedding layer, the head of the causal language model, and the adapter weights to obtain a trained causal language model; Wherein, the training of the word embedding layer and the joint training respectively adopt Chinese data sets.

2. The method according to claim 1, characterized in that The attention module is a multi-head attention module, which adopts a variety of variant modules with different attention mechanisms; Among them, the multiple different attention mechanisms include an expanded attention mechanism and an enhanced local attention mechanism, and the attention results of the multi-head attention module adopt a fusion calculation of distributed attention.

3. The method according to claim 2, characterized in that Sparse calculation is adopted between different heads in the multi-head attention module.

4. The method according to claim 2, characterized in that: The query vectors between different heads in the multi-head attention module are calculated independently, and the key vectors and value vectors are shared between the different heads.

5. The method according to claim 1, characterized in that The decoder of the causal language model includes a self-attention layer, a cross-attention layer and a feedforward network layer, and the self-attention layer, the cross-attention layer and the feedforward network layer integrate representations based on a local attention structure.

6. The method according to claim 5, characterized in that The attention weight parameters of the integrated representation are used to replace the weight parameters of the feedforward network layer.

7. A model pre-training device, characterized in that: include: An input processing unit, used for inputting the Chinese data set as training data into the causal language model to be trained; A training control unit, used to fix encoder parameters of the causal language model and train a word embedding layer of the causal language model; In response to completing the training of the word embedding layer, setting adapter weights in the attention module of the causal language model, and jointly training the word embedding layer, the head of the causal language model, and the adapter weights to obtain a trained causal language model; wherein the training of the word embedding layer and the joint training respectively use the Chinese data set.

8. The device according to claim 7, characterized in that The attention module is a multi-head attention module, which adopts a variant module of a plurality of different attention mechanisms; wherein the plurality of different attention mechanisms include an expanded attention mechanism and an enhanced local attention mechanism, the attention result of the multi-head attention module adopts a fusion calculation of distributed attention, sparse calculation is adopted between different heads in the multi-head attention module, the query vectors between different heads in the multi-head attention module are calculated independently, and the key vectors and value vectors are shared between the different heads; The decoder of the causal language model includes a self-attention layer, a cross-attention layer and a feedforward network layer. The self-attention layer, the cross-attention layer and the feedforward network layer integrate representations based on a local attention structure, and the attention weight parameters of the integrated representations are used to replace the weight parameters of the feedforward network layer.

9. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Multi-task language processing model training method and device and related equipment

    CN120598061A