Training Method, Inference Method, Device and Equipment of Large Language Model
By conducting block text training on large language models, it can use the block attention mechanism to perform inference, which solves the problem of low inference efficiency of large language models, and achieves the effect of reducing the number of memory reads and improving inference efficiency.
Patent Information
- Application Number
- CN202510152395.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-12
AI Technical Summary
During the inference process of large language models, each token inference needs to read the cache of key-value pairs in memory once, resulting in inference efficiency.
The large language model is trained by text fragments containing chunked text, so that the model can identify which knowledge is relatively simple. Therefore, in the inference process, the chunked attention mechanism is used for simple content for processing. The large language model can output multiple tokens at once, and in the process of outputting the token, it only needs to read the KV cache from memory once.
Reduces the number of memory reads and improves the inference efficiency of large language models.
Smart Images

Figure CN119621941B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology. Specifically, it relates to a training method, an inference method, a device, and a device for a large language model. Background Art
[0002] Large Language Model (LLM) is abbreviated as large model. It is a type of deep learning model trained based on a large amount of text data. They are particularly powerful in the field of Natural Language Processing (NLP). Large models usually have many computational layers and a large number of parameters (which may reach billions or even tens of billions), and can understand and generate natural language text. After training, the large model can not only generate natural language text through inference, but also deeply understand the text meaning, process various natural language tasks, such as text summarization, classification, question answering, translation, and document summarization, etc., with strong generalization ability.
[0003] Currently, during the inference process of the large language model, each time a token is inferred, it is necessary to read the key-value pair cache (KV cache) in memory once, resulting in low inference efficiency of the large language model. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a training method, an inference method, a device, and a device for a large language model to improve the inference efficiency of the large language model.
[0005] In a first aspect, the embodiments of this application provide a training method for a large language model. The large language model includes multiple layers of Transformer modules. Each layer of the Transformer module includes an attention mechanism module and a feed-forward neural network module. The attention mechanism module includes a standard attention mechanism module and a chunked attention mechanism module. The method includes:
[0006] Obtain training text, where the training text includes an unlabeled first text segment and a second text segment labeled as chunked text, and the first text segment and the second text segment are arranged alternately;
[0007] Input the training text into the large language model, and process the second text segment through the chunked attention mechanism module in the large language model; process the first text segment through the standard attention mechanism module in the large language model;
[0008] Calculate a first loss based on the prediction result of the first text segment output by the large language model; calculate a second loss based on the prediction result of the second text segment output by the large language model; determine the total loss according to the first loss and the second loss;
[0009] Optimize the parameters of the large language model based on the total loss to obtain a trained large language model.
[0010] In the embodiments of the present application, the large language model is trained through text fragments containing chunked text, enabling the large language model to identify which knowledge is relatively simple. Thus, during the inference process of the large language model, for simple content, a chunked attention mechanism can be used for processing. The large language model can output multiple tokens at once, and during the output of these tokens, only one read of the KV cache from memory is required, reducing the number of memory reads and improving the inference efficiency of the large language model.
[0011] In any embodiment, processing the first text fragment through the standard attention mechanism module in the large language model includes:
[0012] Calculate the first attention value between each token in the first text fragment and the tokens before it in the training text through the standard attention mechanism module.
[0013] In the embodiments of the present application, attention calculation is performed on the first text fragment through the standard attention mechanism module. Since the first text fragment is considered complex content, prediction is performed token by token through the standard attention mechanism module, improving the inference accuracy of the large language model.
[0014] In any embodiment, processing the second text fragment through the chunked attention mechanism module in the large language model includes:
[0015] Calculate the second attention value between each token within the second text fragment and the tokens before it within the second text fragment through the chunked attention mechanism module;
[0016] Calculate the third attention value between the second text fragment and the tokens before the second text fragment in the training text through the chunked attention mechanism module.
[0017] In the embodiments of the present application, based on the processing of the second text fragment by the chunked attention mechanism module, internal attention is calculated, enabling the large language model to better understand the relevance and dependence of these tokens in the local context; by calculating external attention, the large language model can capture and utilize broader context information, improving the inference accuracy of the large language model.
[0018] In any embodiment, after obtaining the second attention value and the third attention value, the method further includes:
[0019] Generate an attention matrix based on the first attention value, the second attention value, and the third attention value;
[0020] Input the attention matrix into the feed-forward neural network module in the large language model, and process it through the feed-forward neural network module to obtain the prediction result corresponding to the training text output by the large language model.
[0021] In the embodiment of the present application, the feed-forward neural network module processes the attention matrix to obtain the prediction result corresponding to the training text for subsequent loss calculation.
[0022] In any embodiment, calculating the first loss based on the prediction result of the first text segment output by the large language model includes:
[0023] According to the formula Calculate to obtain the first loss;
[0024] Wherein, is the first loss; is the conditional probability calculated when the internal parameters of the large language model are ; is the i-th token in the training text; is the probability that the -th token appears under the premise of knowing ; is the first text segment.
[0025] In the embodiment of the present application, the loss is calculated through the cross-entropy loss function, which is convenient for optimizing the internal parameters of the large language model.
[0026] In any embodiment, calculating the second loss based on the prediction result of the second text segment output by the large language model includes:
[0027] According to the formula Calculate the block loss between the second text segment and the tokens in the training text before the first text segment;
[0028] According to the formula Calculate the intra-block loss between each token in the second text segment and the token before it in the second text segment;
[0029] Take the sum of the block loss and the intra-block loss as the second loss;
[0030] Wherein, is the block loss; is the conditional probability calculated when the internal parameters of the large language model are ; is the i-th token in the training text; is the probability that the -th token appears under the premise of knowing the first text segment The probability of occurrence; is the intra-block loss; within the second text segment, given that under the premise The probability of occurrence; is the first text segment; is the second text segment.
[0031] In the embodiments of the present application, the intra-block loss focuses on the token correlation within the second text segment and can accurately evaluate the performance of the model when processing local context, while the chunk loss considers the correlation between the second text segment and the broader context, which helps to evaluate the model's ability to capture global context information. The dual constraints help to improve the overall performance of the model when processing complex text tasks.
[0032] In any embodiment, calculating a second loss based on the prediction result of the second text segment output by the large language model includes:
[0033] Calculating to obtain the second loss according to the formula ;
[0034] where is the second loss; is the conditional probability calculated when the internal parameters of the large language model are ; is the th token in the training text; given that under the premise, the probability of occurrence of the first text segment ; is the second text segment; =0, 1, 2, …, N; where N is the length of the first text segment, ; where is the length of the second text segment.
[0035] When calculating the loss of the model in the embodiments of the present application, the loss corresponding to the chunked text and the loss corresponding to the ordinary sequence text are considered, so that a more accurate loss of the model can be obtained, and then the internal parameters of the model can be better optimized based on this loss.
[0036] In a second aspect, an inference method for a large language model provided by the embodiments of the present application includes:
[0037] Receiving a processing request, where the processing request includes data to be processed;
[0038] Input the data to be processed into the large language model. The large language model determines to use the standard attention mechanism module and / or the chunked attention mechanism module in the large language model for processing based on the data to be processed, and obtains the inference result output by the large language model;
[0039] Among them, the large language model is trained by using the method described in the first aspect.
[0040] In a third aspect, an embodiment of the present application provides a training device for a large language model. The large language model includes multiple layers of Transformer modules, and each layer of the Transformer module includes an attention mechanism module and a feed-forward neural network module; the attention mechanism module includes a standard attention mechanism module and a chunked attention mechanism module; the method includes:
[0041] A text acquisition module, configured to acquire training text, where the training text includes an unlabeled first text segment and a second text segment labeled as a chunked text, and the first text segment and the second text segment are arranged alternately;
[0042] A processing module, configured to input the training text into the large language model, and process the second text segment through the chunked attention mechanism module in the large language model; process the first text segment through the standard attention mechanism module in the large language model;
[0043] A loss calculation module, configured to calculate a first loss based on the prediction result of the first text segment output by the large language model; calculate a second loss based on the prediction result of the second text segment output by the large language model; determine the total loss according to the first loss and the second loss;
[0044] A parameter optimization module, configured to optimize the parameters of the large language model based on the total loss to obtain a trained large language model.
[0045] In a fourth aspect, an embodiment of the present application provides an electronic device, including: a processor, a memory, and a bus, where,
[0046] The processor and the memory complete communication with each other through the bus;
[0047] The memory stores program instructions executable by the processor, and the processor can execute the methods of the first aspect or the second aspect by invoking the program instructions.
[0048] In a fifth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium, including:
[0049] The non-transitory computer-readable storage medium stores computer instructions that cause the computer to execute the method of the first aspect or the second aspect.
[0050] In a sixth aspect, an embodiment of the present application provides a computer program product, including computer program instructions that, when read and run by a processor, execute the method of the first aspect or the second aspect.
[0051] Other features and advantages of the present application will be described in the subsequent specification, and part of them will become obvious from the specification, or can be understood by implementing the embodiments of the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Description of the Drawings
[0052] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0053] Figure 1 It is a schematic diagram of a large language model architecture provided by an embodiment of the present application;
[0054] Figure 2 It is a schematic diagram of the process flow of a training method for a large language model provided by an embodiment of the present application;
[0055] Figure 3 It is a masked image of attention provided by an embodiment of the present application;
[0056] Figure 4 It is a schematic diagram of the process flow of an inference method for a large language model provided by an embodiment of the present application;
[0057] Figure 5 It is a schematic diagram of the structure of a training device for a large language model provided by an embodiment of the present application;
[0058] Figure 6 It is a schematic diagram of the physical structure of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0059] The following will describe in detail the embodiments of the technical solutions of the present application in conjunction with the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present application, so they are only examples and cannot be used to limit the protection scope of the present application.
[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above description of the drawings are intended to cover non-exclusive inclusion.
[0061] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity, specific order, or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, the meaning of "a plurality of" is more than two, unless otherwise specifically defined.
[0062] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0063] In the description of the embodiments of this application, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the token " / " herein generally represents an "or" relationship between the associated objects before and after.
[0064] In the description of the embodiments of this application, the term "a plurality of" refers to more than two (including two). Similarly, "a plurality of groups" refers to more than two groups (including two groups), and "a plurality of pieces" refers to more than two pieces (including two pieces).
[0065] In the description of the embodiments of this application, unless otherwise clearly specified and limited, technical terms such as "install", "connect", "link", "fix", etc. should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or integrated; it can also be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of this application can be understood according to specific circumstances.
[0066] To improve the processing efficiency of large language models, the operation of large language models (LLMs) is usually based on the combination of GPUs and CPUs. A GPU is a hardware device specifically designed for processing graphics and image computing. It was originally developed to accelerate graphics rendering and later for executing large batches of calculations. A GPU can provide many parallel processing cores, has strong parallel processing capabilities, is good at executing parallel computing tasks, and can execute most of the inference calculation parts of an LLM. The CPU is responsible for managing the operating system, processing input / output requests, and executing tasks related to the user interface. The CPU and GPU can work together through an interaction process. The CPU performs computational processing based on CPU memory, which is a random access memory (RAM) that directly interacts with the CPU. The GPU performs computational processing based on video memory, which is a graphics card that directly interacts with the GPU and belongs to special-purpose memory.
[0067] Time-Per-Output-Token is one of the metrics for evaluating the inference performance of large language models. TBT represents the latency per output token on average from after the first token is output in a request until the output ends. This metric directly reflects the response speed of the model when generating continuous text and is crucial for scenarios such as chat applications and real-time translation. A lower TBT means that users can receive the model's response faster, thus enhancing the user experience.
[0068] During the inference process of large language models, a token is a basic unit in the text. This unit can be a character, word, phrase, symbol, or any other form of text segment. In different contexts, the definition of a token may vary, but usually they are the basic units for text analysis and processing. Tokens are unsigned integers obtained by effectively splitting or tokenizing the text token string and mapping them one by one.
[0069] It should be noted that large models can process various data types including images and text. Therefore, the data to be processed input to the large model can be an image, text, or a combination of image and text, etc. A processing request is a request to be processed for the large model to perform inference calculations. This processing request can be sent to the large model through a user's operation or through the processing of a certain program. This application does not limit the above content.
[0070] Current popular large models based on the autoregressive Transformer architecture, such as Baichuan, LLaMA, Qwen, GPT, and GLM, are neural network architectures based on Self-Attention. When making inferences, large models under this architecture mainly include two stages of computation, namely the Prefill stage and the Decode stage, which are executed serially.
[0071] The Prefill stage includes the following processes: After the user inputs the data to be processed into the LLM, the LLM performs autoencoding on the data to be processed to generate the attention information of the data to be processed. For example, it generates the KV cache of the tokens in the data to be processed, caches the inference result data in the GPU video memory, and generates the first token. The inference result data is used in the Decode stage. When the large language model contains N computational layers, the Prefill stage needs to sequentially execute the computations of N computational layers to obtain the KV cache data of each token in each computational layer.
[0072] The Decode stage refers to the process where the LLM actually generates the text to be output. In this stage, given the context information (i.e., attention information) of the Prefill stage, the model inference continues, and after passing through the computations of N computational layers again, the tokens of the subsequent text to be output are generated one by one. Each token of the output text needs to be obtained through the computations of N computational layers until the inference ends. During the execution of the Decode stage, the KV cache data cached in the video memory during the Prefill stage needs to be accessed during the inference computation of each computational layer, and the attention information (Key, Value data) of the newly generated Tokens in the Decode stage will be concatenated with the Keys and Values data cached in the Prefill stage for inference.
[0073] During the Attention calculation process of large models, in order to reduce computational operations, the attention information (including Key, Value) of the already generated Tokens is cached, forming an occupied video memory space, which is the KV cache. The KV cache occupies a large amount of GPU video memory, thus trading video memory space for computational time.
[0074] In the decoding stage of large model inference, for a single user, it is mainly the calculation of general matrix-vector multiplication (GEMV). Since the computing power and bandwidth of the hardware are relatively large, it causes waste of computing power resources. In order to effectively utilize the high computing power resources of the tensor cores in the GPU, large language model service providers usually adopt a multi-user mode, converting the GEMV operation of a single user into a GEMM operation of multiple users. However, multi-user large model inference, especially in the context of long contexts, often brings the problem of insufficient HBM storage in the GPU, resulting in a high TBT. Especially in the current scenarios where strong inference of large models and the use of agents are all in long contexts, it becomes very important to improve the GPU utilization rate in the single-user mode.
[0075] Based on this, the embodiments of the present application provide a training method, an inference method, an apparatus, and a device for a large language model. The embodiments of the present application have made improvements in both the architecture and the training method of the large language model. Figure 1 The following is a schematic diagram of the architecture of a large language model provided by the embodiments of the present application, as Figure 1 shown. The large language model adopts a decoder-only architecture of Transformer, that is, the large language model includes L stacked Transformer modules, and the value of L can be set according to actual situations. Each Transformer module includes an attention mechanism module and a feed-forward network module (FFN). The attention mechanism module includes a standard attention mechanism module and a chunked attention mechanism module. It should be noted that Figure 1 in there is no distinction between the standard attention mechanism module and the chunked attention mechanism module. Structurally, the standard attention mechanism module and the chunked attention mechanism module are the same, only with different internal parameters. And Figure 1 in the attention mechanism module shown, the architecture of a standard attention mechanism module and / or a chunked attention mechanism module is schematically drawn. It includes three different weight matrices ( ) for linear transformation to generate corresponding query vector Q, key vector K, and value vector V, and then uses the Softmax function to obtain the attention weights Among them, during the training process, the standard attention mechanism module is used to calculate the attention scores between the current token and the tokens before the current token in the training text. The chunked attention mechanism module is used to calculate the attention scores between each token and the tokens before it within the current text chunk in the training text, as well as to calculate the attention scores between the current text chunk and the tokens before the current text chunk in the training text. Therefore, in the training text of the embodiments of the present application, there are ordinary sequential texts and chunked texts divided into text chunks. In a specific implementation, for a continuous text, a certain continuous text in the text can be used as a text chunk.
[0076] The feed-forward neural network module (FFN) includes two fully connected layers. The first fully connected layer expands the input dimension (for example, from 512 dimensions to 2048 dimensions), followed by an activation function (usually ReLU or GELU), and then the second fully connected layer reduces the dimension from the expanded dimension back to the original dimension (for example, from 2048 dimensions back to 512 dimensions).
[0077] The large language model also includes an input layer and an output layer. The input layer includes an embedding layer for vectorizing the input content; the output layer includes a fully connected layer .
[0078] Figure 2 It is a schematic flowchart of a training method for a large language model provided by the embodiments of the present application. As Figure 2 shown, the large language model is executed by a GPU and a CPU. The method includes:
[0079] Step 201: Obtain the training text, where the training text includes an unannotated first text segment and a second text segment annotated as a chunked text, and the first text segment and the second text segment are arranged alternately;
[0080] Step 202: Input the training text into the large language model, and process the second text segment through the chunked attention mechanism module in the large language model; process the first text segment through the standard attention mechanism module in the large language model;
[0081] Step 203: Calculate the first loss based on the prediction result of the first text segment output by the large language model; calculate the second loss based on the prediction result of the second text segment output by the large language model; determine the total loss according to the first loss and the second loss;
[0082] Step 204: Optimize the parameters of the large language model based on the total loss to obtain a trained large language model.
[0083] The following provides a detailed description of each of the above steps.
[0084] In step 201, the training text can be some text from the Internet or customized. The training text includes a first text segment that is not labeled and a second text segment that is labeled as a chunked text. The first text segment may include one or more characters. In most cases, the second text segment includes multiple characters. The first text segment and the second text segment are arranged alternately to form a continuous training text.
[0085] For ease of understanding, the following is an example: The training text is: On a quiet early morning, "The sun shines through the thin clouds on the earth", a gentle breeze gently blows through the treetops, bringing a faint fragrance of flowers. The birds are singing sweetly on the branches, as if greeting the arrival of a new day. The distant mountains are looming in the morning mist, "giving a hazy beauty. The farmers in the fields start their day's work. They bend down and carefully cultivate the fields full of dew". Life is ordinary but full of warmth and hope. Everyone is working hard silently for their dreams in their hearts, "looking forward to the beauty of the future". In the above training text, the text within "“”" is the second text segment, and the rest are the first text segments. It can be understood that which texts are divided into text chunks can be manually labeled according to actual needs. For example, simple information or parts with low information density can be labeled as text chunks. It can be understood that text with high information density refers to text that contains more abstract, complex concepts or requires reasoning. For example, technical articles, academic papers, complex conversations, etc. These texts usually require the model to understand a large amount of context information, reasoning ability, and detailed language understanding. Information entropy can be used to quantify the complexity of the text. Texts with high information entropy tend to have more variables and potential ways of understanding, usually corresponding to higher information density. Texts with high information density usually contain a more diverse vocabulary and complex syntactic structures. Such texts will have a higher information entropy. Conversely, it is called low information density.
[0086] To enable the large language model to identify which are text chunks and which are ordinary continuous texts, identifiers can be set at the beginning and end of the text chunks. For example: <|start block|> can be used to identify the beginning of the text chunk, and <|endblock|> can be used to identify the end of the text chunk. It should be noted that the delimiters used to identify the start and end of the chunked text can also be other custom symbols, and the embodiments of this application do not make specific limitations in this regard.
[0087] In step 202, the above training text is input into the large language model to be trained. The standard attention mechanism module in the large language model processes the first text segment, and the chunked attention module processes the second text segment. The specific processing process will be described in detail in the subsequent embodiments. The standard attention mechanism module is the main module used to generate tokens one by one, and it can specifically be a self-attention mechanism module, a linear attention mechanism module, etc.
[0088] In step 203, when the prediction result of the first text segment is output by the large language model, it is predicted and output token by token. When the prediction result of the second text segment is output by the large language model, it is output as a whole block, that is, a whole block of prediction results corresponding to the second text segment is output. It should be noted that during the inference process of the large language model, if it is predicted and output token by token, each time a token is output, it is necessary to access the KV cache from the memory. However, if some simple text is chunked and output by blocks, then only one memory access is required for the output of this block of text. Therefore, in the embodiments of the present application, the reason for dividing some text segments in the training text into text blocks is to enable the large language model to learn which content can be output by blocks and which content needs to be output token by token. For example: through a large amount of training, the large language model can master some knowledge. During the inference process, for the large language model, relatively simple content can be output by blocks, and relatively difficult content can be output token by token. In this way, while improving the inference efficiency of the model, the accuracy of the output content is taken into account.
[0089] Therefore, when calculating the loss, the first loss corresponding to the first text segment and the second loss corresponding to the second text segment are accumulated as the total loss of the large language model.
[0090] In step 204, after obtaining the total loss, the internal parameters in the large language model are optimized in reverse using the total loss to obtain an optimized large language model.
[0091] It should be noted that in order to improve the inference performance of the large language model, a large number of training texts can be used to perform multiple trainings according to the method of steps 201 - 204. The condition for the end of the training can be to reach a preset number of iterative trainings, or the change rate of the loss meets the conditions, etc.
[0092] Embodiments of the present application train a large language model through text fragments containing chunked text, enabling the large language model to identify which knowledge is relatively simple. Thus, during the inference process of the large language model, for simple content, a chunked attention mechanism can be used for processing. The large language model can output multiple tokens at once, and during the output of these tokens, only one read of the KV cache from memory is required, reducing the number of memory reads and improving the inference efficiency of the large language model.
[0093] Based on the above embodiments, processing the first text fragment through the standard attention mechanism module in the large language model includes:
[0094] Calculating the first attention value between each token in the first text fragment and the previously output tokens through the standard attention mechanism module.
[0095] In a specific implementation process, the standard attention mechanism module can be a self-attention mechanism. The self-attention mechanism is an attention mechanism that associates different positions of a single sequence to calculate the representation of the same sequence. It allows the model to dynamically adjust the degree of attention to each element when processing sequence data, thereby capturing complex dependencies within the sequence. The core of the self-attention mechanism is that it does not rely on external information but conducts information interaction and integration among the internal elements of the sequence. It should be noted that the standard attention mechanism module can also be a linear attention mechanism module, etc.
[0096] After the first text fragment is input into the large language model, it can first be preprocessed. For example, each token in the first text fragment is converted into a vector representation. This is usually achieved through an embedding layer, which maps the token to a vector in a high-dimensional space.
[0097] For each token vector in the input sequence, linear transformations are performed using three different weight matrices ( ) to generate the corresponding query vector Q, key vector K, and value vector V. These weight matrices are learnable parameters during the training process.
[0098] For each token in the sequence, calculate the dot product using its query vector Q and all key vectors K to obtain attention scores. To stabilize training and improve generalization ability, the attention scores are usually divided by the square root of the key vector dimension (i.e., the scaling factor). Apply the Softmax function to the scaled attention scores to obtain attention weights. Softmax ensures that the sum of all output weights is 1, enabling the model to learn the importance of each token pair. Use the obtained attention weights to perform a weighted sum of the value vectors V to generate the final output vector. This step can effectively integrate the information within the sequence, such that the output representation of each token contains the context information of the entire sequence.
[0099] Apply the above process to all tokens in the sequence to construct the first attention value.
[0100] In the embodiment of the present application, the standard attention mechanism module is used to calculate the attention for the first text segment, which is considered complex content. Therefore, by predicting token by token through the standard attention mechanism module, the accuracy of the large language model inference is improved.
[0101] On the basis of the above embodiment, the second text segment is processed through the block attention mechanism module in the large language model, including:
[0102] Calculate the second attention value between each token in the second text segment and the tokens before this token in the second text segment through the block attention mechanism module;
[0103] Calculate the third attention value between the second text segment and the tokens before the second text segment in the training text through the block attention mechanism module.
[0104] In the specific implementation process, first, identify and extract the second text segment from the entire training text, which includes a series of tokens. Since identifiers are set at both the beginning and the end of the second text segment, the second text segment can be extracted based on the identifiers.
[0105] The block attention mechanism module aims to reduce the number of times of reading the KV cache from memory when processing long sequence data, that is, during the model inference process, the generation of the second text segment only needs to read the KV cache from memory once. And the large language model outputs the second text segment at one time.
[0106] The standard attention mechanism module calculates the attention between pairwise tokens by token, forming a standard lower triangular attention matrix. Different from the standard attention mechanism module, the block attention mechanism module calculates the attention between pairwise tokens within the current token block and the attention between the tokens within the block and the previous tokens in batches, and the formed attention matrix is not a standard lower triangular attention matrix. The formation of the attention matrix is obtained through attention calculation with a mask, Figure 3 is the mask image of the attention mask for the embodiments of the present application, as Figure 3 shown. The value on the blank box is 0, and the box in the shaded area is 1. The training text is arranged alternately with the first text segment and the second text segment. Therefore, looking from top to bottom, the first 4 tokens are the first text segment, and the attention matrix of the traditional attention mechanism transformer decoder-only architecture is lower triangular. When calculating the 5-7th tokens as the second text segment, since the 5-7th tokens are output together and in a BERT-like manner, the attention values of the second text segment form a 3x3 matrix at positions 5-7, and the mask value of the 3x3 attention matrix is all 1, which is a non-lower triangular matrix attention form; when calculating the mutual attention between the second text segment and the first text segment, due to the decoder-only architecture, the overall attention matrix formed by the second text block and the first text block is lower triangular, and it is non-lower triangular at the position of the block text output.
[0107] After obtaining the second text segment, calculate the attention values within the block and the attention values outside the block through the block attention mechanism module. The so-called attention value within the block refers to the attention value between any two tokens within the second text segment. For example: the second text segment includes token1, token2, and token3, and the block attention mechanism module needs to calculate the attention value between token1 and token2, the attention value between token1 and token3, the attention value between token2 and token1, the attention value between token2 and token3, the attention value between token3 and token1, and the attention value between token3 and token2.
[0108] The out-of-block attention value refers to the attention value between the second text segment and the tokens that appear before the second text segment in the training text. For example: The training text includes token1, token2, token3, and token4. Among them, token1 and token2 are ordinary consecutive texts, that is, the first text segment. Token3 and token4 are a text block, that is, the second text segment. When calculating the out-of-block attention value, it is necessary to calculate the attention value between token3 and token4 as a whole and token1, and the attention value between token3 and token4 as a whole and token2.
[0109] It should be noted that the method of calculating each element value in the matrix for the second attention value and the third attention value is similar to the method of calculating each element value in the first attention value, and they are also both obtained by using the query vector Q, the key vector K, and the value vector V. Only the tokens participating in the calculation are different, so it will not be elaborated here.
[0110] The embodiment of this application processes the second text segment based on the block attention mechanism module to calculate the internal attention, enabling the large language model to better understand the relevance and dependence of these tokens in the local context; by calculating the external attention, the large language model can capture and utilize more extensive context information. This improves the accuracy of the large language model's inference.
[0111] Based on the above embodiments, after obtaining the second attention value and the third attention value, the method further includes:
[0112] Generating an attention matrix based on the first attention value, the second attention value, and the third attention value;
[0113] Inputting the attention matrix into the feed-forward neural network module in the large language model, and processing it through the feed-forward neural network module to obtain the prediction result corresponding to the training text output by the large language model.
[0114] In the specific implementation process, after obtaining the first attention value, the second attention value, and the third attention value, a corresponding attention matrix can be generated, where the attention matrix is similar to the Figure 3 aforementioned masked image. The difference is that the element value corresponding to the blank box is 0, and the element value of the box in the shaded area is the corresponding attention value.
[0115] After obtaining the attention matrix, the attention matrix is input into the feed-forward neural network in the large language model. The feed-forward neural network module usually consists of multiple layers of linear transformations and non-linear activation functions. Each layer receives the output of the previous layer and generates a new output after linear transformation and non-linear activation. The feed-forward neural network module provided in the embodiments of the present application may include two layers of linear transformations and two layers of non-linear activation functions. It can be understood that the specific number of layers of linear transformations and non-linear activation functions included in the feed-forward neural network module can be adjusted according to actual situations, and the embodiments of the present application do not make specific limitations in this regard. In the Transformer architecture, the feed-forward neural network model is usually located after the attention mechanism module and is used to further process the vector representation output by the attention mechanism module.
[0116] After the feed-forward neural network module receives the attention matrix, it performs a series of linear transformation and non-linear activation operations on it. These operations are designed to capture higher-order features in the input sequence and generate a more rich representation. After being processed by the feed-forward neural network module, the output will contain a deeper understanding and analysis of the input sequence.
[0117] At the output end of the feed-forward neural network module, an output layer is usually connected. The output layer is responsible for converting the output of the feed-forward neural network into the final prediction result. The output layer may include components such as linear transformation and Softmax function, which are used to convert the output of the neural network into a probability distribution or other forms of prediction results.
[0118] The embodiments of the present application process the content output by the attention mechanism module based on the feed-forward neural network module to obtain a prediction result, which is used for the calculation of the loss and provides a basis for subsequent optimization of the internal parameters of the large language model.
[0119] Based on the above embodiments, the loss function of the embodiments of the present application includes the classic and commonly used LLM training loss function, intra-block loss function, and block loss function.
[0120] Among them, the LLM training loss function is calculated based on the prediction result of the first text segment, that is, the first loss. The intra-block loss function is the loss between each token in the second text segment and the token before it in the second text segment, that is, the intra-block loss. The block loss function is the loss between the second text segment and the tokens between the second text segment in the training text, that is, the block loss. Each loss function is introduced below:
[0121] The calculation formula of the first loss is as follows:
[0122]
[0123] The calculation formula for the loss within a block is as follows:
[0124]
[0125] The calculation formula for the block loss is as follows:
[0126]
[0127] Wherein, is the first loss; is the conditional probability calculated when the internal parameters of the large language model are ; is the i-th token in the training text; is the probability that the -th token appears under the premise of knowing ; is the block loss; is the probability that the first text segment appears under the premise of knowing ; is the loss within a block; is the probability that appears within the second text segment under the premise of knowing ; is the first text segment; is the second text segment.
[0128] The second loss is the sum of the block loss and the loss within a block. After obtaining the first loss and the second loss, the sum of the first loss and the second loss is used as the total loss. That is, the total loss = + + .
[0129] It should be noted that there may be multiple first text segments and multiple second text segments in the training text. The first loss is the sum of the losses of multiple first text segments, and the second loss is the sum of the losses of multiple second text segments.
[0130] After obtaining the total loss, the internal parameters of the large language model are optimized in reverse based on the total loss to obtain an optimized large language model.
[0131] After completing the model training, the large language model can be deployed in a computer or a server. When deploying, it can be compatible with common low-bit quantization precisions. The methods of low-bit quantization include: common LLM quantization techniques such as LLM.int8(), soomthquant, etc.; and the deployment of classical large models, supporting inference frameworks such as Vllm, SGLang, etc.; at the cluster deployment level, it supports technologies such as PD separation.
[0132] It should be noted that the large language question-answering model can be trained through the above training method, that is, by inputting questions to the large language model, the large language model can return corresponding answers. It can also be used to train multi-modal understanding models.
[0133] Based on the above embodiments, calculating the second loss based on the prediction result of the second text segment output by the large language model includes:
[0134] According to the formula Calculate to obtain the second loss.
[0135] Among them, is the second loss; is the conditional probability calculated when the internal parameters of the large language model are ; is the th token in the training text; is the probability that the first text segment appears on the premise of knowing ; is the second text segment; = 0, 1, 2, …, N; where N is the length of the first text segment, ; among them, is the length of the second text segment.
[0136] Therefore, based on the above embodiments, it can be seen that the second loss in the embodiments of the present application only considers the block loss. The total loss of the model = + , is the first loss, and its calculation formula can be seen in the above embodiments, which will not be elaborated here. Compared with the solution in the above embodiments that considers both block loss and intra-block loss, the solution in the embodiments of the present application reduces the calculation amount.
[0137] When calculating the loss of the model in the embodiments of the present application, the loss corresponding to the segmented text and the loss corresponding to the ordinary sequence text are considered, so that the loss of the model can be obtained more accurately, and then the internal parameters of the model can be better optimized based on this loss.
[0138] Based on the large language model training method obtained above, the large language model can execute the following inference method, Figure 4 is a schematic flow diagram of an inference method of a large language model provided by an embodiment of the present application, as Figure 4 shown, the method includes:
[0139] Step 401: Receive a processing request, and the processing request includes data to be processed;
[0140] Step 402: Input the data to be processed into the large language model. The large language model determines to use the standard attention mechanism module and / or the chunked attention mechanism module in the large language model for processing based on the data to be processed, and obtains the inference result output by the large language model.
[0141] In a specific implementation process, the data to be processed can be text, numbers, images, files, or other formats, specifically depending on the requirements of the application scenario. For example, in the scenario of natural language processing, the data to be processed can be a piece of text; in the scenario of image processing, the data to be processed can be one or more images. Of course, the data to be processed can also be multi-modal data. For example, it can include images and text, can also include files and text, and can also include images, files, and text, etc.
[0142] The processing request can be submitted by the user through the user interface, or can be sent by other systems or services through the API interface.
[0143] Input the received data to be processed into the large language model trained by the training method provided in the above embodiment. The large language model performs inference based on the data to be processed. During the inference process, the large language model can identify whether to use the chunked attention mechanism module for inference or the standard attention mechanism module for inference. If the chunked attention mechanism module is used for inference, the large language model can output multiple tokens simultaneously; if the standard attention mechanism module is used for inference, the large language model outputs tokens one by one. Therefore, during the inference process, the large language model may only use the standard attention mechanism module for inference, may only use the chunked attention mechanism module for inference, or may use both the standard attention mechanism module and the chunked attention mechanism module. For the case of using both the standard attention mechanism module and the chunked attention mechanism module, among the tokens output by the large language model, some tokens are output one by one, and some consecutive tokens are output at once. It should be noted that the above descriptions of only using the chunked attention mechanism module and only using the standard attention mechanism module mean for the attention mechanism module. In the actual inference process, in addition to passing through the attention mechanism module (chunked attention mechanism module and / or standard attention mechanism module), other modules such as the feed-forward neural network module also need to be used for processing.
[0144] In the embodiment of the present application, during the inference process of the large language model, it can effectively locate when it is necessary to use the chunked attention mechanism module for processing, generate text in batches, and improve the inference efficiency of the model.
[0145] Figure 5A schematic structural diagram of a training device for a large language model provided by an embodiment of the present application. This device can be a module, a program segment, or code on an electronic device. It should be understood that this device corresponds to the above Figure 2 method embodiment and can execute Figure 2 each step involved in the method embodiment. The specific functions of this device can be referred to the description above. To avoid repetition, the detailed description is appropriately omitted here. The device includes: a text acquisition module 501, a processing module 502, a loss calculation module 503, and a parameter optimization module 504, where:
[0146] The text acquisition module 501 is used to acquire training text, where the training text includes an unlabeled first text segment and a second text segment labeled as a chunk text, and the first text segment and the second text segment are arranged alternately;
[0147] The processing module 502 is used to input the training text into the large language model, and process the second text segment through the chunk attention mechanism module in the large language model; process the first text segment through the standard attention mechanism module in the large language model;
[0148] The loss calculation module 503 is used to calculate a first loss based on the prediction result of the first text segment output by the large language model; calculate a second loss based on the prediction result of the second text segment output by the large language model; determine the total loss according to the first loss and the second loss;
[0149] The parameter optimization module 504 is used to optimize the parameters of the large language model based on the total loss to obtain a trained large language model.
[0150] Based on the above embodiment, the processing module 502 is specifically used for:
[0151] Calculate a first attention value between each token in the first text segment and the token before this token in the training text through the standard attention mechanism module.
[0152] Based on the above embodiment, the processing module 502 is specifically used for:
[0153] Calculate a second attention value between each token within the second text segment and the token before this token within the second text segment through the chunk attention mechanism module;
[0154] Calculate a third attention value between the second text segment and the token before the second text segment in the training text through the chunk attention mechanism module.
[0155] Based on the above embodiments, the processing module is further configured to:
[0156] Generate an attention matrix based on the first attention value, the second attention value, and the third attention value;
[0157] Input the attention matrix into the feed-forward neural network module in the large language model, and process it through the feed-forward neural network module to obtain the prediction result corresponding to the training text output by the large language model.
[0158] Based on the above embodiments, the loss calculation module 503 is specifically configured to:
[0159] Calculate the first loss according to the formula ;
[0160] where, is the first loss; is the conditional probability calculated when the internal parameters of the large language model are ; is the i-th token in the training text; is the probability that the -th token appears under the premise of knowing ; is the first text segment.
[0161] Based on the above embodiments, the loss calculation module 503 is specifically configured to:
[0162] Calculate the block loss between the second text segment and the tokens in the training text before the first text segment according to the formula ;
[0163] Calculate the intra-block loss between each token in the second text segment and the token before it in the second text segment according to the formula ;
[0164] Take the sum of the block loss and the intra-block loss as the second loss;
[0165] where, is the block loss; is the conditional probability calculated when the internal parameters of the large language model are ; is the i-th token in the training text; is the probability that the first text segment appears under the premise of knowing ; is the intra-block loss; within the second text segment, under known premises appearance probability; is the first text segment; is the second text segment.
[0166] Based on the above embodiments, the loss calculation module 503 is specifically configured to:
[0167] According to the formula calculate to obtain the second loss;
[0168] wherein, is the second loss; is the conditional probability calculated when the internal parameters of the large language model are ; is the th token in the training text; is, under the known premises, the appearance probability of the first text segment ; is the second text segment; = 0, 1, 2, …, N; wherein, N is the length of the first text segment, ; wherein, is the length of the second text segment.
[0169] Figure 6 This is a schematic diagram of the entity structure of the electronic device provided by the embodiments of the present application. As Figure 6 shown, the electronic device includes: a processor 601, a memory 602, and a bus 603; wherein,
[0170] The processor 601 and the memory 602 communicate with each other through the bus 603;
[0171] The processor 601 is configured to call program instructions in the memory 602 to execute the methods provided in the foregoing method embodiments. For example, it includes: obtaining a training text, where the training text includes an unannotated first text segment and a second text segment annotated as a chunk text, and the first text segment and the second text segment are arranged alternately; inputting the training text into a large language model, and processing the second text segment through a chunk attention mechanism module in the large language model; processing the first text segment through a standard attention mechanism module in the large language model; calculating a first loss based on the prediction result of the first text segment output by the large language model; calculating a second loss based on the prediction result of the second text segment output by the large language model; determining a total loss according to the first loss and the second loss; and optimizing the parameters of the large language model based on the total loss to obtain a trained large language model.
[0172] The processor 601 may be an integrated circuit chip with signal processing capabilities. The foregoing processor 601 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0173] The memory 602 may include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0174] This embodiment discloses a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above method embodiments. For example, it includes: obtaining a training text, where the training text includes an unannotated first text segment and a second text segment annotated as a chunked text, and the first text segment and the second text segment are arranged alternately; inputting the training text into a large language model, and processing the second text segment through a chunked attention mechanism module in the large language model; processing the first text segment through a standard attention mechanism module in the large language model; calculating a first loss based on the prediction result of the first text segment output by the large language model; calculating a second loss based on the prediction result of the second text segment output by the large language model; determining a total loss according to the first loss and the second loss; optimizing the parameters of the large language model based on the total loss to obtain a trained large language model.
[0175] This embodiment provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions. The computer instructions cause the computer to execute the methods provided in the above method embodiments. For example, it includes: obtaining a training text, where the training text includes an unannotated first text segment and a second text segment annotated as a chunked text, and the first text segment and the second text segment are arranged alternately; inputting the training text into a large language model, and processing the second text segment through a chunked attention mechanism module in the large language model; processing the first text segment through a standard attention mechanism module in the large language model; calculating a first loss based on the prediction result of the first text segment output by the large language model; calculating a second loss based on the prediction result of the second text segment output by the large language model; determining a total loss according to the first loss and the second loss; optimizing the parameters of the large language model based on the total loss to obtain a trained large language model.
[0176] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. Also, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0177] In addition, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0178] Furthermore, in each embodiment of the present application, the various functional modules may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.
[0179] In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0180] The above are only the embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for training a large language model, characterized in that: The large language model includes a multi-layer Transformer module, and each layer of the Transformer module includes an attention mechanism module and a feedforward neural network module; The attention mechanism module includes a standard attention mechanism module and a block attention mechanism module; the method includes: Acquire a plurality of training texts, each of the training texts comprising a first unlabeled text segment and a second text segment labeled as a block text, and the first text segment and the second text segment are alternately arranged to form a continuous training text; The training text is input into a large language model, and the first text segment is processed by a standard attention mechanism module in the large language model; the standard attention mechanism module is used to calculate a first attention value between each token in the first text segment and a previously output token; the second text segment is processed by a block attention mechanism module in the large language model, and after obtaining the second text segment, the block attention mechanism module is used to calculate the attention value within the block, and the attention value outside the block; the block attention mechanism module is used to calculate a second attention value between each token in the second text segment and a token before the token in the second text segment, and to calculate a third attention value between the second text segment and a token before the second text segment in the training text; the standard attention mechanism module calculates the attention of two tokens per token to form a standard lower triangular attention matrix; the block attention mechanism module batch calculates the attention of two tokens in the current token block and the attention of tokens in the block and previous tokens, and the attention matrix formed is not a standard lower triangular attention matrix; Calculating a first loss based on the prediction result of the first text segment output by the large language model; calculating a second loss based on the prediction result of the second text segment output by the large language model; determining a total loss based on the first loss and the second loss; Parameters of the large language model are optimized based on the total loss to obtain a trained large language model.
2. The method according to claim 1, characterized in that The processing of the first text segment by a standard attention mechanism module in the large language model includes: A first attention value between each token in the first text segment and a token preceding the token in the training text is calculated by the standard attention mechanism module.
3. The method according to claim 2, characterized in that The processing of the second text segment by the block attention mechanism module in the large language model includes: Calculating a second attention value between each token in the second text segment and a token before the token in the second text segment through the block attention mechanism module; A third attention value between the second text segment and a token preceding the second text segment in the training text is calculated by the block attention mechanism module.
4. The method according to claim 3, characterized in that After obtaining the second attention value and the third attention value, the method further includes: generating an attention matrix based on the first attention value, the second attention value, and the third attention value; The attention matrix is input into a feedforward neural network module in the large language model, and is processed by the feedforward neural network module to obtain a prediction result corresponding to the training text output by the large language model.
5. The method according to claim 1, characterized in that The calculating a first loss based on the prediction result of the first text segment output by the large language model includes: According to the formula Calculate and obtain the first loss; in, is the first loss; The internal parameters of the large language model are Calculate the conditional probability obtained when is the first Tokens; For the known Under the premise, The probability of a token appearing; is the first text segment; =0,1,2,…,N, where N is the length of the first text segment.
6. The method according to any one of claims 1 to 5, characterized in that: The calculating a second loss based on the prediction result of the second text segment output by the large language model includes: According to the formula Calculating a block loss between the second text segment and a token in the training text that precedes the first text segment; According to the formula Calculating the intra-block loss between each token in the second text segment and the token before the token in the second text segment; taking the sum of the block loss and the intra-block loss as the second loss; in, is the block loss; The internal parameters of the large language model are Calculate the conditional probability obtained when is the first Tokens; For the known Under the premise that the first text fragment Probability of occurrence; is the loss within the block; For the second text fragment, in the known Prerequisites Probability of occurrence; is the first text segment; is the second text segment; =0,1,2,…,N; wherein N is the length of the first text segment, ;in, is the length of the second text segment.
7. The method according to any one of claims 1 to 5, characterized in that: The calculating a second loss based on the prediction result of the second text segment output by the large language model includes: According to the formula Calculate and obtain the second loss; in, for the second loss; The internal parameters of the large language model are Calculate the conditional probability obtained when is the first Tokens; For the known Under the premise that the first text fragment Probability of occurrence; is the second text segment; =0,1,2,…,N; wherein N is the length of the first text segment, ;in, is the length of the second text segment.
8. A method for reasoning about a large language model, characterized in that: include: receiving a processing request, wherein the processing request includes data to be processed; Inputting the data to be processed into a large language model, the large language model determining, based on the data to be processed, to adopt a standard attention mechanism module and / or a block attention mechanism module in the large language model for processing, and obtaining an inference result output by the large language model; Wherein, the large language model is obtained by training using the method described in any one of claims 1-7.
9. A large language model training device, characterized in that: The large language model includes a multi-layer Transformer module, and each layer of the Transformer module includes an attention mechanism module and a feedforward neural network module; The attention mechanism module includes a standard attention mechanism module and a block attention mechanism module; the device includes: A text acquisition module, used for acquiring a plurality of training texts, each of the training texts comprising a first unlabeled text segment and a second text segment labeled as a block text, and the first text segment and the second text segment are alternately arranged to form a continuous training text; A processing module, used for inputting the training text into a large language model, and processing the first text segment through a standard attention mechanism module in the large language model; the standard attention mechanism module is used to calculate a first attention value between each token in the first text segment and a previously output token; the second text segment is processed through a block attention mechanism module in the large language model, and after obtaining the second text segment, the block attention mechanism module is used to calculate the attention value within the block, and the attention value outside the block; the block attention mechanism module is used to calculate a second attention value between each token in the second text segment and a token before the token in the second text segment, and to calculate a third attention value between the second text segment and a token before the second text segment in the training text; the standard attention mechanism module calculates the attention of two tokens per token to form a standard lower triangular attention matrix; the block attention mechanism module batch calculates the attention of two tokens in the current token block and the attention of tokens in the block and previous tokens, and the attention matrix formed is not a standard lower triangular attention matrix; a loss calculation module, configured to calculate a first loss based on a prediction result of the first text segment output by the large language model; calculate a second loss based on a prediction result of the second text segment output by the large language model; and determine a total loss based on the first loss and the second loss; A parameter optimization module is used to optimize the parameters of the large language model based on the total loss to obtain a trained large language model.
10. An electronic device, characterized in that: include: processor, memory and bus, wherein, The processor and the memory communicate with each other via the bus; The memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the method according to any one of claims 1 to 7.
11. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 7.
12. A computer program product, characterized in that The method comprises computer program instructions, and when the computer program instructions are read and executed by a processor, the method according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Information record data blocking method based on pre-training language model
CN117763093A