Large language model performance optimization method and apparatus, storage medium, and electronic device

WO2026199694A1PCT designated stage Publication Date: 2026-10-01X STAR TECHNOLOGY PTE LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/095829
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-25
Filing Date
2025-05-19
Publication Date
2026-10-01

Smart Images

  • Figure CN2025095829_01102026_PF_FP_ABST
    Figure CN2025095829_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of natural language processing, and provides a large language model performance optimization method and apparatus, a storage medium, and an electronic device. The electronic device processes generated text by means of a large language model, to obtain hidden state information transmitted to a main decoding head, wherein the large language model further comprises a plurality of slave decoding heads parallel to the main decoding head, and a preset arrangement order among the plurality of slave decoding heads represents an arrangement order among decoding results; then, sequence information of each slave decoding head is combined with the hidden state information, to obtain information to be decoded of each slave decoding head; and finally, an optimal predicted text subsequent to the generated text is obtained on the basis of a candidate word set decoded from each piece of information to be decoded. In this way, by combining the sequence information of each slave decoding head with the hidden state information, the plurality of slave decoding heads decode, in parallel, respective information to be decoded, thereby significantly improving the inference speed of the large language model while maintaining the inference accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, devices, storage media and electronic devices for optimizing the performance of large language models

[0001] Cross-reference to related applications

[0002] This disclosure claims priority to Chinese Patent Application No. 2025103542429, filed on March 25, 2025, entitled “Method, Apparatus, Storage Medium and Electronic Device for Optimizing the Performance of Large Language Models”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the field of natural language processing, and more specifically, to a method, apparatus, storage medium, and electronic device for optimizing the performance of large language models. Background Technology

[0004] Since the advent of ChatGPT, generative large language models have become a hot topic of research and discussion. Most current generative large language models use an autoregressive approach to generate text. Specifically, during the generation process, the model predicts the next most likely word based on the generated text sequence, adds it to the sequence, and then repeats this process until the entire text is completed.

[0005] For example, the process will be illustrated more intuitively below with reference to Figure 1. Taking the input text sequence "I go to the basketball court" as an example, the text sequence "I go to the basketball court" is input into the large language model, referred to as the target model. The target model generates a token for the next word "ball" through a forward inference (for ease of description, "word" will be used as the descriptive unit below, but "token" is used in actual processing). Then, the newly generated "ball" is appended to the previous input, and the resulting text sequence "I go to the basketball court" is input into the target model. The target model generates the next word "court" through a second forward inference, and the newly generated word "court" is appended to the previous input. The resulting text sequence "I go to the basketball court" is then used as the next input. This process continues until the complete output text of the target model is obtained.

[0006] Therefore, this method relies on the probability distribution of the model, that is, to calculate the conditional probability of each possible word based on the context and select the word with the highest probability as the output; however, it is limited to only one word at a time, resulting in low inference efficiency.

[0007] In light of this, several technical solutions for accelerating inference were proposed during the research process. The most representative of these is the multi-word prediction method, which generates multiple candidate words at once and selects the best predicted text from them. However, practical experience revealed that the prediction accuracy of multiple candidate words is often insufficient, leading to significant fluctuations in the acceptance rate. Furthermore, when the acceptance rate decreases, the acceleration effect on inference weakens significantly, potentially even negating the efficiency gains from multi-word prediction. While other related techniques proposed during the research can improve the accuracy of multi-word prediction, they come at the cost of inference speed. Therefore, how to significantly improve the inference speed of large language models while maintaining inference accuracy has become a crucial problem that urgently needs to be solved.

[0008] Application content

[0009] To overcome at least one deficiency in the prior art, this disclosure provides a method, apparatus, storage medium, and electronic device for optimizing the performance of large language models, specifically including:

[0010] This disclosure provides a performance optimization method for a large language model, wherein the large language model includes a main decoding head, and the method includes:

[0011] The generated text is processed by a large language model to obtain hidden state information transmitted to the main decoding head. The large language model also includes multiple slave decoding heads running in parallel with the main decoding head. The preset arrangement order among the multiple slave decoding heads represents the arrangement order among the decoding results.

[0012] The sequence information of each decoder head is combined with the hidden state information to obtain the information to be decoded for each decoder head.

[0013] Based on the candidate word set decoded from each piece of information to be decoded, the best predicted text for the subsequent generated text is obtained.

[0014] Secondly, this disclosure provides a large language model performance optimization device, wherein the large language model includes a main decoding head, and the device includes:

[0015] The text feature module is configured to process the generated text through a large language model to obtain hidden state information transmitted to the main decoding head. The large language model also includes multiple slave decoding heads running in parallel with the main decoding head. The preset arrangement order among the multiple slave decoding heads represents the arrangement order among the decoding results.

[0016] The feature optimization module is configured to combine the sequence information of each decoder head with the hidden state information to obtain the information to be decoded for each decoder head.

[0017] The feature decoding module is configured to obtain the best predicted text following the generated text based on the candidate word set decoded from each piece of information to be decoded.

[0018] Thirdly, this disclosure provides a storage medium storing a computer program that, when executed by a processor, implements the large language model performance optimization method.

[0019] Fourthly, this disclosure provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program, and the computer program, when executed by the processor, implements the large language model performance optimization method.

[0020] Compared with the prior art, this disclosure has the following beneficial effects:

[0021] The large language model performance optimization method, apparatus, storage medium, and electronic device disclosed herein involve processing the generated text using a large language model to obtain hidden state information transmitted to the main decoder. The large language model also includes multiple slave decoders running in parallel with the main decoder, with a preset arrangement order representing the order of decoding results. Then, the sequence information of each slave decoder is combined with the hidden state information to obtain the information to be decoded for each slave decoder. Finally, based on the candidate word set decoded from each piece of information to be decoded, the best predicted text for the subsequent generated text is obtained. Thus, by combining the sequence information of each slave decoder with the hidden state information to obtain the information to be decoded for each slave decoder, and by having multiple slave decoders decode their respective information to be decoded in parallel, the inference speed of the large language model can be significantly improved while maintaining inference accuracy. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of this disclosure and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 is a schematic diagram of the principle of the autoregressive text generation method provided in the embodiments of this disclosure;

[0024] Figure 2 is a schematic diagram illustrating the interaction principle between the draft model and the target model provided in the embodiments of this disclosure;

[0025] Figure 3 is one of the schematic diagrams illustrating the principle of reusing hidden state information according to an embodiment of this disclosure;

[0026] Figure 4 is a second schematic diagram illustrating the principle of reusing hidden state information according to an embodiment of this disclosure;

[0027] Figure 5 is a flowchart illustrating the large language model performance optimization method provided in this embodiment of the present disclosure;

[0028] Figure 6 is a schematic diagram illustrating the principle of combining hidden state information and sequence information provided in the embodiments of this disclosure;

[0029] Figure 7 is one of the detailed schematic diagrams of the large language model performance optimization method provided in the embodiments of this disclosure;

[0030] Figure 8 is a second detailed schematic diagram of the large language model performance optimization method provided in the embodiments of this disclosure;

[0031] Figure 9 is a tree-shaped combination diagram provided in the embodiments of this disclosure;

[0032] Figure 10 shows the attention mask matrix provided in an embodiment of this disclosure;

[0033] Figure 11 is a schematic diagram of the shielding block provided in an embodiment of this disclosure;

[0034] Figure 12 is a schematic diagram illustrating the principle of shielding block shielding calculation provided in the embodiments of this disclosure;

[0035] Figure 13 is a comparison chart of acceptance rates based on Qwen2-7B provided in the embodiments of this disclosure;

[0036] Figure 14 is a comparison chart of acceptance rates based on Qwen2-72B provided in the embodiments of this disclosure;

[0037] Figure 15 is a schematic diagram of the large language model performance optimization device provided in an embodiment of this disclosure;

[0038] Figure 16 is a schematic diagram of the structure of the electronic device provided in the embodiment of this disclosure. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure (hereinafter referred to as "the embodiments") clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0040] Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely to illustrate selected embodiments of the disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0041] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0042] In the description of this application, it should be noted that the terms "first," "second," "third," etc., are used only for distinguishing descriptions and should not be construed as indicating or implying relative importance. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0043] Based on the above statement, as introduced in the background technology, although multi-word prediction and other accelerated reasoning techniques can improve efficiency, insufficient prediction accuracy leads to fluctuations in acceptance rate and weakened acceleration effect. Moreover, techniques that improve accuracy often sacrifice speed. Therefore, how to significantly improve the reasoning speed of large language models while maintaining accuracy has become a key issue.

[0044] To make the objectives, technical solutions, and advantages of the embodiments disclosed herein easier to understand, the research process for proposing this solution will be described in detail below.

[0045] During the research, it was considered that the large number of model parameters in the target model was the main reason for the slow word-by-word inference speed. Therefore, if multiple candidate words could be inferred with a smaller amount of computation for the target model to filter, the inference speed could be improved. Based on this idea, a draft model, called the target model, was proposed to use a model with relatively fewer model parameters to quickly generate multiple candidate words. Then, the target model would verify these candidate words and accept some of them based on the verification results.

[0046] For example, continue to take the text sequence "wo qu lan" (I go to basket[ball]) shown in FIG. 1 as an example. As shown in FIG. 2, the text sequence "wo qu lan" is input into a draft model, and the draft model generates 4 candidate words "qiu chang da jia" (basketball court fight) through 4 forward inferences, the 4 candidate words are appended to the text sequence "wo qu lan", to obtain the text sequence "wo qu lan qiu chang da jia" (I go to fight on the basketball court). Then, the text sequence "wo qu lan qiu chang da jia" is input into a target model at one time, and the target model generates 5 candidate words "qiu chang da lan le" (basketball court basketballing done) through one forward inference. Finally, the 4 candidate words "qiu chang da jia" generated by the draft model and the 5 candidate words "qiu chang da lan le" generated by the target model are compared in sequence, the comparison is stopped when different candidate words are compared, and the same part "qiu chang da lan" (basketball court basketball) is intercepted as the optimal subsequent prediction text for the text sequence "wo qu lan".

[0047] It should be noted that the "lan" (basketball) in the optimal prediction text is a correction to the draft model generated by the target model. Since the generation of each word depends on the preceding words, when there is an error in the words generated earlier, the words generated later are also wrong. In this example, since "lan" is generated by inputting "wo qu lan" spliced with "qiu chang da" into the target model, and none of these words are problematic, "lan" can be used as the optimal prediction text. However, "le" (done) is generated by inputting "wo qu lan" spliced with "qiu chang da jia" into the target model, and "jia" (fight) has been determined by the target model as an erroneous word, so "le" cannot be used as the optimal prediction text.

[0048] In this way, the target model only needs to generate 3 words through one forward inference. It should be understood that the draft model usually uses a model of the same series as the target model with a smaller number of parameters, and its inference speed is much faster than that of the target model. When the acceptance rate of candidate words generated by the draft model is relatively high, the inference speed of the target model can be significantly improved.

[0049] However, it has been found in practice that, although the above-mentioned draft model does not have a large number of parameters, it still occupies a certain amount of computing power resources (including GPU resources and memory resources) during inference, thereby limiting its popularization and application, especially in the case of a shortage of computing power.

[0050] Therefore, this research also proposes a technique for reusing the hidden state information of the target model to generate multiple candidate words. Before explaining this technique, it is necessary to understand that in a large language model based on the Transformer architecture, the Transformer module is the basic unit that constitutes the large language model. As shown in Figure 3, multiple slave decoders are added in parallel with the main decoder head (Lm_Head) after the last hidden layer of the Transformer module. Each slave decoder head can generate a candidate word set containing at least one candidate word; then, the candidate words in these candidate word sets are verified through a forward inference process. The main decoder head (Lm_Head) converts the hidden state information of the model into actual words in the vocabulary. The slave decoders function similarly to the main decoder head, decoding the same hidden state information in parallel to generate more candidate words, thereby improving inference efficiency and the quality of candidate text.

[0051] In practice, it was found that while the newly added decoders have a pre-defined order to represent the sequence of decoding results, they are not aware of their own sequence position during decoding. This leads to inconsistent acceptance rates for generated candidate words, and the acceleration effect becomes less noticeable when the acceptance rate decreases. Specifically, with the addition of multiple decoders, each generating at least one candidate word, the number of permutations and combinations increases. Consequently, more candidate words need to be verified in forward inference. However, each additional candidate sequence increases the computational load of the attention mechanism. If the acceptance rate is unstable, it can easily lead to a waste of computational resources, even though only a small number of candidate words can be accepted, thus slowing down the inference speed.

[0052] Therefore, inspired by the Long Short-Term Memory (LSTM) network model, which introduces time-step dependencies to carry sequence information, a solution as shown in Figure 4 is proposed. As shown in Figure 4, the predicted word output by the main decoder and the hidden state information are input into the first slave decoder. The candidate word set output by the first slave decoder, along with the hidden state information, is then input into the second slave decoder to obtain the candidate word set output by the second decoder. This process is repeated to generate multiple candidate word sets. Finally, the candidate words in these sets are validated through a forward inference process.

[0053] However, in practice, it was found that because decoding needs to be performed sequentially from the decoding head, and each decoder also needs to perform a certain amount of computation, the inference speed will be significantly reduced after the number of decoding heads increases to a certain extent.

[0054] In conclusion, how to significantly improve the reasoning speed of large language models while maintaining reasoning accuracy has become a key issue that urgently needs to be addressed.

[0055] Based on the discovery of the aforementioned technical problems, the following technical solutions are proposed through creative effort to solve or improve these problems. It should be noted that the deficiencies in the solutions of the prior art are the result of practical experience and careful research. Therefore, the discovery process of the aforementioned problems and the solutions proposed in the embodiments of this disclosure below should be considered as contributions to this disclosure during the inventive process, and should not be construed as technical content known to those skilled in the art.

[0056] In view of the above-mentioned technical problems, this embodiment provides a method for optimizing the performance of large language models. As shown in Figure 5, the method includes:

[0057] S1 processes the generated text using a large language model to obtain the hidden state information that is transmitted to the main decoding head.

[0058] The large language model also includes multiple slave decoders running in parallel with the main decoder. The preset arrangement order among the multiple slave decoders represents the arrangement order among the decoding results.

[0059] S2, combine the sequence information of each decoder head with the hidden state information to obtain the information to be decoded for each decoder head.

[0060] S3, based on the candidate word set decoded from each piece of information to be decoded, obtain the best predicted text for the subsequent generated text.

[0061] In this way, by combining the sequence information of each decoder head with the hidden state information, the information to be decoded for each decoder head is obtained, and the information to be decoded for each decoder head is decoded in parallel by multiple decoders, thereby significantly improving the inference speed of large language models while maintaining inference accuracy.

[0062] It should be understood that, for the large language model performance optimization method provided in this embodiment, the electronic device implementing the method can be, but is not limited to, a mobile terminal, tablet computer, laptop computer, desktop computer, server, etc. The server can be a single server or a group of servers. The server group can be centralized or distributed (e.g., the server can be a distributed system). In some embodiments, the server can be local or remote relative to the user terminal. In some embodiments, the server can be implemented on a cloud platform; by way of example only, the cloud platform can include private cloud, public cloud, hybrid cloud, community cloud, distributed cloud, inter-cloud, multi-cloud, etc., or any combination thereof. In some embodiments, the server can be implemented on an electronic device having one or more components.

[0063] To make the solution provided in this embodiment clearer, the following describes in detail each step of the method shown in Figure 5, using a server as the electronic device implementing the method. However, it should be understood that the operations in the flowchart may not be implemented in sequence, and steps without logical contextual relationships may be reversed in order or implemented simultaneously. Furthermore, those skilled in the art, guided by this disclosure, may add one or more other operations to the flowchart, or remove one or more operations from the flowchart. Referring again to Figure 5, the method includes:

[0064] S1 processes the generated text using a large language model to obtain the hidden state information that is transmitted to the main decoding head.

[0065] The large language model also includes multiple slave decoders running in parallel with the main decoder. The preset arrangement order of the slave decoders represents the arrangement order of the decoding results. In other words, in this embodiment of the large language model, in addition to the main decoder, there are multiple parallel slave decoders; these slave decoders work in parallel with the main decoder and participate together in the model's decoding process.

[0066] Furthermore, different decoders generate different candidate words, and the order in which these candidate words are combined into the text is determined by the order in which the decoders are arranged. For example, if decoder A is listed before decoder B, then when candidate word A generated by decoder A is combined with candidate word B generated by decoder B, candidate word A needs to be listed before candidate word B.

[0067] In this embodiment, the aforementioned generated text refers to the portion of text content that has already been generated and output by the large language model. Hidden state information refers to the intermediate representation extracted during the large language model's processing of the generated text, containing contextual information accumulated by the model during text processing and its current state. During the large language model's processing, the server transmits this hidden state information to the main decoder, enabling the inference of subsequent predicted words from the generated text.

[0068] As shown in Figure 6, the multiple slave decoders are led out from the position before the master decoder in the forward propagation path, so that the multiple slave decoders can simultaneously decode the hidden state information transmitted to the master decoder.

[0069] In this embodiment, the branch containing the main decoder is referred to as the backbone model of the large language model. As an optional implementation, the slave decoders can adopt a similar structure to the main decoder. For example, each slave decoder can be structured as a single or multiple ResBlock layers connected to a Final Linear layer (a linear layer that transforms the latent space to the dictionary dimension). The structure of each ResBlock layer is x + SiLU(Linear(x)), where x is the latent space input, SiLU is the activation function, and Linear is the linear layer. Therefore, compared to training an additional draft model, only simple fine-tuning of each slave decoder is required. During training, only the final hidden state information of the backbone model is used, without needing the logit processing of the FinalLinear layer in the main predictor.

[0070] Based on the above embodiments' description of the generated text and hidden state information in step S1, referring to Figure 5, step S2 in Figure 5 will be explained below:

[0071] S2, combine the sequence information of each decoder head with the hidden state information to obtain the information to be decoded for each decoder head.

[0072] Referring to Figure 6, the server can map the position of each decoder head to the sequence information of each decoder head using a sine or cosine function. This mapping method borrows from the positional encoding technique in the Transformer model. Its core principle is to use sine or cosine functions to generate position vectors, which are added to the input word embedding vectors. In this way, the model can understand the positional relationships in the input sequence without explicit positional information.

[0073] In this embodiment, a similar method is used, but applied to the arrangement of the decoder heads. Specifically, the server converts the arrangement of each decoder head into corresponding sequence information using a sine or cosine function. This allows the model to understand the relative positions of the decoder heads during the decoding process, thereby performing the decoding operation more efficiently.

[0074] Then, the server can combine the sequence information of each decoder with the hidden state information to obtain the information to be decoded for each decoder. This can be understood as the hidden state information being combined with the position-encoded sequence information to form the information to be decoded. This information not only contains contextual information but also allows each decoder to perceive its position in the sequence. The sequence information helps the decoders better understand the positional relationships of candidate words in the sequence, enhancing the model's sequential dependence during decoding. Thus, the model can more accurately select appropriate candidate words, thereby improving the candidate word acceptance rate and decoding accuracy.

[0075] Based on the above explanation of the sequence information and the information to be decoded from the decoding head in step S2, the following explanation will continue with step S3 in Figure 5:

[0076] S3, based on the candidate word set decoded from each piece of information to be decoded, obtain the best predicted text for the subsequent generated text.

[0077] Each candidate word set includes at least one candidate word. This can be understood as the server decoding each candidate word set from each piece of information to be decoded, each containing at least one candidate word. These candidate word sets are obtained by multiple decoding heads respectively; then, by evaluating and filtering these candidate word sets, at least a portion of the candidate word sets are selected. From each of these selected candidate word sets, one candidate word is chosen and combined in the order they are arranged to form the best predicted text for the subsequent generated text.

[0078] For example, Figure 6 will be used to illustrate this process more intuitively. Figure 6 shows two decoders. For the generated text "Take one basketball, we", the first decoder decodes candidate words including "going" and "very", while the second decoder decodes candidate words including "to", "off", and "happy". The final best predicted text is "going to". It can be seen that "going" is selected from "going" and "very", and "to" is selected from "to", "off", and "happy".

[0079] Regarding step S3 in Figure 5, several implementation methods were proposed during the research process. As shown in Figure 7, step S3 may include:

[0080] S3-1A combines the candidate words in each candidate word set in a permutation order to obtain multiple candidate texts.

[0081] For example, continuing to refer to Figure 6, "going, very" and "to, off, happy" can be combined in the following order to produce six candidate texts: "going to", "going off", "going happy", "very to", "very off", and "very happy". The goal is to find the best predicted text from these six candidate texts. However, it should be understood that when each candidate text contains a large number of candidate words, the best predicted text may be a partial segment of one of the candidate texts, rather than the entire candidate text.

[0082] S3-2A determines the best predicted text from multiple candidate texts.

[0083] The research revealed that, because the backbone model often contains a large number of parameters, performing a single verification of candidate texts by the backbone model does not actually improve inference speed; in fact, it slows it down. This is because, as shown in Figure 6, six candidate texts are obtained, meaning the backbone model needs to perform six forward inferences, while the final verified candidate text only contains two candidate words. However, using the most conventional inference method, six forward inferences would yield six accurate predicted words. Further research revealed that some candidate texts exhibit significant semantic and grammatical issues. For example, in the six candidate texts in Figure 6, "very to" clearly does not conform to English grammar. Initially, the research considered accumulating the probabilities of candidate words in each candidate text to obtain a score for each text. However, after practice, it was found that this method did not consider the semantic information of the context, resulting in the retention of the candidate words with the highest probabilities in each set. Often, the candidate text composed of the highest-probability candidate words is not the optimal text. Therefore, in this embodiment, a model with relatively small parameters can be used to quickly score the semantics and syntax of the above 6 candidate texts, and the one with the highest score can be selected for verification by the backbone model.

[0084] Based on the above concept, the server can concatenate at least some text fragments from the generated text with each candidate text to obtain multiple texts to be scored; the pre-trained text scoring model scores each of the multiple texts to be scored to obtain a score for each text to be scored; and the best predicted text is determined based on the score of each text to be scored.

[0085] It should be understood that this embodiment concatenates at least a portion of the generated text with each candidate text primarily to provide more contextual information, thereby making the scoring more accurate. Furthermore, the aforementioned text scoring model can be fine-tuned using the BERT model. As a pre-trained deep learning model, BERT is capable of capturing bidirectional contextual information from text, making it highly suitable for handling semantic and grammatical tasks. Therefore, it is only necessary to utilize a pre-trained BERT model and train it for a specific task to adapt it to specific scoring requirements.

[0086] In this way, each candidate text is scored by a text scoring model with smaller parameters, thereby quickly obtaining the score of each candidate text. The candidate text with the highest score is then selected for further validation by the backbone model, thereby determining the best predicted text for the subsequent generated text.

[0087] As another optional implementation of step S3 above, a tree-structured attention mask matrix can be used to perform forward inference on candidate words from multiple candidate word sets to complete the verification. As shown in Figure 8, step S3 may also include:

[0088] S3-1B, each candidate word set is multiplied sequentially according to the arrangement order to obtain multiple subsequences to be verified.

[0089] In this case, the number of candidate words in each subsequence to be verified is equal to the number of candidate words in the previous subsequence to be verified.

[0090] For example, continuing with the predicted word "are" output by the main decoder in Figure 6 and the two candidate word sets decoded from the decoder, and using Figure 9 as a reference, this process will be explained more intuitively. As shown in the tree-like combination diagram in Figure 9, "are" can be combined with "going" and "very". Since there is only one "are", the subsequence to be verified obtained by doubling "going" and "very" includes one candidate word set "going" and "very". Based on the subsequence to be verified obtained by doubling "going" and "very", "going" and "very" can be combined with "to", "off", and "happy" respectively. Therefore, the subsequence to be verified obtained by doubling "to", "off", and "happy" includes two candidate word sets "to", "off", and "happy".

[0091] Based on the above embodiments describing the subsequence to be verified, the following explanation will continue with step S3-2B in Figure 8:

[0092] S3-2B concatenates the generated text with multiple subsequences to be verified into a complete text sequence, and constructs an attention mask matrix based on the combination order of candidate words in the text sequence.

[0093] In this context, the rows and columns of the attention mask matrix correspond one-to-one with the candidate words in the text sequence. The attention mask matrix includes unrelated elements marked with masks, and the candidate words associated with the unrelated elements do not have an adjacent combination relationship.

[0094] For example, continuing with the tree-like combination diagram shown in Figure 9, and further illustrated in Figure 10, this process will be explained more intuitively. The text sequence obtained based on the tree-like combination diagram in Figure 9 is "going, very, to, off, happy, to, off, happy". Figure 10 shows the corresponding attention mask matrix. The eight rows of this attention mask matrix correspond one-to-one with the eight candidate words in the text sequence, and the eight columns also correspond one-to-one with the eight candidate words in the text sequence. Similar to a conventional attention mask matrix, the elements located in the upper triangle shown in Figure 10 are all marked as unrelated elements. Unlike a conventional attention mask matrix, in the lower triangle shown in Figure 10, elements filled with patterns represent adjacent combination relationships between candidate words located in the element's row and column positions; therefore, elements without patterns can be marked with a mask as unrelated elements.

[0095] Based on the description of the attention mask matrix in step S3-2B of the above embodiments, the following will continue to explain step S3-3B in Figure 8:

[0096] S3-3B uses the attention mask matrix as a constraint for each decoding layer, and processes the feature information of the text sequence through the self-attention mechanism in each decoding layer in turn; and determines the best predicted text based on the decoding result of the last layer among multiple decoding layers.

[0097] Thus, the attention mask matrix ensures that during self-attention calculations, the decoding layer only allows each candidate word to see the previous candidate word in its combination path, avoiding information from other paths and preventing interference between different paths. This mechanism ensures that each candidate word can only focus on the candidate word preceding it, guaranteeing independence and orderliness in the decoding process. It can be understood that the attention mask matrix serves as a constraint for each decoding layer, allowing verification of all candidate words in the entire candidate word set with a single forward inference. Therefore, the attention mask matrix not only improves decoding efficiency but also guarantees that each candidate word relies only on its preceding correct path during decoding, thereby improving the overall accuracy and reliability of the decoding.

[0098] The research also revealed that, compared to conventional attention mask matrices, the attention mask matrix constructed in this embodiment is a sparse matrix, meaning it contains a large number of unrelated elements. Furthermore, during the computation of the self-attention mechanism, even if the weight scores for the positions of unrelated elements are calculated, they are not subsequently included in the weighted summation. This can be understood as the weight scores corresponding to these unrelated elements being ignored during the calculation process, but not affecting the final weighted summation result. Therefore, this embodiment also provides the following optional implementation methods for steps S3-3B to reduce the computational load during the reasoning process of large language models, thereby improving the reasoning speed of large language models.

[0099] To make the implementation of S3-3B easier to understand, the structure of the GPU and the optimization methods for its structure will be explained below.

[0100] First, it's important to understand that GPUs utilize two different levels of memory: SRAM (Static Random-Access Memory) and HBM (High Bandwidth Memory). These differ significantly in performance, capacity, and access speed. SRAM resides on the GPU chip, offering extremely high access speeds but with smaller capacities, making it suitable for storing frequently accessed, small-scale data. HBM, on the other hand, is a high-bandwidth off-chip memory with larger capacities but relatively slower access speeds, making it suitable for storing large-scale data.

[0101] Furthermore, the inference process of large language models often involves a large number of self-attention mechanism operations, including generating query, key, and value matrices, calculating attention scores, normalizing attention weights using the Softmax function, and finally performing a weighted summation of the value matrices to generate the output. Due to the space limitations of SRAM, conventional self-attention mechanisms require frequent data transfer between SRAM and HBM during computation. Frequent reading and writing of data from HBM consumes a significant amount of time, especially when processing large-scale datasets.

[0102] Due to the limited capacity of SRAM, it is difficult to perform calculations directly using SRAM when the data size is large. Therefore, related technologies adopt a block-based strategy to divide the matrix to be calculated into blocks that can be stored in SRAM, so that all calculations can be completed in SRAM step by step, thereby reducing the number of read and write interactions with HBM and reducing the time required for I / O operations.

[0103] Based on the block operation described above, this embodiment provides the following implementation method for step S3-3B:

[0104] S3-3B-1 divides the attention mask matrix into multiple sub-blocks of a preset size, and determines the masking block from the multiple sub-blocks.

[0105] In this context, each element in the masked block is an unrelated element. For example, as shown in Figure 11, for the attention mask matrix shown in Figure 10, assuming a preset size of 2×2, the lower triangular part shown in Figure 10 can be divided into 10 sub-blocks, of which 10 sub-blocks include 3 masked blocks.

[0106] Based on the above description of the sub-blocks and shielding blocks in step S3-3B-1, step S3-3B further includes:

[0107] S3-3B-2, for each decoding layer, divides the query matrix mapped from the text sequence into multiple query sub-blocks according to a preset size.

[0108] In this process, the size of each query sub-block satisfies the matrix operation constraint relationship with the preset size. For example, Figure 12 will be used to illustrate this process more intuitively. As shown in Figure 12, the self-attention mechanism involves a query matrix, a key-value matrix, and an attention weight matrix. The attention weight matrix and the attention mask matrix have the same dimensions, and their elements correspond one-to-one. Assuming that the size of each masked block in the attention mask matrix is ​​2×3, the size of each corresponding block in the attention weight matrix is ​​also 2×3. To obtain a 2×3 sub-block in the attention weight matrix shown in Figure 12, the query matrix in Figure 12 can be divided into 2×4 query sub-blocks, and the key-value matrix can be adaptively divided into 4×3 sub-blocks.

[0109] Based on the above embodiment's explanation of the relationship between query sub-blocks and masked blocks, step S3-3B further includes:

[0110] S3-3B-3, determine at least one target block from multiple query sub-blocks that has no corresponding relationship with the masked block.

[0111] S3-3B-4, based on at least one target block, obtain the attention weight matrix.

[0112] S3-3B-5, based on the attention weight matrix, obtains the decoding result of the decoding layer.

[0113] Thus, as shown in Figure 12, the correspondence between the shielded block and the query sub-block is such that this embodiment only performs calculations on target blocks among the multiple query sub-blocks that have no corresponding relationship with the shielded block, thereby reducing invalid calculations and improving the inference speed of the large language model.

[0114] Furthermore, in practice, it was found that the construction process of the attention mask matrix is ​​closely related to the candidate word set and its size. As the number of candidate words in the candidate word set and the number of candidate words in each set increases, the dimensionality of the attention mask matrix increases significantly, directly leading to larger models requiring more computational resources during inference. However, in-depth analysis of the candidate word set revealed a considerable proportion of redundant terms that conflict with the contextual semantics. Since these redundant terms not only increase computational complexity but may also affect the model's inference accuracy, instead of directly using each initial candidate word decoded from the decoder as the candidate word set decoded from the information to be decoded, the candidate word set in this embodiment can also be obtained by filtering multiple directly decoded initial candidates.

[0115] Therefore, between steps S3 shown in Figure 5, the server can also decode its own information to be decoded from each decoding head to obtain multiple initial word sets, wherein each initial word set includes multiple initial candidate words; then, at least a portion of the text fragments in the generated text are concatenated with the multiple initial word sets to form a sequence to be filtered; the sequence to be filtered is processed by a pre-trained word filtering model to determine the redundant words in each initial word set, wherein the redundant words in each initial word set represent initial candidate words that conflict with the context semantics; finally, the redundant words in each initial word set are removed to obtain the candidate word set decoded from each piece of information to be decoded.

[0116] For example, assume that the candidate word set "going, very" and "to, off, happy" shown in Figure 6 are both initial word sets that have not been filtered. If the server processes the input word filtering model with the sequence to be filtered after concatenating "Take one basketball, we" with "going, very" and "to, off, happy", and the output result shows that "happy" is a redundant word, then it is removed from "to, off, happy". Subsequently, the server uses the remaining "to, off" and "going, very" as the final candidate word set. In this way, by removing redundant words that conflict with the context semantics, not only is the quality of the candidate word set optimized, but the amount of computation required in the subsequent inference process is also significantly reduced, thereby improving the overall efficiency of the model.

[0117] The word selection model described above can also be fine-tuned using a pre-trained model with fewer parameters. For example, using BERT as the base model, a classification layer is added and configured to output the probability of whether each candidate word is redundant. Then, a training dataset is constructed, where each sample includes a concatenated sequence of the generated text fragment and the initial candidate word set, with redundant terms (1 indicating redundancy, 0 indicating non-redundancy) labeled to indicate semantic conflict with the context. Finally, the concatenated text sequence is input into the BERT model to obtain the contextual representation of each candidate word, and the model is trained under supervision using labeled data to obtain the word selection model described above.

[0118] The performance optimization method for the large language model provided in this embodiment was also compared and tested with models of different parameter scales. As shown in Figures 13 and 14, based on the two open-source models Qwen2-7B and Qwen2-72B, Qwen2-7B and Qwen2-72B with sequence information introduced are represented by a darker color, while Qwen2-7B and Qwen2-72B without sequence information are represented by a lighter color. Furthermore, the horizontal axis in the figures represents the selected test dataset, and the vertical axis represents the acceptance rate of candidate words, defined as the average number of candidate words accepted per inference step generated from the decoder. The test results show that after introducing sequence information, an average of 2.1 to 2.8 tokens are obtained per inference step, compared to only 1 token per inference step without sequence information. The extra 1.1 to 1.8 tokens represent the average number of candidate words accepted per step, indicating a significant improvement in the acceptance rate. It should be noted that the large language model performance optimization method provided in this embodiment can be used not only for the Qwen2 series of large language models, but also for other large language models, such as llama, baichuan and other large language models.

[0119] Based on the same inventive concept as the large language model performance optimization method provided in this embodiment, this embodiment also provides a large language model performance optimization device, wherein the large language model includes a main decoding head. It should be understood that the device includes at least one software functional module that can be stored in memory or embedded in an electronic device. The processor in the electronic device is configured to execute the executable module stored in memory. For example, the software functional modules and computer programs included in the device. Referring to Figure 15, functionally, the device may include:

[0120] The text feature module 11 is configured to process the generated text through a large language model to obtain the hidden state information transmitted to the main decoding head. The large language model also includes multiple slave decoding heads running in parallel with the main decoding head. The preset arrangement order among the multiple slave decoding heads represents the arrangement order among the decoding results.

[0121] The feature optimization module 12 is configured to combine the sequence information and hidden state information of each decoder head to obtain the information to be decoded for each decoder head;

[0122] The feature decoding module 13 is configured to obtain the best predicted text for the subsequent generation of the generated text based on the candidate word set decoded from each piece of information to be decoded.

[0123] In this embodiment, the text feature module 11 is configured to implement step S1 in Figure 5, the feature optimization module 12 is configured to implement step S2 in Figure 5, and the feature decoding module 13 is configured to implement step S3 in Figure 5. Therefore, for a detailed description of each of the above modules, please refer to the specific implementation of the corresponding steps.

[0124] Furthermore, it should be understood that, since the large language model performance optimization method provided in this embodiment has the same inventive concept, the large language model performance optimization device can also implement other steps or sub-steps of the method through the above modules.

[0125] Optionally, the feature decoding module 13 is further configured to:

[0126] The candidate words in each candidate word set are combined in a permutation order to obtain multiple candidate texts;

[0127] The best predicted text is determined from multiple candidate texts.

[0128] Optionally, the feature decoding module 13 is further configured to:

[0129] At least some text fragments from the generated text are concatenated with each candidate text to obtain multiple texts to be scored;

[0130] A pre-trained text scoring model is used to score multiple texts to be scored, and a score is obtained for each text to be scored.

[0131] The best predicted text is determined based on the score of each text to be scored.

[0132] Optionally, the branch containing the main decoding head is called the backbone model of the large language model, which includes multiple decoding layers connected in series; the feature decoding module 13 is further configured as follows:

[0133] Each candidate word set is multiplied sequentially according to the order of arrangement to obtain multiple subsequences to be verified. The number of candidate words in each subsequence to be verified is equal to the number of candidate words in the previous subsequence to be verified.

[0134] The generated text is concatenated with multiple subsequences to be verified into a complete text sequence. An attention mask matrix is ​​constructed based on the combination order of candidate words in the text sequence. The rows and columns of the attention mask matrix correspond one-to-one with the candidate words in the text sequence. The attention mask matrix includes unrelated elements marked with masks. Candidate words associated with unrelated elements do not have adjacent combination relationships.

[0135] The attention mask matrix is ​​used as a constraint for each decoding layer. The feature information of the text sequence is processed sequentially through the self-attention mechanism in each decoding layer. The best predicted text is determined based on the decoding result of the last layer among multiple decoding layers.

[0136] Optionally, the feature decoding module 13 is further configured to:

[0137] The attention mask matrix is ​​divided into multiple sub-blocks of a preset size, and a masking block is determined from the multiple sub-blocks, wherein each element in the masking block is an unrelated element;

[0138] For each decoding layer, the query matrix mapped from the text sequence is divided into multiple query sub-blocks according to a preset size, wherein the size of each query sub-block satisfies the constraint relationship of matrix operation with respect to the preset size;

[0139] Identify at least one target block from multiple query sub-blocks that has no corresponding relationship with the masked block;

[0140] Based on at least one target block, obtain the attention weight matrix;

[0141] Based on the attention weight matrix, the decoding result of the decoding layer is obtained.

[0142] Optionally, before obtaining the best predicted text following the generated text based on the candidate word set decoded from each piece of information to be decoded, the feature decoding module 13 is further configured to:

[0143] By decoding the information to be decoded by each decoder head, multiple initial word sets are obtained, where each initial word set includes multiple initial candidate words;

[0144] Concatenate at least some text fragments from the generated text with multiple initial word sets to form a sequence to be filtered;

[0145] The pre-trained word selection model processes the sequence to be selected and identifies redundant words in each initial word set. The redundant words in each initial word set represent initial candidate words that conflict with the context semantics.

[0146] Redundant words are removed from each initial word set to obtain a candidate word set for each piece of information to be decoded.

[0147] Optionally, the feature optimization module 12 is further configured as follows:

[0148] The arrangement position of each decoder head is mapped to the sequence information of each decoder head using a sine or cosine function;

[0149] The sequence information of each decoder head is combined with the hidden state information to obtain the information to be decoded for each decoder head.

[0150] In addition, the functional modules in the various embodiments of this disclosure can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0151] It should also be understood that if the above embodiments are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure.

[0152] Therefore, this embodiment also provides a storage medium, which is a computer-readable storage medium. This storage medium stores a computer program, which, when executed by a processor, implements the large language model performance optimization method provided in this embodiment. The storage medium can be any medium capable of storing program code, such as a USB flash drive, external hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0153] This embodiment provides an electronic device for implementing a large language model performance optimization method. As shown in FIG16, the electronic device may include a processor 22 and a memory 21. The memory 21 stores a computer program, and the processor implements the large language model performance optimization method provided in this embodiment by reading and executing the computer program corresponding to the above-described embodiments in the memory 21.

[0154] Referring again to Figure 14, the electronic device also includes a communication unit 23. The memory 21, processor 22, and communication unit 23 are electrically connected to each other directly or indirectly via system bus 24 to realize data transmission or interaction.

[0155] The memory 21 can be an information recording device based on any electronic, magnetic, optical, or other physical principles, configured to record execution instructions, data, etc. In some embodiments, the memory 21 can be, but is not limited to, volatile memory, non-volatile memory, memory drive, etc.

[0156] In some embodiments, the volatile memory may be random access memory (RAM); in some embodiments, the non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc.; in some embodiments, the storage drive may be a disk drive, solid-state drive, any type of storage disk (such as optical disc, DVD, etc.), or similar storage media, or a combination thereof.

[0157] The communication unit 23 is configured to send and receive data via a network. In some embodiments, the network may include a wired network, a wireless network, a fiber optic network, a telecommunications network, an intranet, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network, a ZigBee network, or a near field communication (NFC) network, or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include wired or wireless network access points, such as base stations and / or network switching nodes, through which one or more components of the service request processing system can connect to the network to exchange data and / or information.

[0158] The processor 22 may be an integrated circuit chip with signal processing capabilities, and may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the processor described above may include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC) computer, or a microprocessor, or any combination thereof.

[0159] It is understood that the structure shown in Figure 14 is for illustrative purposes only. Electronic devices may have more or fewer components than those shown in Figure 14, or may have different configurations. The components shown in Figure 14 may be implemented using hardware, software, or a combination thereof.

[0160] It should be understood that the apparatus and methods disclosed in the above embodiments can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions configured to perform a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0161] The above descriptions are merely various embodiments of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims. Industrial applicability

[0162] This disclosure provides a method, apparatus, storage medium, and electronic device for optimizing the performance of large language models, which can significantly improve the inference speed of large language models while maintaining inference accuracy.

Claims

1. A method for optimizing the performance of a large language model, characterized in that, The large language model includes a main decoding head, and the method includes: The generated text is processed by a large language model to obtain hidden state information transmitted to the main decoding head. The large language model also includes multiple slave decoding heads running in parallel with the main decoding head. The preset arrangement order among the multiple slave decoding heads represents the arrangement order among the decoding results. The sequence information of each decoder head is combined with the hidden state information to obtain the information to be decoded for each decoder head. Based on the candidate word set decoded from each piece of information to be decoded, the best predicted text for the subsequent generated text is obtained.

2. The method for optimizing the performance of a large language model according to claim 1, characterized in that, Based on the candidate word set decoded from each piece of information to be decoded, the best predicted text following the generated text is obtained, including: The candidate words in each of the candidate word sets are combined according to the arranged order to obtain multiple candidate texts; The best predicted text is determined from the multiple candidate texts.

3. The method for optimizing the performance of a large language model according to claim 2, characterized in that, The best predicted text is determined from the plurality of candidate texts, including: At least a portion of the text fragments in the generated text are concatenated with each of the candidate texts to obtain multiple texts to be scored; The multiple texts to be scored are scored by a pre-trained text scoring model to obtain a score for each text to be scored. The best predicted text is determined based on the score of each of the texts to be scored.

4. The method for optimizing the performance of a large language model according to any one of claims 1-3, characterized in that, The branch containing the main decoding head is called the backbone model of the large language model, which includes multiple decoder layers connected in series. Based on the candidate word set decoded from each piece of information to be decoded, the best predicted text following the generated text is obtained, including: Each candidate word set is multiplied sequentially according to the arrangement order to obtain multiple subsequences to be verified, wherein the number of candidate words in each subsequence to be verified is equal to the number of candidate words in the previous subsequence to be verified; The generated text is concatenated with the multiple subsequences to be verified to form a complete text sequence. An attention mask matrix is ​​constructed based on the combination order of candidate words in the text sequence. The rows and columns of the attention mask matrix correspond one-to-one with the candidate words in the text sequence. The attention mask matrix includes unrelated elements marked with masks. The candidate words associated with the unrelated elements do not have an adjacent combination relationship. The attention mask matrix is ​​used as a constraint for each decoding layer, and the feature information of the text sequence is processed sequentially through the self-attention mechanism in each decoding layer; and the best predicted text is determined based on the decoding result of the last layer among the multiple decoding layers.

5. The method for optimizing the performance of a large language model according to claim 4, characterized in that, The attention mask matrix is ​​used as a constraint for each decoding layer, and the feature information of the text sequence is processed sequentially through the self-attention mechanism in each decoding layer, including: The attention mask matrix is ​​divided into multiple molecular blocks of a preset size, and a shielding block is determined from the multiple molecular blocks, wherein each element in the shielding block is the unrelated element; For each of the decoding layers, the query matrix mapped from the text sequence is divided into multiple query sub-blocks according to the preset size, wherein the size of each query sub-block satisfies the constraint relationship of matrix operation with respect to the preset size; From the plurality of query sub-blocks, at least one target block that has no corresponding relationship with the shielded block is determined; Based on the at least one target block, an attention weight matrix is ​​obtained; The decoding result of the decoding layer is obtained based on the attention weight matrix.

6. The method for optimizing the performance of a large language model according to any one of claims 1-5, characterized in that, Before obtaining the best predicted text following the generated text based on the candidate word set decoded from each of the aforementioned information to be decoded, the method further includes: By decoding its own information from each decoding head, multiple initial word sets are obtained, wherein each initial word set includes multiple initial candidate words; At least a portion of the text fragments in the generated text are concatenated with the plurality of initial word sets to form a sequence to be filtered; The sequence to be screened is processed by a pre-trained word selection model to determine redundant words in each initial word set, wherein the redundant words in each initial word set represent initial candidate words that conflict with the context semantics. Redundant words are removed from each initial word set to obtain a candidate word set for each piece of information to be decoded.

7. The method for optimizing the performance of a large language model according to any one of claims 1-6, characterized in that, The hidden state information is an intermediate representation extracted by the large language model during the processing of the generated text, which includes the context information and current state accumulated by the large language model during the processing of the generated text.

8. The method for optimizing the performance of a large language model according to any one of claims 1-7, characterized in that, The sequence information of each decoder head is combined with the hidden state information to obtain the information to be decoded for each decoder head, including: The arrangement position of each decoder head is mapped to the sequence information of each decoder head using a sine or cosine function; The sequence information of each decoder head is combined with the hidden state information to obtain the decoding information of each decoder head.

9. A performance optimization device for large language models, characterized in that, The large language model includes a main decoding head, and the device includes: The text feature module is configured to process the generated text through a large language model to obtain hidden state information transmitted to the main decoding head. The large language model also includes multiple slave decoding heads running in parallel with the main decoding head. The preset arrangement order among the multiple slave decoding heads represents the arrangement order among the decoding results. The feature optimization module is configured to combine the sequence information of each decoder head with the hidden state information to obtain the information to be decoded for each decoder head. The feature decoding module is configured to obtain the best predicted text following the generated text based on the candidate word set decoded from each piece of information to be decoded.

10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the large language model performance optimization method according to any one of claims 1-8.

11. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing a computer program, which, when executed by the processor, implements the large language model performance optimization method according to any one of claims 1-8.