Large Language Model Performance Optimization Method, Device, Storage Medium and Electronic Device

By introducing a method that combines parallel from decoding terminals and sequence information in the large language model, the decoding process is optimized, and the contradiction between the inference speed and accuracy of the generative large language model is solved, achieving more efficient inference speed and more accurate text generation.

CN119862913BActive Publication Date: 2025-07-22SHANGHAI XULU INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510354242.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-22
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing generative large language model is difficult to balance between inference speed and accuracy. Multi-word prediction methods have problems with insufficient accuracy and fluctuations in acceptance rate, and the addition of decoding terminals leads to waste of computing resources.

Method used

By introducing multiple parallel slave decoding terminals into the large language model, using the combination of hidden state information and sequence information, a parallel decoding candidate word set is generated, and the decoding process is optimized through the text scoring model and attention mask matrix, redundant word terms are eliminated, and the inference speed and accuracy are improved.

Benefits of technology

While maintaining inference accuracy, it significantly improves the inference speed of large language models, reduces the waste of computing resources, and improves the acceptance rate of candidate words and the accuracy of decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862913B_ABST
    Figure CN119862913B_ABST
Patent Text Reader

Abstract

The present application provides a method, apparatus, storage medium, and electronic device for optimizing the performance of a large language model, which relates to the field of natural language processing. The electronic device processes the generated text through the large language model to obtain the hidden state information transmitted to the main decoding head; wherein, the large language model further includes a plurality of slave decoding heads parallel to the main decoding head, and the preset arrangement order among the plurality of slave decoding heads represents the arrangement order among the decoding results; then, the sequence information of each slave decoding head is combined with the hidden state information to obtain the information to be decoded for each slave decoding head; finally, according to the candidate word sets decoded from each piece of information to be decoded, the best predicted text subsequent to the generated text is obtained. In this way, after combining the sequence information of each slave decoding head with the hidden state information, the respective information to be decoded is decoded in parallel by a plurality of slave decoding heads, so that the inference speed of the large language model can be significantly improved while maintaining the inference accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing. Specifically, it relates to a method, device, storage medium, and electronic device for optimizing the performance of large language models. Background Art

[0002] Since the advent of ChatGPT, generative large language models have become a hot topic of research and discussion. Most current generative large language models generate text in an autoregressive manner. Specifically, during the generation process, the model predicts the next most likely word based on the generated text sequence, adds it to the sequence, and then repeats this process until the entire text is completed.

[0003] Exemplarily, the following combines Figure 1 to illustrate this process more intuitively. Taking the input text sequence "I go to the basket" as an example, the text sequence "I go to the basket" is input into a large language model called the target model. The target model generates the token of the next word "ball" (for ease of description, hereinafter "word" is used as the description unit, and in actual processing, it is a token) through a forward inference. Then, the newly generated "ball" is appended to the previous input, and the resulting text sequence "I go to the basketball" is input into the target model. The target model generates the next word "court" through the second forward inference, and the newly generated word "court" is appended to the previous input. Continuing, the resulting text sequence "I go to the basketball court" is used as the next input. And so on, to obtain the complete output text of the target model.

[0004] Therefore, this method relies on the probability distribution of the model, that is, calculates the conditional probability of each possible word based on the context and selects the word with the highest probability as the output. However, it is limited to outputting only one word at a time, resulting in low inference efficiency.

[0005] In view of this, several technical solutions for accelerating inference have been proposed during the research process. The most representative one is the multi-word prediction method, that is, generating multiple candidate words at once and selecting the best predicted text from them. After practice, it is found that the prediction accuracy of multiple candidate words is often not ideal, resulting in large fluctuations in the acceptance rate of candidate words. And when the acceptance rate decreases, the inference acceleration effect will be significantly weakened, and may even offset the efficiency improvement brought by multi-word prediction. Other related technologies proposed during the research process, although they can improve the accuracy of multi-word prediction, are at the cost of sacrificing the inference speed. Therefore, how to significantly improve the inference speed of large language models while maintaining the inference accuracy has become a key problem to be solved urgently. Summary of the Invention

[0006] To overcome at least one deficiency in the prior art, this application provides a method, device, storage medium, and electronic device for optimizing the performance of large language models, specifically including:

[0007] In a first aspect, the present application provides a method for optimizing the performance of a large language model, where the large language model includes a main decoding head, and the method includes:

[0008] Processing the generated text through the large language model to obtain hidden state information transmitted to the main decoding head, where the large language model further includes a plurality of slave decoding heads in parallel with the main decoding head, and the preset arrangement order among the plurality of slave decoding heads represents the arrangement order among the decoding results;

[0009] Combining the sequence information of each slave decoding head with the hidden state information to obtain the information to be decoded for each slave decoding head;

[0010] Based on the candidate word sets decoded from each piece of the information to be decoded, obtaining the best predicted text for the subsequent part of the generated text.

[0011] In a second aspect, the present application provides an apparatus for optimizing the performance of a large language model, where the large language model includes a main decoding head, and the apparatus includes:

[0012] A text feature module for processing the generated text through the large language model to obtain hidden state information transmitted to the main decoding head, where the large language model further includes a plurality of slave decoding heads in parallel with the main decoding head, and the preset arrangement order among the plurality of slave decoding heads represents the arrangement order among the decoding results;

[0013] A feature optimization module for combining the sequence information of each slave decoding head with the hidden state information to obtain the information to be decoded for each slave decoding head;

[0014] A feature decoding module for obtaining the best predicted text for the subsequent part of the generated text based on the candidate word sets decoded from each piece of the information to be decoded.

[0015] In a third aspect, the present application provides a storage medium storing a computer program, and when the computer program is executed by a processor, the method for optimizing the performance of the large language model as described above is implemented.

[0016] In a fourth aspect, the present application provides an electronic device, where the electronic device includes a processor and a memory, the memory stores a computer program, and when the computer program is executed by the processor, the method for optimizing the performance of the large language model as described above is implemented.

[0017] Compared with the prior art, the present application has the following beneficial effects:

[0018] In the large language model performance optimization method, device, storage medium, and electronic device provided by this application, the electronic device processes the generated text through the large language model to obtain the hidden state information transmitted to the main decoding head. Among them, the large language model also includes multiple slave decoding heads parallel to the main decoding head, and the preset arrangement order among the multiple slave decoding heads represents the arrangement order among the decoding results. Then, the sequence information of each slave decoding head is combined with the hidden state information to obtain the information to be decoded for each slave decoding head. Finally, the best predicted text following the generated text is obtained according to the candidate word sets decoded from each piece of information to be decoded. In this way, by combining the sequence information of each slave decoding head with the hidden state information, the information to be decoded for each slave decoding head is obtained, and multiple slave decoding heads decode their respective information to be decoded in parallel, thereby being able to significantly improve the inference speed of the large language model while maintaining the inference accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of this application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 It is a schematic diagram of the principle of the autoregressive text generation method provided by the embodiment of this application;

[0021] Figure 2 It is a schematic diagram of the interaction principle between the draft model and the target model provided by the embodiment of this application;

[0022] Figure 3 It is one of the schematic diagrams of the principle of reusing hidden state information provided by the embodiment of this application;

[0023] Figure 4 It is the second schematic diagram of the principle of reusing hidden state information provided by the embodiment of this application;

[0024] Figure 5 It is a schematic flowchart of the large language model performance optimization method provided by the embodiment of this application;

[0025] Figure 6 It is a schematic diagram of the principle of combining hidden state information and sequence information provided by the embodiment of this application;

[0026] Figure 7A It is one of the detailed schematic diagrams of the large language model performance optimization method provided by the embodiment of this application;

[0027] Figure 7BSchematic diagram II of the details of the large language model performance optimization method provided by the embodiments of the present application;

[0028] Figure 8 Tree-shaped combination relationship diagram provided by the embodiments of the present application;

[0029] Figure 9 Attention mask matrix provided by the embodiments of the present application;

[0030] Figure 10 Schematic diagram of the shielding block provided by the embodiments of the present application;

[0031] Figure 11 Schematic diagram of the principle of shielding block shielding calculation provided by the embodiments of the present application;

[0032] Figure 12A Acceptance rate comparison chart based on Qwen2-7B provided by the embodiments of the present application;

[0033] Figure 12B Acceptance rate comparison chart based on Qwen2-72B provided by the embodiments of the present application;

[0034] Figure 13 Schematic diagram of the structure of the large language model performance optimization device provided by the embodiments of the present application;

[0035] Figure 14 Schematic diagram of the structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present application (hereinafter simply referred to as this embodiment) clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0037] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0038] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0039] In the description of the present application, it should be noted that the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance. In addition, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.

[0040] Based on the above statement, as introduced in the background technology, although accelerated reasoning technologies such as multi-word prediction can improve efficiency, the insufficient prediction accuracy leads to fluctuations in acceptance rate and weakened acceleration effect, and technologies to improve accuracy often sacrifice speed. Therefore, how to significantly improve the reasoning speed of large language models while maintaining accuracy has become a key issue at present.

[0041] In order to make the objectives, technical solutions and advantages of the embodiments of the present application easier to understand, the research process of proposing this solution is described in detail below.

[0042] During the research, it was considered that the large model parameters of the target model were the main reason for the slow word-by-word reasoning speed. Therefore, if multiple candidate words could be inferred with a smaller amount of calculation for the target model to screen, the reasoning speed could be improved. Based on this concept, it was proposed to use a model with relatively few model parameters, called the draft model of the target model, to quickly generate multiple candidate words. Then, the target model verifies these candidate words and accepts some of them based on the verification results.

[0043] For example, continue with Figure 1 Take the text sequence "I go to the basket" as an example. Figure 2As shown, the text sequence "I go to the basketball court" is input to the draft model, and the draft model generates 4 candidate words "court fight" after 4 forward reasonings. The 4 candidate words are appended to the text sequence "I go to the basketball court" to obtain the text sequence "I go to the basketball court fight". Then, the text sequence "I go to the basketball court fight" is input to the target model at one time, and the target model generates 5 candidate words "court play basketball" after one forward reasoning. Finally, the 4 candidate words "court fight" generated by the draft model and the 5 candidate words "court play basketball" generated by the target model are compared in turn until different candidate words are compared, and the comparison is stopped, and the same part "court play basketball" is intercepted as the best predicted text for the subsequent text sequence "I go to the basketball court". In this way, the target model only needs to generate 4 words after one forward reasoning. It should be understood that the draft model usually uses a model with a small number of parameters in the same series as the target model, and the reasoning speed is much faster than the target model. When the acceptance rate of the candidate words generated by the draft model is large, the reasoning speed of the target model can be significantly improved.

[0044] However, in practice, it was found that although the above draft model has a small number of parameters, it still takes up a certain amount of computing resources (including GPU resources and memory resources) during reasoning, which limits its promotion and application, especially when computing power is scarce.

[0045] In view of this, the research process further proposed a technology to reuse the hidden state information of the target model to generate multiple candidate words. Before explaining this technology, it is necessary to first understand that the Transformer module is the basic unit of the large language model based on the Transformer architecture. Figure 3 As shown in the figure, multiple slave decoding heads are added in parallel with the main decoding head (Lm_Head) after the last hidden layer of the Transformer module. Each slave decoding head can generate a candidate word set containing at least one candidate word; then, the candidate words in these candidate word sets are verified through a forward reasoning process. Among them, the main decoding head (Lm_Head) is used to convert the hidden state information of the model into actual words in the vocabulary. The role of the slave decoding head is similar to that of the main decoding head, which decodes the same hidden state information in parallel to generate more candidate words, thereby improving the reasoning efficiency and the quality of the candidate text.

[0046] For the above-mentioned multiple slave decoder heads, it is further found in the practical process that there is a preset arrangement order among the newly added multiple slave decoder heads, which is used to represent the arrangement order of decoding results. However, the slave decoder heads cannot perceive their own sequence positions during decoding, resulting in unstable acceptance rates of candidate words sometimes. When the acceptance rate of candidate words becomes low, the acceleration effect is not obvious. Specifically, due to the addition of multiple slave decoder heads, each slave decoder head can generate at least one candidate word. As the number of slave decoder heads and candidate words increases, the number of generated permutation combinations also increases. Therefore, the more candidate words need to be verified in forward inference. However, each additional candidate sequence will increase the computational complexity of the attention mechanism. If the acceptance rate is unstable, it is easy to waste a large amount of computing resources, but actually only a small number of candidate words can be accepted, resulting in a slower inference speed.

[0047] Therefore, inspired by the Long Short-Term Memory (LSTM) model, which carries sequential information between sequences by introducing dependencies of time steps, the following solution is proposed as Figure 4 shown. As Figure 4 shown, the predicted word output by the master decoder head and the hidden state information are input into the first slave decoder head, and the candidate word set output by the first slave decoder head and the hidden state information are together input into the second slave decoder head to obtain the candidate word set output by the second decoder head. And so on, multiple candidate word sets can be generated in sequence. Finally, the candidate words in these candidate word sets are verified through a single forward inference process.

[0048] However, it is found in the practical process that since the slave decoder heads need to perform decoding sequentially, and each decoder also needs to perform a certain amount of operations, when the number of slave decoder heads increases to a certain extent, the inference speed will be significantly reduced.

[0049] In summary, how to significantly improve the inference speed of large language models while maintaining inference accuracy has become a key problem to be solved urgently.

[0050] Based on the discovery of the above technical problems, through creative work, the following technical solutions are proposed to solve or improve the above problems. It should be noted that the defects existing in the above solutions in the prior art are the results obtained through practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by the embodiments of the present application below for the above problems should be regarded as contributions made to the present application during the invention and creation process, rather than being understood as well-known technical content in the art.

[0051] In view of the above technical problems, this embodiment provides a method for optimizing the performance of a large language model. As Figure 5 shown, the method includes:

[0052] S1. Process the generated text through a large language model to obtain the hidden state information transmitted to the main decoding head.

[0053] Among them, the large language model further includes a plurality of slave decoding heads parallel to the main decoding head, and the preset arrangement order among the plurality of slave decoding heads represents the arrangement order among the decoding results.

[0054] S2. Combine the sequence information of each slave decoding head with the hidden state information to obtain the information to be decoded for each slave decoding head.

[0055] S3. Obtain the best predicted text for the subsequent generated text according to the candidate word sets decoded from each piece of information to be decoded.

[0056] In this way, by combining the sequence information of each slave decoding head with the hidden state information, the information to be decoded for each slave decoding head is obtained, and each slave decoding head decodes its respective information to be decoded in parallel, so that the inference speed of the large language model can be significantly improved while maintaining the inference accuracy.

[0057] It should be understood that for the method for optimizing the performance of the large language model provided in this embodiment, the electronic device implementing this method may be, but is not limited to, a mobile terminal, a tablet computer, a laptop computer, a desktop computer, a server, etc. The server may be a single server or a server group. The server group may be centralized or distributed (for example, the server may be a distributed system). In some embodiments, the server may be local or remote relative to the user terminal. In some embodiments, the server may be implemented on a cloud platform; by way of example only, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an inter-cloud, a multi-cloud, etc., or any combination thereof. In some embodiments, the server may be implemented on an electronic device having one or more components.

[0058] To make the solution provided in this embodiment clearer, the following takes the server as the electronic device implementing this method to Figure 5 elaborate on each step of the method shown in detail. However, it should be understood that the operations in the flowchart may not be implemented in sequence, and steps without a logical context relationship may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of this application. Continuing to refer to Figure 5 , the method includes:

[0059] S1. Process the generated text through a large language model to obtain the hidden state information transmitted to the main decoder head.

[0060] Among them, the large language model also includes multiple slave decoder heads parallel to the main decoder head, and the preset arrangement order among the multiple slave decoder heads represents the arrangement order among the decoding results. It can be understood that in the large language model of this embodiment, in addition to the main decoder head, there are multiple parallel slave decoder heads; these slave decoder heads work in parallel with the main decoder head and jointly participate in the decoding process of the model. In addition, different slave decoder heads will generate different candidate words, and the order in which these candidate words are combined into text is determined by the arrangement order of the slave decoder heads. For example, if slave decoder head A is arranged before slave decoder head B, then when the candidate word A generated by slave decoder head A is combined with the candidate word B generated by slave decoder head B, candidate word A needs to be arranged in front of candidate word B.

[0061] In addition, the above-mentioned generated text refers to the part of the text content that has been generated and output by the large language model currently. In addition, the hidden state information refers to the intermediate representation form extracted during the process of the large language model processing the generated text, which contains the context information and the current state accumulated by the model when processing the text. During the processing of the large language model, the server transmits this hidden state information to the main decoder head, and then the subsequent predicted words of the generated text can be inferred.

[0062] As Figure 6 shown, for the above-mentioned multiple slave decoder heads, they are led out from the position before the main decoder head in the forward propagation path, so that the multiple slave decoder heads can simultaneously decode the hidden state information transmitted to the main decoder head. In this embodiment, the branch where the main decoder head is located is called the backbone model of the large language model. As an alternative implementation, the slave decoder head can adopt a structure similar to that of the main decoder head. For example, the structure of each slave decoder head is a single layer or multiple connected layers (linear layer for converting the hidden space to the dictionary dimension). Each layer has a structure of , where is the input of the hidden space, is the activation function, and is the linear layer. Therefore, compared with training an additional draft model, only simple fine-tuning is required for each slave decoder head. During training, only the final hidden state information of the backbone model needs to be used, without going through the layers in the main prediction head processing.

[0063] Based on the introduction of the generated text and hidden state information in step S1 in the above embodiment, continue to refer to Figure 5 , and continue to Figure 5The description of step S2 in

[0064] S2. Combine the sequence information of each slave decoding head with the hidden state information to obtain the information to be decoded for each slave decoding head.

[0065] Regarding this, continue to refer to Figure 6 , the server can map the arrangement position of each slave decoding head to the sequence information of each slave decoding head through a sine function or a cosine function. This mapping method draws on the position encoding technology in the Transformer model. Its core principle is to use a sine or cosine function to generate position vectors, and these position vectors are added to the input word embedding vectors. In this way, the model can understand the position relationship in the input sequence without explicit position information. In this embodiment, a similar method is adopted, but applied to the arrangement positions of the slave decoding heads. That is, the server converts the arrangement position of each slave decoding head into the corresponding sequence information through a sine or cosine function. In this way, the model can understand the relative position of the slave decoding heads during the decoding process, so as to perform the decoding operation more effectively.

[0066] Then, the server further combines the sequence information of each slave decoding head with the hidden state information to obtain the information to be decoded for each slave decoding head. It can be understood that the hidden state information is combined with the sequence information after position encoding to form the information to be decoded. This information to be decoded not only contains context information but also enables each slave decoding head to perceive its position in the arrangement order. Among them, the sequence information helps the slave decoding head better understand the position relationship of the candidate words in the sequence, enhancing the order dependence of the model during the decoding process. In this way, the model can more accurately select the appropriate candidate words, thereby improving the acceptance rate of the candidate words and the accuracy of decoding.

[0067] Based on the above description of the sequence information and the information to be decoded of the slave decoding heads in step S2, next, continue to describe Figure 5 the step S3 in

[0068] S3. According to the candidate word sets decoded from each piece of information to be decoded, obtain the best predicted text for the subsequent generated text.

[0069] Among them, each candidate word set includes at least one candidate word. It can be understood that the server, according to the candidate word sets decoded from each piece of information to be decoded, each candidate word set contains at least one candidate word. These candidate word sets are decoded by multiple decoding heads respectively; then, by evaluating and screening these candidate word sets, at least some of the candidate word sets are selected. Each of the selected candidate word sets will select one candidate word and combine them in the arrangement order to form the best predicted text for the subsequent generated text.

[0070] Exemplarily, the following combines Figure 6 to more intuitively illustrate this process. Figure 6 As shown in Figure 6 , for the generated text "Take one basketball, we", the first candidate word set decoded from the decoder includes "going, very", and the second candidate word set decoded from the decoder includes "to, off, happy"; and the final best predicted text is "going to". It can be seen that "going" is selected from "going, very", and "to" is selected from "to, off, happy".

[0071] Regarding Figure 5 step S3 in Figure 5 , a variety of implementation manners are proposed during the research process. As Figure 7A shown, step S3 may include:

[0072] S3-1A, combining the candidate words in each candidate word set in a permutation order to obtain multiple candidate texts.

[0073] Exemplarily, continuing to refer to Figure 6 , "going, very" and "to, off, happy" can be combined in a permutation order to obtain a total of 6 candidate texts: "going to", "going off", "going happy", "very to", "very off", "very happy". It is necessary to find the best predicted text from these 6 candidate texts. However, it should be understood that when the number of candidate words in each candidate text is very large, the best predicted text may be a partial segment of one of the candidate texts, rather than the entire candidate text.

[0074] S3-2A, determining the best predicted text from multiple candidate texts.

[0075] During the research process, it is found that since the backbone model often contains a large number of parameters, if the backbone model validates the candidate text once, it actually cannot improve the inference speed, but will instead reduce the inference speed. The reason is that Figure 6 6 candidate texts can be obtained in Figure 6 , which means that the backbone model needs to perform 6 forward inferences, and the finally verified candidate text only contains 2 candidate words. However, according to the most conventional inference method, 6 forward inferences can obtain 6 accurate predicted words. In this regard, after further research, it is found that some candidate texts have obvious semantics and grammar. For example, Figure 6Among the six candidate texts obtained, "very to" obviously does not conform to the grammar of English expressions. During the research process, it was first thought to accumulate the probabilities of the candidate words in each candidate text respectively to obtain the score of each candidate text. However, after practice, it was found that this method did not consider the semantic information of the context, resulting in the final retention of the candidate word with the highest probability in each candidate word set. But in many cases, the candidate text composed of the candidate words with the highest probability is not the best text. Therefore, in this embodiment, a model with a relatively small parameter can be used to quickly score the semantics and grammar of the above six candidate texts, and select the one with the highest score for the backbone model to verify.

[0076] Based on the above concept, the server can splice at least some text fragments in the generated text with each candidate text to obtain multiple texts to be scored; use a pre-trained text scoring model to score each text to be scored respectively to obtain the score of each text to be scored; determine the best predicted text according to the scores of each text to be scored.

[0077] It should be understood that in this embodiment, splicing at least some text fragments in the generated text with each candidate text is mainly to provide more context information to make the scoring more accurate. In addition, for the above text scoring model, it can be obtained by fine-tuning the BERT model. As a pre-trained deep learning model, BERT already has the ability to capture the bidirectional context information of the text and is very suitable for processing semantic and grammar tasks. Therefore, only need to use the pre-trained BERT model and train it for specific tasks on this basis to make it adapt to specific scoring requirements.

[0078] In this way, use a text scoring model with a small parameter to score each candidate text, so as to quickly obtain the score of each candidate text, and select the candidate text with the highest score for the backbone model to continue to verify, and determine the best predicted text for the subsequent generated text from it.

[0079] As another optional implementation manner of the above step S3, a tree-shaped attention mask matrix can also be used to complete the verification of the candidate words in multiple candidate word sets through one forward inference. As Figure 7B shown, step S3 may further include:

[0080] S3-1B, multiply each candidate word set in sequence according to the arrangement order to obtain multiple subsequences to be verified.

[0081] Among them, the number of candidate word sets in each subsequence to be verified is equal to the number of candidate words in the previous subsequence to be verified. Exemplarily, continue with Figure 6 the predicted word "are" output by the main solution dock in Figure 8A more intuitive illustration of this process is as follows. As Figure 8 shown in the tree-shaped combination relationship diagram, "are" can be combined with "going, very". Since there is only one "are", the subsequences to be verified obtained by doubling "going, very" include one candidate word set "going, very". Based on the subsequences to be verified obtained by doubling "going, very", "going" and "very" in them can be respectively combined with "to, off, happy". Therefore, the subsequences to be verified obtained by doubling "to, off, happy" include two candidate word sets "to, off, happy".

[0082] Based on the description of each subsequence to be verified in the above embodiments, continue to describe Figure 7B step S3-2B in

[0083] S3-2B: Concatenate the generated text with multiple subsequences to be verified into a complete text sequence, and construct an attention mask matrix according to the combination order of candidate words in the text sequence.

[0084] Among them, the rows and columns of the attention mask matrix respectively correspond one-to-one to the candidate words in the text sequence. The attention mask matrix includes unassociated elements marked with masks, and there is no adjacent combination relationship between the candidate words associated with the unassociated elements.

[0085] Exemplarily, continue to take the Figure 8 shown tree-shaped combination relationship diagram as an example, and combine it with Figure 9 to give a more intuitive illustration of this process. Based on the text sequence obtained from the Figure 8 shown tree-shaped combination relationship diagram, which is "going, very, to, off, happy, to, off, happy", Figure 9 shows the corresponding attention mask matrix. The 8 rows of this attention mask matrix correspond one-to-one to the 8 candidate words in the text sequence, and the 8 columns of the attention mask matrix also correspond one-to-one to the 8 candidate words in the text sequence. Similar to the conventional attention mask matrix, the elements in the Figure 9 shown upper triangle are all marked as unassociated elements. Different from the conventional attention mask matrix, Figure 9 in the Figure 9 shown lower triangle, the elements filled with patterns represent that there is an adjacent combination relationship between the candidate words at the row and column positions of the elements; therefore, masks can be marked on the unmarked elements as unassociated elements.

[0086] Based on the description of the attention mask matrix in step S3-2B in the above embodiments, next continue to describe Figure 7B step S3-3B in

[0087] S3-3B. Use the attention mask matrix as a constraint condition for each decoding layer, and sequentially process the feature information of the text sequence through the self-attention mechanism in each decoding layer; and determine the best predicted text according to the decoding results of the last layer in multiple decoding layers.

[0088] In this way, by the above-mentioned attention mask matrix, when the decoding layer performs self-attention calculation, each candidate word is only allowed to see the previous candidate word in its combination path, and will not see the information of other paths, thus avoiding interference between different paths. Through this mechanism, each candidate word can only focus on the candidate words in front of it, ensuring the independence and sequentiality in the decoding process. It can be understood that by using the attention mask matrix as a constraint condition for each decoding layer, all candidate words in the entire candidate word set can be verified in one forward inference. Therefore, through the above-mentioned attention mask matrix, not only the decoding efficiency is improved, but also it is ensured that each candidate word only depends on the correct path before it during the decoding process, thereby improving the overall decoding accuracy and reliability.

[0089] During the research process, it was also found that compared with the conventional attention mask matrix, the attention mask matrix constructed in this embodiment is a sparse matrix, that is, there are a large number of uncorrelated elements. In the calculation process of the self-attention mechanism, even if the weight scores at the positions of the uncorrelated elements are calculated, they will not participate in the weighted summation subsequently. It can be understood that the weight scores corresponding to these uncorrelated elements will be ignored during the calculation process, but will not affect the final weighted summation result. In view of this, this embodiment also provides the following optional implementation manner of step S3-3B to reduce the calculation amount during the inference process of the large language model, thereby improving the inference speed of the large language model. To make the implementation manner of S3-3B easier to understand, the structure of the GPU and the optimization method for its structure will be explained first below.

[0090] First, it should be understood that there are two different levels of memory in the GPU, namely SRAM (Static Random Access Memory) and HBM (High Bandwidth Memory). There are significant differences in their performance, capacity, and access speed. Among them, SRAM is located on the chip of the GPU, has an extremely high access speed, but a small capacity, and is suitable for storing small-scale data that is frequently accessed. HBM is a high-bandwidth off-chip memory with a large capacity but a relatively slow access speed, and is suitable for storing large-scale data.

[0091] In addition, the inference process of large language models often involves a large number of self-attention mechanism operations, including generating query, key, and value matrices, calculating attention scores, normalizing them through the Softmax function to obtain attention weights, and finally performing weighted summation on the value matrix to generate the output. Limited by the space of SRAM, during the calculation process, the conventional self-attention mechanism needs to frequently transfer data between SRAM and HBM. Frequent reading and writing of data from HBM takes a lot of time, especially when dealing with large-scale data.

[0092] Limited by the small capacity of SRAM, it is difficult to directly use SRAM for calculation when the data scale is relatively large. Therefore, in related technologies, a block strategy is adopted to divide the matrix to be calculated into blocks that SRAM can store, so as to gradually complete all calculations in SRAM, thereby reducing the number of read and write interactions with HBM and the time required for I / O operations.

[0093] Based on the block operation described above, the following implementation manner of step S3-3B is provided in this embodiment:

[0094] S3-3B-1, divide the attention mask matrix into multiple sub-blocks of a preset size, and determine the masked blocks from the multiple sub-blocks.

[0095] Among them, each element in the masked block is an irrelevant element. Exemplarily, as Figure 10 shown, for the Figure 9 shown attention mask matrix, assuming the preset size is , then the lower triangular part shown in Figure 9 can be divided into 10 sub-blocks, and 3 of the 10 sub-blocks are masked blocks.

[0096] Based on the above introduction of the sub-blocks and masked blocks in step S3-3B-1, step S3-3B further includes:

[0097] S3-3B-2, for each decoding layer, divide the query matrix mapped from the text sequence into multiple query sub-blocks according to the preset size.

[0098] Among them, the size of each query sub-block satisfies the constraint relationship of matrix operations with the preset size. Exemplarily, the following combines Figure 11 to illustrate this process more intuitively. As Figure 11 shown, the figure shows the query matrix, key-value matrix, and attention weight matrix involved in the self-attention mechanism. The attention weight matrix has the same dimension as the attention mask matrix, and the elements correspond one by one. Assume that the size of each masked block in the attention mask matrix is , then the size of each block corresponding to it in the attention weight matrix is also . If you want to obtain Figure 11 the sub-block with size in the shown attention weight matrix, you can divide the query matrix in Figure 11 into query sub-blocks of , and adaptively divide the key-value matrix into sub-blocks of .

[0099] Based on the description of the relationship between the query sub-block and the masking block in the above embodiments, step S3-3B further includes:

[0100] S3-3B-3, determining at least one target block that has no corresponding relationship with the masking block from multiple query sub-blocks.

[0101] S3-3B-4, obtaining the attention weight matrix according to at least one target block.

[0102] S3-3B-5, obtaining the decoding result of the decoding layer according to the attention weight matrix.

[0103] In this way, just as Figure 11 shows the corresponding relationship between the masking block and the query sub-block, in this embodiment, only the target blocks that have no corresponding relationship with the masking block among multiple query sub-blocks are calculated, thereby reducing invalid calculations and improving the inference speed of the large language model.

[0104] In addition, it is also found in the practical process that the construction process of the attention mask matrix is closely related to the candidate word set and its scale. As the candidate word set and the number of candidate words in each candidate word set increase, the dimension of the attention mask matrix will increase significantly, which directly leads to the large model consuming more computing resources during the inference process. However, through in-depth analysis of the candidate word set, it is found that there are a considerable proportion of redundant word items that conflict with the context semantics. Since these redundant word items not only increase the computational complexity but also may affect the inference accuracy of the model, therefore, compared with directly using each of the multiple initial candidates directly decoded from the decoder head as the candidate word set decoded from the information to be decoded, the candidate word set in this embodiment can also be obtained by screening the multiple initial candidates directly decoded. Therefore, in Figure 5Between the steps S3 shown, the server can also decode its own information to be decoded through each slave decoding head to obtain a plurality of initial word sets, where each initial word set includes a plurality of initial candidate words; then, at least some text segments in the generated text are spliced with the plurality of initial word sets into a sequence to be screened; the sequence to be screened is processed by a pre-trained word screening model to determine redundant words in each initial word set, where the redundant words in each initial word set represent initial candidate words that conflict with the context semantics; finally, the redundant words in each initial word set are removed to obtain a candidate word set decoded from each piece of information to be decoded.

[0105] Exemplarily, assume Figure 6 The candidate word sets "going, very", "to, off, happy" shown are all initial word sets that have not been screened. If the server inputs the sequence to be screened after splicing "Take one basketball, we" with "going, very", "to, off, happy" into the word screening model for processing, and the output result shows that "happy" is a redundant word, then it is removed from "to, off, happy". Subsequently, the server takes the remaining "to, off" and "going, very" as the final candidate word set. In this way, by removing redundant words that conflict with the context semantics, not only the quality of the candidate word set is optimized, but also the computational amount required in the subsequent inference process is significantly reduced, thereby improving the overall efficiency of the model.

[0106] The above word screening model can also be obtained by fine-tuning a pre-trained model with fewer model parameters. Exemplarily, still taking BERT as an example, the BERT model is used as the basic model, and a classification layer is added on this basis to output the probability of whether each candidate word is a redundant word. Then, a training data set is constructed. Each sample in the training data set includes the spliced sequence of the generated text segment and the initial candidate word set, and the redundant word items that conflict with the context semantics are marked (1 represents redundant, 0 represents non-redundant). Finally, the spliced text sequence is input into the BERT model to obtain the context representation of each candidate word, and the model is supervised and trained using the marked data to obtain the above word screening model.

[0107] For the large language model performance optimization method provided in this embodiment, it is also compared and tested on models with different parameter scales. Such as Figure 12A and Figure 12BAs shown, based on the two open source models of Qwen2-7B and Qwen2-72B, Qwen2-7B and Qwen2-72B with sequence information introduced are represented in blue, and Qwen2-7B and Qwen2-72B without sequence information are represented in green. In addition, the horizontal axis in the figure represents the selected test data set, and the vertical axis represents the acceptance rate of the candidate words, which is defined as the average number of candidate words accepted per step. The test results show that after the introduction of sequence information, on average, 2.1~2.8 candidate words can be obtained per reasoning, and the acceptance rate has been significantly improved. It should be noted that the large language model performance optimization method provided in this embodiment can not only be used for the large language model of the Qwen2 series, but also can be applied to other large language models, for example, large language models such as llama and baichuan.

[0108] Based on the same inventive concept as the large language model performance optimization method provided in this embodiment, this embodiment also provides a large language model performance optimization device, and the large language model includes a main decoding head. It should be understood that the device includes at least one software function module that can be stored in a memory or solidified in an electronic device in the form of software. The processor in the electronic device is used to execute the executable module stored in the memory. For example, the software function module and computer program included in the device. Please refer to Figure 13 , functionally speaking, the device may include:

[0109] The text feature module 11 is used to process the generated text through the large language model to obtain hidden state information transmitted to the main decoding head, wherein the large language model also includes a plurality of slave decoding heads in parallel with the main decoding head, and the preset arrangement order between the plurality of slave decoding heads represents the arrangement order between the decoding results;

[0110] A feature optimization module 12, used to combine the sequence information of each slave decoding head with the hidden state information to obtain the information to be decoded of each slave decoding head;

[0111] The feature decoding module 13 is used to obtain the best predicted text of the generated text according to the candidate word set decoded from each piece of information to be decoded.

[0112] In this embodiment, the text feature module 11 is used to implement Figure 5 In step S1, the feature optimization module 12 is used to implement Figure 5 In step S2, the feature decoding module 13 is used to implement Figure 5 Therefore, for the detailed description of the above modules, please refer to the specific implementation of the corresponding steps.

[0113] In addition, it should also be understood that, due to having the same inventive concept as the large language model performance optimization method provided in this embodiment, the large language model performance optimization device can also implement other steps or sub-steps of this method through the above-mentioned modules.

[0114] Optionally, the feature decoding module 13 is further specifically configured to:

[0115] Combine the candidate words in each candidate word set in a permutation order to obtain multiple candidate texts;

[0116] Determine the best predicted text from the multiple candidate texts.

[0117] Optionally, the feature decoding module 13 is further specifically configured to:

[0118] Concatenate at least some text segments in the generated text with each candidate text to obtain multiple texts to be scored;

[0119] Score each text to be scored respectively through a pre-trained text scoring model to obtain the score of each text to be scored;

[0120] Determine the best predicted text according to the scores of each text to be scored.

[0121] Optionally, the branch where the main decoding head is located is called the backbone model of the large language model, and the backbone model includes a plurality of decoding layers connected in series; the feature decoding module 13 is further specifically configured to:

[0122] Multiply each candidate word set in turn according to the permutation order to obtain multiple subsequences to be verified, where the number of candidate word sets in each subsequence to be verified is equal to the number of candidate words in the previous subsequence to be verified;

[0123] Concatenate the generated text with the multiple subsequences to be verified into a complete text sequence, and construct an attention mask matrix according to the combination order of the candidate words in the text sequence, where the rows and columns of the attention mask matrix correspond to the candidate words in the text sequence one by one, and the attention mask matrix includes unassociated elements marked with masks, and there is no adjacent combination relationship between the candidate words associated with the unassociated elements;

[0124] Use the attention mask matrix as the constraint condition for each decoding layer, and process the feature information of the text sequence through the self-attention mechanism in each decoding layer in turn; and determine the best predicted text according to the decoding result of the last decoding layer among the multiple decoding layers.

[0125] Optionally, the feature decoding module 13 is further specifically configured to:

[0126] The attention mask matrix is divided into multiple sub-blocks of a preset size, and a shielding block is determined from the multiple sub-blocks, where each element in the shielding block is an uncorrelated element;

[0127] For each decoding layer, the query matrix mapped from the text sequence is divided into multiple query sub-blocks according to a preset size, where the size of each query sub-block satisfies the constraint relationship of matrix operations with the preset size;

[0128] At least one target block that has no corresponding relationship with the shielding block is determined from the multiple query sub-blocks;

[0129] An attention weight matrix is obtained according to the at least one target block;

[0130] A decoding result of the decoding layer is obtained according to the attention weight matrix.

[0131] Optionally, before obtaining the best predicted text for the subsequent generated text according to the candidate word sets decoded from each piece of information to be decoded, the feature decoding module 13 is further configured to:

[0132] Each decoding head decodes its own information to be decoded to obtain multiple initial word sets, where each initial word set includes multiple initial candidate words;

[0133] At least some text segments in the generated text are concatenated with the multiple initial word sets to form a sequence to be screened;

[0134] The sequence to be screened is processed by a pre-trained word screening model to determine the redundant words in each initial word set, where the redundant words in each initial word set represent the initial candidate words that conflict with the context semantics;

[0135] The redundant words in each initial word set are removed to obtain the candidate word sets decoded from each piece of information to be decoded.

[0136] Optionally, the feature optimization module 12 is further specifically configured to:

[0137] Map the arrangement position of each decoding head to the sequence information of each decoding head through a sine function or a cosine function;

[0138] The sequence information of each decoding head is combined with the hidden state information to obtain the information to be decoded of each decoding head.

[0139] In addition, in each embodiment of the present application, the various functional modules may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.

[0140] It should also be understood that if the above embodiments are implemented in the form of software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.

[0141] Therefore, this embodiment also provides a storage medium, which is a computer-readable storage medium. This storage medium stores a computer program, and when the computer program is executed by a processor, it realizes the large language model performance optimization method provided by this embodiment. Among them, this storage medium can be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc.

[0142] An electronic device for implementing the large language model performance optimization method provided by this embodiment. As Figure 14 described, this electronic device may include a processor 22 and a memory 21. And, the memory 21 stores a computer program, and the processor realizes the large language model performance optimization method provided by this embodiment by reading and executing the computer program corresponding to the above embodiments in the memory 21.

[0143] Continue to refer to Figure 14 , this electronic device also includes a communication unit 23. Each element of the memory 21, the processor 22, and the communication unit 23 is directly or indirectly electrically connected through a system bus 24 to realize data transmission or interaction.

[0144] Among them, the memory 21 can be an information recording device based on any electronic, magnetic, optical, or other physical principles, and is used to record execution instructions, data, etc. In some embodiments, the memory 21 can be, but is not limited to, a volatile memory, a non-volatile memory, a storage drive, etc.

[0145] In some embodiments, the volatile memory may be a Random Access Memory (RAM); in some embodiments, the non-volatile memory may be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electric Erasable Programmable Read-Only Memory (EEPROM), a flash memory, etc.; in some embodiments, the storage drive may be a disk drive, a solid state drive, any type of storage disk (such as an optical disk, a DVD, etc.), or a similar storage medium, or a combination thereof, etc.

[0146] The communication unit 23 is configured to transmit and receive data via a network. In some embodiments, the network may include a wired network, a wireless network, an optical fiber network, a telecommunication network, an intranet, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Networks (WLAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a Public Switched Telephone Network (PSTN), a Bluetooth network, a ZigBee network, or a Near Field Communication (NFC) network, etc., or any combination thereof. In some embodiments, the network may include one or more network access points. For example, the network may include a wired or wireless network access point, such as a base station and / or a network switching node, and one or more components of the service request processing system may be connected to the network via the access point to exchange data and / or information.

[0147] The processor 22 may be an integrated circuit chip with signal processing capabilities, and the processor may include one or more processing cores (e.g., a single-core processor or a multi-core processor). By way of example only, the above-mentioned processor may include a Central Processing Unit (CPU), an Application Specific Integrated Circuit (ASIC), an Application Specific Instruction-set Processor (ASIP), a Graphics Processing Unit (GPU), a Physics Processing Unit (PPU), a Digital Signal Processor (DSP), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a microcontroller unit, a Reduced Instruction Set Computing (RISC), or a microprocessor, etc., or any combination thereof.

[0148] It can be understood that Figure 14 The structure shown is only illustrative. The electronic device may also have more or fewer components than Figure 14 shown, or have a different configuration from Figure 14 that shown. Figure 14 Each of the components shown may be implemented in hardware, software, or a combination thereof.

[0149] It should be understood that the devices and methods disclosed in the above embodiments can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0150] As described above, these are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for optimizing the performance of a large language model, characterized in that, The large language model includes a main decoding head, and the method includes: Processing the generated text through the large language model to obtain hidden state information transmitted to the main decoding head. The large language model further includes a plurality of slave decoding heads parallel to the main decoding head, and a preset arrangement order among the plurality of slave decoding heads represents the arrangement order among decoding results; Mapping the arrangement position of each slave decoding head to sequence information of each slave decoding head through a sine function or a cosine function; Combining the sequence information of each slave decoding head with the hidden state information to obtain decoding information to be decoded for each slave decoding head; Decoding the decoding information to be decoded for itself through each slave decoding head to obtain a plurality of initial word sets, where each initial word set includes a plurality of initial candidate words; Concatenating at least some text segments in the generated text with the plurality of initial word sets to form a sequence to be screened; Processing the sequence to be screened through a pre-trained word screening model to determine redundant words in each initial word set, where the redundant words in each initial word set represent initial candidate words conflicting with the context semantics; Removing the redundant words in each initial word set to obtain a candidate word set decoded from each decoding information to be decoded; Obtaining the best predicted text subsequent to the generated text according to the candidate word sets decoded from each decoding information to be decoded, where these candidate word sets are decoded by the plurality of decoding heads respectively.

2. The method for optimizing the performance of a large language model according to claim 1, wherein Obtaining the best predicted text subsequent to the generated text according to the candidate word sets decoded from each decoding information to be decoded, includes: Combining the candidate words in each candidate word set in the arrangement order to obtain a plurality of candidate texts; Determining the best predicted text from the plurality of candidate texts.

3. The method for optimizing the performance of a large language model according to claim 2, wherein Determining the best predicted text from the plurality of candidate texts, includes: Concatenating at least some text segments in the generated text with each candidate text to obtain a plurality of texts to be scored; Scoring each of the plurality of texts to be scored through a pre-trained text scoring model to obtain the score of each text to be scored; Determining the best predicted text according to the scores of each text to be scored.

4. The method for optimizing the performance of a large language model according to claim 1, characterized in that, The branch where the main decoding head is located is called the backbone model of the large language model, and the backbone model includes a plurality of decoding layers connected in series; obtaining the best predicted text subsequent to the generated text according to the candidate word sets decoded from each decoding information to be decoded, includes: Doubling each candidate word set in sequence according to the arrangement order to obtain a plurality of subsequences to be verified, where the number of candidate word sets in each subsequence to be verified is equal to the number of candidates in the previous subsequence to be verified; Concatenate the generated text with the multiple subsequences to be verified into a complete text sequence, and construct an attention mask matrix according to the combination order of candidate words in the text sequence. The rows and columns of the attention mask matrix respectively correspond one-to-one to the candidate words in the text sequence. The attention mask matrix includes unassociated elements marked with masks, and there is no adjacent combination relationship between the candidate words associated with the unassociated elements; Use the attention mask matrix as a constraint condition for each decoding layer, and sequentially process the feature information of the text sequence through the self-attention mechanism in each decoding layer; and determine the best predicted text according to the decoding result of the last layer in the multiple decoding layers.

5. The method for optimizing the performance of a large language model according to claim 4, characterized in that Using the attention mask matrix as a constraint condition for each decoding layer, and sequentially processing the feature information of the text sequence through the self-attention mechanism in each decoding layer includes: Divide the attention mask matrix into multiple sub-blocks of a preset size, and determine a shielding block from the multiple sub-blocks. Each element in the shielding block is the unassociated element; For each decoding layer, divide the query matrix mapped from the text sequence into multiple query sub-blocks according to the preset size, where the size of each query sub-block satisfies the constraint relationship of matrix operations with the preset size; Determine at least one target block from the multiple query sub-blocks that has no corresponding relationship with the shielding block; Obtain an attention weight matrix according to the at least one target block; Obtain the decoding result of the decoding layer according to the attention weight matrix.

6. An apparatus for optimizing the performance of a large language model, characterized in that, The large language model includes a main decoding head, and the device includes: A text feature module for processing the generated text through the large language model to obtain the hidden state information transmitted to the main decoding head. The large language model also includes multiple slave decoding heads parallel to the main decoding head, and the preset arrangement order between the multiple slave decoding heads represents the arrangement order between the decoding results; A feature optimization module for mapping the arrangement position of each slave decoding head to the sequence information of each slave decoding head through a sine function or a cosine function; combining the sequence information of each slave decoding head with the hidden state information to obtain the information to be decoded of each slave decoding head; A feature decoding module for decoding the information to be decoded of each slave decoding head through each slave decoding head to obtain multiple initial word sets, where each initial word set includes multiple initial candidate words; Concatenate at least some text segments in the generated text with the multiple initial word sets into a sequence to be screened; Process the sequence to be screened through a pre-trained word screening model to determine the redundant words in each initial word set, where the redundant words in each initial word set represent the initial candidate words that conflict with the context semantics; Remove the redundant words in each initial word set to obtain the candidate word set decoded from each piece of information to be decoded; Based on the candidate word sets decoded from each of the to-be-decoded information, obtain the best predicted text following the generated text, where these candidate word sets are respectively decoded by the multiple decoding heads.

7. A storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the large language model performance optimization method according to any one of claims 1-5.

8. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory stores a computer program, and when the computer program is executed by a processor, it implements the large language model performance optimization method according to any one of claims 1-5.