Decoding acceleration method and system of large language model, electronic equipment and storage medium
By weighting the output data of each network layer of the large language model and aborting the calculation in advance in the inference stage, the problem of large amount of calculation in the decoding process in the prior art is solved, and the decoding efficiency is improved.
Patent Information
- Application Number
- CN202510159850.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-17
AI Technical Summary
The existing large language models have huge computational volumes during the decoding process, resulting in slow decoding speed and high cost.
During the training process of the large language model, the output data of each network layer is weighted, and the real-time output data meets the preset conditions during the inference stage, and subsequent calculations are aborted in advance to achieve decoding acceleration.
It effectively reduces the computational volume, improves decoding efficiency, and is suitable for large language models of different architectures.
Smart Images

Figure CN120163238A_ABST
Abstract
Description
Background Art
[0002] The latest progress in large language models (LLMs) has achieved breakthroughs in language understanding and language generation, covering almost all natural language processing (NLP) tasks widely used in this field today. In particular, the release of OpenAI ChatGPT and GPT-4 has sparked a widespread research boom. This autoregressive language modeling provides a flexible framework for solving complex tasks, with a unified natural language input and output data format, while also relaxing the need for large-scale task-specific data collection and training.
[0003] However, in recent years, the development trend of language models has also been obvious. People have tried to continuously increase the number of model parameters and the amount of training data in order to obtain more powerful decoding capabilities. Although some recent progress enables LLMs to be effectively trained on large amounts of data and can indeed achieve impressive results, the models trained in this way still face huge inference pressures in actual use, with slow decoding speeds and high costs.
[0004] Because as an autoregressive language model, when generating text during inference, it is similar to the process of human speech or writing in the present invention (word by word). For example, in machine translation, this is particularly obvious in the autoregressive decoding process, in which the entire stack of Transformer layers is repeatedly calculated for each input data. That is, it is necessary to further predict the next word based on the previously generated words. This is a serial process, and the language model cannot parallelize this process. The present invention cannot generate a coherent sentence jumpily. Considering the currently common language models, which contain at least hundreds of millions to trillions of parameters, decoding each word requires a large amount of computation, which is undoubtedly a huge challenge. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a decoding acceleration method, system, electronic device, and storage medium for large language models in view of the deficiencies of the prior art, specifically as follows:
[0006] 1) In the first aspect, the present invention provides a decoding acceleration method for large language models, and the specific technical solution is as follows:
[0007] During the training of the large language model, weight processing is performed on the output data of each network layer in the large language model;
[0008] In the process of processing the current input data using the trained large language model, whenever the output data of a network layer in the trained large language model is obtained, it is determined whether the output data of the network layer obtained in real time meets the preset conditions until the network layer with a judgment result of "yes" is determined. The prediction result is determined according to the output data of this network layer, and the prediction result is combined with the current input data and then re-input into the trained large language model.
[0009] The beneficial effects of a decoding acceleration method for a large language model provided by the present invention are as follows:
[0010] It improves the existing decoding process of the large language model, determines whether the output data of the network layer obtained in real time meets the preset conditions, and can effectively reduce the calculation amount during the inference process using the trained large language model, thereby improving the decoding efficiency. Moreover, it is applicable to large language models with different architectures and has strong applicability.
[0011] On the basis of the above solution, a decoding acceleration method for a large language model of the present invention can also be improved as follows.
[0012] Further, during the training process of the large language model, weight processing is performed on the output data of each network layer in the large language model, including:
[0013] During the training process of the large language model, the output data of each network layer in the large language model is weighted using the modified loss function.
[0014] Further, when the output data of each network layer in the large language model is weighted using the modified loss function, among every two adjacent network layers in the large language model, the weight of the output data of the previous network layer is higher than the weight of the output data of the subsequent network layer.
[0015] Further, the large language model is a large language model based on the Transformer architecture, a large language model based on the Seq2Seq architecture, or a large language model based on a pure decoder architecture.
[0016] 2) In the second aspect, the present invention also provides a decoding acceleration system for a large language model, and the specific technical solution is as follows:
[0017] It includes a weight processing module and a data processing module;
[0018] The weight processing module is used for: during the training process of the large language model, performing weight processing on the output data of each network layer in the large language model;
[0019] The data processing module is used for: during the process of processing the current input data by using the trained large language model, whenever the output data of the network layer in the trained large language model is obtained, it is determined whether the output data of the network layer obtained in real time meets the preset conditions until the network layer whose judgment result is "yes" is determined, and the prediction result is determined according to the output data of this network layer, and the prediction result is combined with the current input data and then re-input into the trained large language model.
[0020] On the basis of the above solution, a decoding acceleration system for a large language model of the present invention can also be improved as follows.
[0021] Further, the weighting processing module is specifically used for:
[0022] During the training process of the large language model, the output data of each network layer in the large language model is weighted by using the modified loss function.
[0023] Further, when the output data of each network layer in the large language model is weighted by using the modified loss function, among every two adjacent network layers in the large language model, the weight of the output data of the previous network layer is higher than the weight of the output data of the next network layer.
[0024] Further, the large language model is a large language model based on the Transformer architecture, a large language model based on the Seq2Seq architecture, or a large language model based on a pure decoder architecture.
[0025] 3) Thirdly, the present invention also provides an electronic device, which includes a processor. The processor is coupled with a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor so that the electronic device implements any one of the above decoding acceleration methods for the large language model.
[0026] 4) Fourthly, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above decoding acceleration methods for the large language model.
[0027] It should be noted that for the beneficial effects obtained by the technical solutions and corresponding possible implementation manners of the second to fourth aspects of the present invention, reference can be made to the technical effects of the first aspect and its corresponding possible implementation manners above, which will not be elaborated here. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention:
[0029] Figure 1 It is: a schematic flowchart of a decoding acceleration method for a large language model according to an embodiment of the present invention;
[0030] Figure 2 It is: a schematic diagram of the decoding process of a large language model based on the Transformer architecture in the prior art;
[0031] Figure 3 It is: a schematic diagram of the decoding process of a large language model based on the Transformer architecture in the present invention;
[0032] Figure 4 It is: a schematic structural diagram of a decoding acceleration system for a large language model according to an embodiment of the present invention;
[0033] Figure 5 It is: a schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0034] The principles and features of the present invention are described below. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0035] The technical solutions of the present invention and how the technical solutions of the present invention solve the above technical problems are described in detail below with specific embodiments. These several specific embodiments can be combined with each other. Concepts or processes that are the same or similar may not be repeated in some embodiments. The embodiments of the present invention will be described below with reference to the accompanying drawings.
[0036] As Figure 1 shown, a decoding acceleration method for a large language model according to an embodiment of the present invention includes the following steps:
[0037] S1. During the training of the large language model, weight processing is performed on the output data of each network layer in the large language model;
[0038] Among them, the large language model can be: a large language model for device credibility evaluation, a large language model for personalized education and service information recommendation, a large language model for product recommendation, a large language model for knowledge base query, a large language model for product recommendation based on dialogue analysis, etc. The datasets used in the training of large language models applied to different fields are different and can be set according to actual situations. For example, the dataset used in the training of a large language model for product recommendation based on dialogue analysis is: dialogue texts and corresponding product recommendation information.
[0039] S2. During the process of processing the current input data using the trained large language model, whenever the output data of a network layer in the trained large language model is obtained, it is determined whether the real-time obtained output data of the network layer meets the preset conditions until a network layer with a judgment result of "yes" is determined. According to the output data of this network layer, a prediction result is determined, and the prediction result is combined with the current input data and then re-input into the trained large language model.
[0040] Optionally, in the above technical solution, during the training process of the large language model, weighted processing is performed on the output data of each network layer in the large language model, including:
[0041] During the training process of the large language model, the modified loss function is used to perform weighted processing on the output data of each network layer in the large language model.
[0042] Optionally, in the above technical solution, when using the modified loss function to perform weighted processing on the output data of each network layer in the large language model, among every two adjacent network layers in the large language model, the weight of the output data of the previous network layer is higher than the weight of the output data of the subsequent network layer.
[0043] Optionally, in the above technical solution, the large language model is a large language model based on the Transformer architecture, a large language model based on the Seq2Seq architecture, or a large language model based on a pure decoder architecture.
[0044] Taking the large language model based on the Transformer architecture as an example, and combining the training method and usage method of the large language model in the prior art, the technical solution and technical effect of the present invention are described as follows.
[0045] As Figure 2 shown, when given the input data ("I") to the Transformer architecture-based large language model, all network layers in the Transformer architecture-based large language model need to participate in the calculation until the last network layer, and the output data ("love") is decoded. Then, repeated decoding is performed to generate the next output data ("spring") until all the content ("I love the spring when all the flowers bloom") is decoded. The computational amount of this method is undoubtedly huge. On this basis, a decoding acceleration method for a large language model of the present invention is proposed, specifically as follows:
[0046] I. Model training stage of the large language model based on the Transformer architecture:
[0047] S101. In the training stage, the present invention first collects some input sentences. For example, in Figure 2 and Figure 3Sentences such as "I love the blooming flowers in spring" in it. Since the large language model based on the Transformer architecture is an autoregressive model, these input sentences can be used both as input and as misaligned output data.
[0048] S102. Sequentially use the tokens in the input sentence (each character in "I love the blooming flowers in spring" can be a token) as input data to calculate the output data of each network layer in the Transformer architecture. For example, the large language model based on the Transformer architecture has a total of 9 network layers ( Figure 2 and Figure 3 the rectangles in it represent network layers). Taking the token "I" as an example, when passing through each network layer (which can also be called a block) in the Transformer architecture, the output data of each network layer will be obtained.
[0049] In the existing model training process, as Figure 2 shown, only the output data of the topmost network layer (the last network layer of the large language model based on the Transformer architecture, and the sequence numbers of all network layers are determined according to the data transmission order) is used. The softmax function is used to calculate the output data of this network layer to obtain the probability distribution result of each vocabulary in the vocabulary table, and the vocabulary at the index position with the highest probability in the output data vocabulary table is used as the prediction result (output) of the entire model. The output data of the intermediate network layers (all network layers except the topmost network layer are intermediate network layers) is only used as the input data of the next network layer for data transmission.
[0050] In the present invention, the output data of each network layer will be collected (in this implementation, 9 output data will be collected). Different from the "way of calculating loss based on the output data of the topmost network layer" in the prior art, the present invention modifies the loss data to obtain a modified loss function. During the training process of the large language model, the modified loss function is used to perform weighted processing on the output data of each network layer in the large language model.
[0051] Among them, the modified loss function is: L represents: the number of network layers in the large language model, ω i represents: the weight assigned to the i-th network layer in the large language model, represents: the loss of the i-th network layer in the large language model, both j and i are positive integers and are not greater than L, represents the calculated loss.
[0052] Through "In every two adjacent network layers in the large language model, the weight of the output data of the previous network layer is higher than that of the output data of the subsequent network layer". For example, when the large language model has 3 network layers, the weight of the first network layer is The weight of the second network layer is (that is ), and the weight of the third network layer is (that is ). The purpose is to encourage the large language model to generate representations meaningful to the prediction results between network layers. Therefore, the modified loss function proposed in this invention no longer uses only the output data of the last network layer for prediction, but uses the weighted average of the predictions of the network layers for prediction. Moreover, it is more inclined to make the weights of the subsequent network layers larger. Therefore, the expressive ability of the later network layers is relatively stronger. Facts have proved that this method can not only improve the prediction ability between network layers, but also ensure the performance of the large language model.
[0053] II. Model inference stage of the large language model based on the Transformer architecture:
[0054] After the model training stage of the above large language model based on the Transformer architecture is completed, a trained large language model is obtained. What this invention expects is that the trained large language model can accelerate decoding in the inference stage. As Figure 3 shown, during the process of decoding using the trained large language model, when the output data of a certain network layer already has the ability to be decoded, the trained large language model will abort the calculation of the subsequent network layers and directly perform decoding, achieving decoding acceleration by reducing unnecessary computational volume. That is to say, during the process of processing the current input data using the trained large language model, whenever the output data of the network layer in the trained large language model is obtained, it is judged whether the real-time obtained output data of the network layer meets the preset conditions until the network layer whose judgment result is yes is determined. The prediction result is determined according to the output data of this network layer, and the prediction result is combined with the current input data and re-input into the trained large language model. The specific implementation process is as follows:
[0055] S103. In the inference stage, the present invention needs to predict the content of the next token according to the given input data. Taking the input data "I" as an example, the input data of the trained large language model is "I", and the predicted output data needs to be "love". Similar to the above training stage, the trained large language model needs to perform layer-by-layer calculations. The output data of the first network layer is used as the input data of the second network layer, the output data of the second network layer is used as the input data of the third network layer, and so on. To achieve the goal of "aborting the calculations of subsequent network layers", the present invention sets a threshold λ and calculates a confidence score S based on the output data of the network layer. Then the preset condition is: S > λ. When S > λ is met, the judgment result is yes; when S > λ is not met, the judgment result is no.
[0056] Among them, the process of obtaining the confidence score S is as follows:
[0057] Whenever the output data of the network layer in the trained large language model is obtained, the softmax function is used to calculate the output data of this network layer to obtain the probability distribution result of each vocabulary in the vocabulary table. The difference between the maximum value and the second maximum value in the probability distribution result is taken as the confidence of this network layer. The reason is that if starting from a certain intermediate network layer, the trained large language model has an absolute confidence advantage in the output of a certain vocabulary, that is, the difference between the maximum value and the second maximum value is very large, then the vocabulary corresponding to the maximum value can be directly used as the prediction result to end the subsequent calculation process in advance.
[0058] Then the prediction result is combined with the current input data and re-input into the trained large language model. For example, if the prediction result is "love", after combining the input data "I" and the prediction result "love", it is used as the new input data to be input into the trained large language model to calculate and obtain the next prediction result until completion.
[0059] The present invention does not adjust the architecture of the large language model. Theoretically, it is applicable to any large language model based on the Transformer architecture. Therefore, it is applicable to generative scenarios in fields such as NLP or CV, and is also applicable to large language models based on the Seq2Seq architecture or pure decoder architecture, with strong generalization ability; since the model architecture is not adjusted, only necessary calculations are performed on the loss function and the output data of the network layer, so it has strong flexibility and can cooperate with architecture modification schemes such as model compression, distillation, and quantization to further accelerate the decoding process of the large language model.
[0060] In another embodiment, it includes:
[0061] 1) Training stage:
[0062] ① During the training phase, the present invention first has some input sentences. For example, in Figure 2 and Figure 3 such as "I love the blooming flowers in spring", since the model is an autoregressive model, these sentences can be used as both inputs and misaligned outputs.
[0063] ② Sequentially use the tokens in the input sentences as input data, and calculate the output of each network layer in the Transformer architecture layer by layer. For example, Figure 2 and Figure 3 both include 9 network layers. Taking the token "I" as an example, after passing through each network layer (block) in the Transformer architecture, the output data corresponding to that block will be obtained.
[0064] ③ In traditional model training, the present invention only uses the output data of the topmost network layer, and based on this output data, calculates the softmax function, and outputs the vocabulary at the position of the maximum probability index in the vocabulary as the prediction result (output), while the output of the intermediate network layer is only used as the input data for the next network layer for data transmission. In the model training of the present invention, for the output of each block of the Transformer architecture, the present invention will collect them all.
[0065] ④ Different from the traditional method of calculating loss based on the top-layer output, here the present invention performs weighted calculation on the collected output data of 9 network layers, which is specifically implemented using a modified loss function. The purpose is to encourage the large language model to generate representations meaningful to the prediction result between network layers. Therefore, the modified loss function proposed by the present invention no longer only uses the output data of the last network layer for prediction, but uses the weighted average of the predictions of the network layers for prediction. Moreover, it is more inclined to make the weights of the later network layers larger, so the expression ability of the later network layers is relatively stronger. Facts have proved that this method can not only improve the prediction ability between network layers, but also ensure the performance of the large language model.
[0066] 2) Inference phase:
[0067] ① In the inference phase, the present invention needs to predict the content of the next token according to the given input data. Taking the input data "I" as an example for illustration, the input data of the trained large language model is "I", and the predicted output data needs to be "love". The same as the above training phase, the trained large language model needs to calculate layer by layer. The output data of the first network layer is used as the input data of the second network layer, the output data of the second network layer is used as the input data of the third network layer, and so on.
[0068] ②Now the problem is how to decide to terminate prematurely. For this, the present invention needs to first define a confidence measure method. Specifically, the present invention sets a threshold λ, and calculates a confidence score S based on the output data of the network layer. Then the preset condition is: S > λ. When S > λ is met, the judgment result is yes; when S > λ is not met, the judgment result is no.
[0069] ③The present invention uses the output data calculated layer by layer by the softmax function, and takes the difference between the maximum value and the second maximum value in the softmax calculation result as the confidence, because:
[0070] If starting from a certain intermediate network layer, the trained large language model has an absolute confidence advantage in the output of a certain vocabulary, that is, the difference between the maximum value and the second maximum value is very large, then the vocabulary corresponding to the maximum value can be directly used as the prediction result to prematurely end the subsequent calculation process.
[0071] ④Once the confidence S of a certain network layer is > λ, the present invention directly calculates the softmax using the output result of this network layer, and performs sampling or applies it to other decoding strategies based on the calculation result.
[0072] ⑤Then take the original input "I" and the newly decoded generated new token "love" as the new input, recalculate and decode the next token, and recalculate the position of the premature termination layer.
[0073] The beneficial effects of the present invention are as follows:
[0074] 1) It improves the decoding process of the original ecological generative model (large language model input generative model). Through the premature prediction of the network layer, the calculation amount in the inference process can be reduced, thereby improving the decoding efficiency.
[0075] 2) It can save about half of the calculation time through the softmax method.
[0076] 3) It has very high versatility. For large language models based on the Transformers architecture, large language models based on the T5 Seq2Seq architecture, or large language models based on the pure decoder architecture of the GPT series, the present invention can be flexibly applied without being limited to a specific large language model.
[0077] 4) Combined with existing model compression technologies (such as distillation, quantization, pruning, etc.), the decoding speed of the generative model can be further improved.
[0078] In the above embodiments, although the steps are numbered as S1, S2, etc., these are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation, which is also within the protection scope of the present invention. It can be understood that in some embodiments, it may include some or all of the above embodiments.
[0079] As Figure 4 shown, a decoding acceleration system 200 of a large language model according to an embodiment of the present invention includes a weighting processing module 201 and a data processing module 202;
[0080] The weighting processing module 201 is configured to: during the training of the large language model, perform weighting processing on the output data of each network layer in the large language model;
[0081] The data processing module 202 is configured to: during the process of using the trained large language model to process the current input data, whenever the output data of the network layer in the trained large language model is obtained, determine whether the real-time obtained output data of the network layer meets the preset conditions until the network layer with a judgment result of yes is determined, determine the prediction result according to the output data of this network layer, combine the prediction result with the current input data, and re-enter it into the trained large language model.
[0082] Optionally, in the above technical solution, the weighting processing module 201 is specifically configured to:
[0083] During the training of the large language model, perform weighting processing on the output data of each network layer in the large language model by using the modified loss function.
[0084] Optionally, in the above technical solution, when performing weighting processing on the output data of each network layer in the large language model by using the modified loss function, among every two adjacent network layers in the large language model, the weight of the output data of the previous network layer is higher than the weight of the output data of the subsequent network layer.
[0085] Optionally, in the above technical solution, the large language model is a large language model based on the Transformer architecture, a large language model based on the Seq2Seq architecture, or a large language model based on a pure decoder architecture.
[0086] It should be noted that the beneficial effects of the decoding acceleration system 200 of the large language model provided in the above embodiments are the same as those of the decoding acceleration method of the large language model, which will not be elaborated here. In addition, when the system provided in the above embodiments implements its functions, only the division of the above functional modules is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments, which will not be elaborated here.
[0087] Among them, the decoding acceleration system of the large language model of the present invention can be a computer program (including program code) running in a computer device. For example, the decoding acceleration system of the large language model of the present invention is an application software, which can be used to execute the corresponding steps in the decoding acceleration method of the large language model of the present invention.
[0088] In some embodiments, the decoding acceleration system of the large language model of the present invention can be implemented in a combination of software and hardware. As an example, the decoding acceleration system of the large language model of the present invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the decoding acceleration method of the large language model of the present invention. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.
[0089] Among them, the modules described in the embodiments of the present invention can be implemented by software or by hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases.
[0090] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the decoding acceleration method of any of the above large language models is implemented. That is to say, an electronic device according to an embodiment of the present invention may include, but is not limited to: a processor and a memory; the memory is used to store a computer program; the processor is used to execute the decoding acceleration method of the large language model shown in any embodiment of the present invention by calling the computer program.
[0091] In an alternative embodiment, an electronic device is provided, as Figure 5 shown, Figure 5 the electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 may be used for data interaction between this electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in actual applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present invention.
[0092] The processor 4001 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules and circuits described in connection with the disclosure of the present invention. The processor 4001 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0093] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect, peripheral component interconnect standard) bus or an EISA (Extended Industry Standard Architecture, extended industry standard structure) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5Only a thick line is used to represent the bus 4002 in the figure, but it does not mean that there is only one bus or one type of bus.
[0094] The memory 4003 can be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0095] The memory 4003 is used to store the application program code (computer program) for implementing the solution of the present invention, and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the application program code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0096] Among them, the electronic device can also be a terminal device, and the terminal device can be any device that can install applications, including at least one of a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart TV, and a smart vehicle-mounted device.
[0097] It should be noted that Figure 5 The shown electronic device is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.
[0098] A computer-readable storage medium according to an embodiment of the present invention, on which a computer program is stored, and when the computer program is executed by a processor, it implements the decoding acceleration method of any one of the above large language models.
[0099] Optionally, the computer-readable storage medium can be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0100] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions that are stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes any one of the above decoding acceleration methods for large language models.
[0101] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0102] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the methods and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0103] The computer-readable storage medium provided by the embodiments of the present invention may be, but is not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared ray, or semiconductor, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program may be used by or in combination with an instruction execution system, device or component.
[0104] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.
[0105] The above description is only a preferred embodiment of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the present invention.
[0106] It should be noted that the terms "first", "second", etc. in the description and claims of this application are used to distinguish similar objects, and do not represent a limitation on a specific order or sequence. In appropriate cases, the use order of similar objects may be interchanged so that the embodiments of this application described herein can be implemented in an order other than the illustrated or described order.
[0107] Those skilled in the art know that the present invention can be implemented as a system, method or computer program product. Therefore, the present invention can be specifically implemented in the following forms: it can be completely hardware, can also be completely software (including firmware, resident software, microcode, etc.), or can also be in the form of a combination of hardware and software, which is generally referred to as "circuit", "module" or "system" herein. In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, and the computer-readable media contains computer-readable program code.
[0108] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A decoding acceleration method for a large language model, characterized in that: include: In the process of training the large language model, weighting the output data of each network layer in the large language model; In the process of processing the current input data using the trained large language model, whenever the output data of the network layer in the trained large language model is obtained, it is judged whether the output data of the network layer obtained in real time meets the preset conditions, until the network layer with a judgment result of yes is determined, and the prediction result is determined according to the output data of the network layer, and the prediction result is combined with the current input data and re-input into the trained large language model.
2. The decoding acceleration method of a large language model according to claim 1, characterized in that: In the process of training the large language model, weighted processing is performed on the output data of each network layer in the large language model, including: In the process of training the large language model, the output data of each network layer in the large language model is weighted by using the transformed loss function.
3. The decoding acceleration method of a large language model according to claim 2, characterized in that: When the output data of each network layer in the large language model is weighted by using the modified loss function, in every two adjacent network layers in the large language model, the weight of the output data of the previous network layer is higher than the weight of the output data of the next network layer.
4. The decoding acceleration method of a large language model according to any one of claims 1 to 3, characterized in that: The large language model is a large language model based on a Transformer architecture, a large language model based on a Seq2Seq architecture, or a large language model based on a pure decoder architecture.
5. A decoding acceleration system for a large language model, characterized in that: It includes a weighted processing module and a data processing module; The weighted processing module is used to: perform weighted processing on the output data of each network layer in the large language model during the training of the large language model; The data processing module is used to: in the process of processing the current input data using the trained large language model, whenever the output data of the network layer in the trained large language model is obtained, determine whether the output data of the network layer obtained in real time meets the preset conditions, until determining the network layer with a judgment result of yes, determine the prediction result according to the output data of the network layer, combine the prediction result with the current input data, and re-input it into the trained large language model.
6. The decoding acceleration system for a large language model according to claim 5, characterized in that: The weighted processing module is specifically used for: In the process of training the large language model, the output data of each network layer in the large language model is weighted by using the transformed loss function.
7. The decoding acceleration system for a large language model according to claim 6, characterized in that: When the output data of each network layer in the large language model is weighted by using the modified loss function, in every two adjacent network layers in the large language model, the weight of the output data of the previous network layer is higher than the weight of the output data of the next network layer.
8. A large language model decoding acceleration system according to any one of claims 5 to 7, characterized in that: The large language model is a large language model based on a Transformer architecture, a large language model based on a Seq2Seq architecture, or a large language model based on a pure decoder architecture.
9. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the decoding acceleration method of a large language model as claimed in any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the decoding acceleration method for a large language model according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Generative model decoding method and device, equipment and medium
CN116911279A