Model calculation method and device
By dividing the large language model into multiple parts and running on multiple computing units, combining attention layer computing optimization and low-bit quantization technology, the problem that the large language model cannot run on a single GPU is solved, and the computing resource utilization and task execution efficiency are improved.
Patent Information
- Application Number
- CN202311668536.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-06
AI Technical Summary
Large language models cannot run on a single GPU due to the large order of parameters, resulting in low computing resource utilization and low task execution efficiency.
The target model is divided into multiple parts by module and deployed and run on multiple computing units, each part is calculated independently, and the computing efficiency is improved by optimizing the calculation process of the attention layer and low-bit quantization technology.
It realizes the operation of large language models on multiple computing units, improves computing resource utilization and task execution efficiency, and reduces data storage and transmission requirements.
Smart Images

Figure CN120104299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a model calculation method and device. Background Art
[0002] A large language model is a language model that is trained on a large-scale text corpus and contains tens of billions of parameters (or more). It is a series of artificial intelligence models designed to understand and generate human language. Large language models are trained on large amounts of text data to solve common language problems such as text classification, question answering, document summarization, and text generation.
[0003] With the advent of ChatGPT, the parameters of language models are getting larger and larger, such as 13B (13 billion), 30B, 70B, 130B, etc. Such a large language model cannot be run on a GPU (Graphics Processing Unit). Summary of the invention
[0004] The embodiments of the present invention provide a model calculation method and device, which can run a large language model on multiple computing units, and at the same time can greatly improve the computing resource utilization of the computing units and improve the efficiency of the large language model in executing target tasks.
[0005] In a first aspect, an embodiment of the present invention discloses a model calculation method, which is applied to an electronic device, wherein the electronic device includes at least two calculation units, and the method includes:
[0006] Divide the target model into N parts according to the modules and deploy them on the computing unit, where N is less than or equal to the number of computing units included in the electronic device; wherein at least one part is deployed on one computing unit, and each part includes at least one module;
[0007] Input the input sequence of the first target task into the target model so that each part of the target model uses the computing unit in which it is located to calculate the input it receives; wherein the input received by the latter part includes the output of the previous part, and the input received by the first part includes the input sequence of the first target task.
[0008] In a second aspect, an embodiment of the present invention discloses a model calculation device, which is applied to an electronic device, wherein the electronic device includes at least two calculation units, and the device includes:
[0009] A partitioning and deployment module, used to divide the target model into N parts according to the modules and deploy them on the computing unit, where N is less than or equal to the number of computing units included in the electronic device; wherein at least one part is deployed on one computing unit, and each part includes at least one module;
[0010] A model calculation module is used to input the input sequence of the first target task into the target model, so that each part of the target model uses the calculation unit in which it is located to calculate the input it receives; wherein the input received by the latter part includes the output of the previous part, and the input received by the first part includes the input sequence of the first target task.
[0011] The embodiments of the present invention include the following advantages:
[0012] The embodiment of the present invention divides the target model into multiple parts and deploys them on multiple different computing units to run independently, thereby realizing the internal parallelization of the target model. When the scale of the target model is large, the target model can be run on multiple computing units, and at the same time, the computing resource utilization rate of the computing unit can be greatly improved, and the efficiency of the target model in performing the target task can be improved. In addition, the embodiment of the present invention performs low-bit quantization on the target data, which can greatly reduce the storage resources required for the target data and increase the data transmission volume, thereby increasing the number of requests that the target model can process at one time. Furthermore, the embodiment of the present invention optimizes the calculation process of the attention layer, and the optimized calculation process can reduce the number of data reads and writes, save data bandwidth, and thus improve the calculation efficiency of the target model. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.
[0014] Figure 1 is a flow chart of steps of an embodiment of a model calculation method of the present invention;
[0015] Figure 2 It is a schematic diagram of dividing and deploying a target model on multiple GPUs according to the present invention;
[0016] Figure 3 is another schematic diagram of dividing and deploying a target model on multiple GPUs according to the present invention;
[0017] Figure 4 yes Figure 2 Schematic diagram of the execution time of the deployment method;
[0018] Figure 5 yes Figure 3 Schematic diagram of the execution time of the deployment method;
[0019] Figure 6 It is a schematic flow chart of the calculation method 1 of the present invention;
[0020] Figure 7 This is a schematic diagram of the reading operation of the attention layer in one calculation process in calculation method 1;
[0021] Figure 8 It is a schematic flow chart of the second calculation method of the present invention;
[0022] Fig. 9 It is a structural block diagram of an embodiment of a model calculation device of the present invention. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] The terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable when appropriate, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, the term "and / or" in the specification and claims is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. In the embodiments of the present invention, the term "multiple" refers to two or more, and other quantifiers are similar.
[0025] In order to facilitate understanding of the technical solution of the present invention, some technical terms involved in the present invention are introduced below.
[0026] Large language models are one of the applications of deep learning, especially in the field of natural language processing (NLP). Large language models need to be trained on a large amount of text data to learn various patterns and structures of language. ChatGPT is an example of a large language model that is trained to understand and generate human language in order to have effective conversations and answer various questions.
[0027] Large language models are trained to solve general (multi-domain) language problems such as text classification, question answering, document summarization, and text generation.
[0028] (1) Text classification: Large language models can analyze and learn from input text to classify it into one or more predefined categories. For example, large language models can be used to classify whether an email is spam or whether a tweet is positive, negative, or neutral.
[0029] (2) Question answering: Large language models can answer natural language questions asked by users. For example, large language models can be used to answer user queries in search engines or to answer user questions in intelligent assistants.
[0030] (3) Document summarization: Large language models can automatically extract the main information in a text to generate a document summary or excerpt. For example, large language models can be used to generate a summary of a news article or to extract key plot points and events from a novel.
[0031] (4) Text generation: Large language models can use previously learned patterns and structures to generate new text. For example, large language models can be used to generate poems, short stories, or articles on a specific topic.
[0032] In specific implementations, a general large language model can be pre-trained first, and then fine-tuned for specific goals to obtain a single model applied to specific tasks or fields.
[0033] In the pre-training phase, large-scale general text data is used for training so that the model can learn the basic structure of the language and various common sense. Then, in the fine-tuning phase, a smaller and more specific data set is used for further training. The data set in the fine-tuning phase is usually targeted at a specific task or field, such as medical text, legal text, or specific conversation data. Fine-tuning allows the model to better understand and generate the language of this specific field, so as to better complete specific tasks.
[0034] Although large language models require a large amount of general text data in the pre-training phase, only relatively small domain-specific data is needed in the fine-tuning phase. This is because the model has learned a lot of language knowledge and common sense in the pre-training phase, and the fine-tuning phase is mainly to adapt the model to a specific task or domain. This enables large language models to perform well in data-scarce domains, which can greatly reduce the complexity and cost of developing and maintaining different models.
[0035] The Transformer model is a deep learning model widely used in the field of natural language processing (NLP). The main feature of the Transformer model is the use of an attention mechanism, which allows the model to take into account the contextual relationships of all elements in the sequence when processing sequence data.
[0036] The Transformer model mainly consists of two parts: encoder and decoder.
[0037] Encoder: The encoder consists of multiple identical modules (Blocks), each of which consists of two layers: the first layer is the attention layer, which can take into account the contextual relationship of all elements in the input sequence; the second layer is the feed forward neural network. Each layer is followed by a residual connection and layer normalization. The task of the encoder is to convert the input sequence into a set of continuous representations that take into account the context of each element in the input sequence.
[0038] Decoder: The decoder is also composed of multiple identical blocks, each of which includes the following three layers: The first layer is the attention layer, which only considers the element and the elements before it when processing the current element, and does not consider the elements after it. This mechanism is called masked attention. The second layer is the encoder-decoder attention layer, which enables the decoder to pay attention to the output of the encoder. The third layer is a feedforward neural network. Each layer is followed by a residual connection and layer normalization. The task of the decoder is to generate the output of the next moment based on the output of the encoder and the output of the decoder at the previous moment.
[0039] Reference Figure 1 , shows a flow chart of steps of an embodiment of a model calculation method of the present invention, the method can be applied to an electronic device, the electronic device includes at least two calculation units, the method may include the following steps:
[0040] Step 101: Divide the target model into N parts according to modules and deploy them on the computing unit, where N is less than or equal to the number of computing units included in the electronic device; wherein at least one part is deployed on one computing unit, and each part includes at least one module;
[0041] Step 102: input an input sequence of the first target task into the target model, so that each part of the target model calculates the input it receives using its computing unit; wherein the input received by the latter part includes the output of the former part, and the input received by the first part includes the input sequence of the first target task.
[0042] The model calculation method provided by the present invention can be applied to electronic devices, and the electronic devices include at least two computing units. The electronic devices may include terminal devices such as smart phones, desktop computers, tablet computers, laptop computers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, intelligent voice interaction devices, smart home appliances, smart wearable devices, and vehicle-mounted terminal devices. The electronic device may also be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0043] The computing unit may be a GPU or an NPU (Neural Network Processing Unit). The embodiment of the present invention does not limit this. For ease of description, the embodiment of the present invention takes the case where the computing unit may be a GPU as an example, that is, the electronic device running the target model includes at least two GPUs.
[0044] The target model may be a large language model, the target model may be a general large language model obtained through pre-training, or the target model may be a single model applied to a specific task or field obtained by fine-tuning the general large language model.
[0045] The target model is composed of M modules (Blocks), M is greater than or equal to 2, and the M blocks are connected in series. The target model can perform target tasks, including but not limited to any one of tasks such as text classification, question answering, document summarization, and text generation.
[0046] When a large language model cannot be run on one GPU, each block of a large language model can be evenly divided into N parts and deployed on N GPUs for calculation. Assuming that a large language model consists of M blocks connected in series, after the N parts of the first block are calculated, the calculation results of these N parts are summarized and sent to the N parts of the second block for calculation. And so on, until the N parts of the Mth block are calculated.
[0047] Reference Figure 2 , which shows a schematic diagram of dividing the target model and deploying it on multiple GPUs. Figure 2 As shown: Step 1, divide each Block into N parts and deploy them on N GPUs. Figure 2 As shown in (a), assuming that the target model contains M blocks (such as Block to BlockM), the electronic device contains N GPUs, and assuming that N is 3 (such as GPU1, GPU2, and GPU3), each Block can be divided into 3 parts. For example, Block1 is divided into the following three parts: 1-1, 1-2, and 1-3; Block2 is divided into the following three parts: 2-1, 2-2, and 2-3; and so on. Step 2, copy the input of the first Block (such as Block1) into N copies and distribute them to the N GPUs respectively, so that the N parts of the first Block can calculate the received input separately, such as Figure 2 As shown in (b) and (c) in the figure. In step 3, after the N parts of the first Block are calculated, the calculation results of these N parts are combined on a certain GPU, such as the first GPU (such as GPU1). After the calculation results of these N parts are sorted out, the first GPU uses them as the input of the second Block (such as Block2). Repeat step 2 until the calculation of the Mth Block is completed, and the output result of the target model is obtained, as shown in Figure 2 As shown in (c) and (d).
[0048] use Figure 2 Although the deployment method shown in the figure can run the target model on multiple GPUs, a large language model is usually composed of a large number of blocks in series. Therefore, when the target model executes a target task, it needs to wait until all N parts of the last block are calculated before the first block can accept new calculation requests, causing the previously completed blocks to be idle, thus affecting the utilization of the GPU. In addition, if the utilization rate of a block on a GPU is originally 100%, but Figure 2In the deployment method shown, each Block is divided into N parts and deployed on N GPUs, so the utilization rate of one Block on one GPU becomes 100% / N, further affecting the utilization rate of the GPU and causing a waste of GPU computing resources.
[0049] In order to solve the above problems, the embodiments of the present invention are Figure 2 The deployment method shown in the figure is improved, and the target model is divided into N parts according to the modules and deployed on the computing unit, where N is less than or equal to the number of computing units included in the electronic device; wherein at least one part is deployed on one computing unit, and each part includes at least one module. Figure 3 , shows another schematic diagram of dividing the target model and deploying it on multiple GPUs. Figure 3 As shown. The M blocks of the target model are divided into N parts to obtain part 1 to part N. For example, a target model has 6 blocks and the electronic device has 3 GPUs. The target model can be divided into 3 parts and deployed on 3 GPUs, with one part deployed on each GPU, and each part containing 2 blocks. For another example, a target model has 7 blocks and the electronic device has 3 GPUs. The target model can be divided into 3 parts and deployed on 3 GPUs, with one part deployed on each GPU; part 1 contains 3 blocks, part 2 contains 2 blocks, and part 3 contains 2 blocks.
[0050] In an optional embodiment of the present invention, the target model includes M modules, M≥N, and dividing the target model into N parts according to the modules and deploying them on the computing unit includes:
[0051] If M is an integer multiple of N, the M modules of the target model are divided into N parts in equal proportion and deployed on the N computing units, with one part deployed on each computing unit, and each part containing M / N modules;
[0052] If M is not an integer multiple of N, the M modules of the target model are divided into N parts and deployed on the N computing units, with one part deployed on each computing unit, each of the N-1 parts containing M / N modules, and the Nth part containing M / N+M%N modules. M / N represents the quotient obtained by dividing M by N, and M%N represents the remainder obtained by dividing M by N.
[0053] For example, a target model has 7 blocks, and the electronic device has 3 GPUs (such as GPU1 to GPU3); the target model can be divided into three parts, part 1 contains 2 blocks, and part 1 is deployed on GPU1, part 2 contains 2 blocks, and part 2 is deployed on GPU2, and part 3 contains 3 blocks, and part 3 is deployed on GPU3; or, the target model can be divided into three parts, part 1 contains 3 blocks, and part 1 is deployed on GPU1, part 2 contains 2 blocks, and part 2 is deployed on GPU2, and part 3 contains 2 blocks, and part 3 is deployed on GPU3.
[0054] For another example, a target model has 5 blocks, and the electronic device has 5 GPUs (such as GPU1 to GPU5). The 5 blocks of the target model can be divided into 3 parts, part 1 contains 2 blocks, and part 1 is deployed on GPU1, part 2 contains 2 blocks, and part 2 is deployed on GPU2, and part 3 contains 1 block, and part 3 is deployed on GPU3. For another example, the 5 blocks of the target model can be divided into 5 parts, each part contains 1 block, and they are deployed on 5 GPUs respectively.
[0055] It should be noted that the above-mentioned method of dividing and deploying the target model on multiple GPUs is only for exemplary description, and the embodiment of the present invention does not limit the method of dividing and deploying the target model.
[0056] The embodiment of the present invention divides the target model into N parts according to modules, where N is less than or equal to the number of computing units included in the electronic device, and at least one part is deployed on one computing unit, and each part includes at least one module. The input of the first part of the N parts is the input sequence of the target task. The input of each part after the first part includes the output of the previous part. Specifically, the input sequence of the first target task is input into the first part of the N parts, so that each part of the target model uses the computing unit where it is located to calculate the received input and output the result; the result output by the last part includes the calculation result of the first target task.
[0057] The target model is composed of M blocks (such as Block1 to BlockM) connected in series. The calculation method of the target model is to input the input sequence of the target task into Block1. After Block1 completes the calculation, it outputs the result to Block2 for calculation, and so on, until BlockM completes the calculation and outputs the calculation result of the target task.
[0058] Assume that the electronic device has 3 GPUs, Figure 2In the deployment method shown, each block of the target model is evenly divided into N parts, which are deployed on GPU1, GPU2, and GPU3 respectively. Therefore, when calculating each block, GPU1, GPU2, and GPU3 need to be occupied at the same time. Only after the last block (BlockM) is calculated, can Block1 occupy GPU1, GPU2, and GPU3 again for new calculations.
[0059] Reference Figure 4 , showing Figure 2 Assume that at time T, the target model needs to be used to execute the first target task (such as the first target task is recorded as session1), then the input sequence of session1 is input into Block1 for calculation; at time T+1, the target model needs to be used to execute the second target task (such as session2); however, since all three GPUs are occupied at this time, the target model cannot receive new calculation requests, and the occupied three GPU computing resources can only be released after BlockM is calculated. That is, Figure 4 As shown in the figure, at time T+4, the target model can receive new computing requests, and Block1 can receive the input sequence of session2 for computing.
[0060] In an optional embodiment of the present invention, the method may further include:
[0061] At a target time, an input sequence of a second target task is input into the target model, so that the portion of the target model that has completed the calculation of the input sequence for the first target task starts to calculate the input sequence for the second target task; the target time is used to indicate that there is a portion of the target model that has completed the calculation of the input sequence for the first target task, and there is a portion that has not yet completed the calculation of the input sequence for the first target task.
[0062] for Figure 3In the deployment method shown, different GPUs are allowed to deploy different parts of the target model, and one part contains at least one Block. For example, part 1 (including Block1 and Block2) is deployed on GPU1, part 2 (including Block3 and Block4) is deployed on GPU2, and part 3 (including Block5 and Block6) is deployed on GPU3. GPU1 is only occupied by part 1, GPU2 is only occupied by part 2, and GPU3 is only occupied by part 3. In this way, when part 1 receives the input sequence of session1, part 1 uses the computing resources of GPU1 for calculation; when part 1 is calculated, the computing resources of GPU1 will be released, and part 1 can receive new calculation requests. The same applies to GPU2 and GPU3. Reference Figure 5 , showing Figure 3 The execution time diagram of the deployment method is as follows: Figure 5 As shown in the figure, Part1 represents part 1, Part2 represents part 2, and Part3 represents part 3. At time T+2 (i.e., the target time), Part1 has completed the calculation of the input sequence for the first target task. Although Part2 and Part3 have not yet completed the calculation of the input sequence for the first target task, Part1 has released the computing resources of GPU1 occupied by it. Therefore, the target model can start to receive new tasks at this time. Part1 can receive the input sequence of session2 and start calculation. The time of receiving new calculation requests is earlier than Figure 4 The T+4 moment.
[0063] In a specific implementation, the second target task can be sent to the target model later than the first target task, or it can be sent to the target model at the same time as the first target task. After the calculation of the first part of the target model for the first target task is completed, there is no need to wait for the calculation of the other parts to be completed, that is, the calculation of the input sequence of the second target task received can be started at the target time.
[0064] It should be noted that the first target task and the second target task are only used to distinguish two different target tasks. In the specific implementation, the target model can receive multiple target tasks at the same time. After the calculation of the first part for the current target task is completed, the calculation of the next target task can be started, and the same is true for other parts. In this way, the parallel calculation of multiple tasks can be realized to improve the efficiency of task execution.
[0065] The embodiment of the present invention divides the target model into N parts and deploys them on N computing units, so that each part runs independently on a different computing unit, thereby allowing parallel processing within the target model and significantly improving the computing resource utilization of the computing units.
[0066] Furthermore, in the embodiment of the present invention, the target model is divided into N parts according to modules and deployed on the computing unit, and a scheduler may be configured for each part, such as scheduler 1 to scheduler N. Each part and its corresponding scheduler are deployed on a GPU to become an independent service, and the N services can be used to execute the target task in parallel.
[0067] For example, when executing the first target task, the input sequence of the first target task is input to scheduler 1. Scheduler 1 arranges the input data of part 1 based on the received input sequence and inputs it to part 1 of the target model for calculation. Part 1 uses the computing resources of GPU1 to calculate the input from scheduler 1, and after the calculation is completed, the result is output to scheduler 2 in GPU2. Scheduler 2 arranges the input data of part 2 based on the received calculation result of part 1 and inputs it to part 2 for calculation. And so on.
[0068] In an optional embodiment of the present invention, at least one module of the target model includes an attention layer, which calculates a query matrix Q, a key matrix K and a value matrix V based on the received input, and calculates the output of the attention layer based on the calculated query matrix Q, key matrix K and value matrix V; wherein the input received by the attention layer includes a vector matrix corresponding to an input sequence of the first target task, or the input received by the attention layer includes the output of a module on the attention layer.
[0069] The target model can be a large language model, and the basic structure of the large language model can be a Transformer structure. The Transformer model is built with Block as the basic unit, and each Block contains a multi-layer self-attention mechanism and a feedforward neural network layer. This modular architecture makes the Transformer model easy to modify, expand and adjust, and multiple Blocks can be freely combined and stacked as needed.
[0070] The main feature of the Transformer model structure is the use of an attention mechanism, that is, at least one module of the target model includes an attention layer. Exemplarily, each module of the target model may include an attention layer. The attention layer calculates a query matrix Q, a key matrix K, and a value matrix V based on the received input, and calculates the output of the attention layer based on the calculated matrix Q, matrix K, and matrix V.
[0071] In an embodiment of the present invention, the target model includes M blocks, and the M blocks include encoder blocks and decoder blocks. Furthermore, each block may include an attention layer, and the input received by the attention layer may be a vector matrix X composed of the representation vectors x of each element in the input sequence of the target task, or the input received by the attention layer may be the output of the previous block. The elements in the input sequence may be words, characters, phrases, or other independent language units in the text.
[0072] In an example, assume that the first target task is a translation task, and the input sequence of the first target task is "I have a cat". The vector matrix X is a 4-row matrix, each row being the representation vector x of each element (that is, each word) in the input sequence. The representation vector x of an element can be obtained by adding or concatenating the word embedding and position embedding of the element. Word embedding represents the encoding of the meaning of the element, and position embedding represents the encoding of the position of the element in the input sequence.
[0073] The query matrix Q, key matrix K and value matrix V are obtained by the attention layer by performing matrix linear transformation on the received input. After obtaining the query matrix Q, key matrix K and value matrix V, the output of the attention layer can be calculated.
[0074] Query matrix Q: used to represent the query vector for each element in the input sequence. Specifically, the i-th row of the query matrix Q represents the query vector for the i-th element in the input sequence, which will be used to calculate the similarity score between other elements in the input sequence and the i-th element.
[0075] Key matrix K: used to represent the key vector of each element in the input sequence. Specifically, the i-th row of the key matrix K represents the key vector of the i-th element in the input sequence, which will be used to calculate the similarity score between the i-th element in the input sequence and other elements.
[0076] Value matrix V: A value vector used to represent each element in the input sequence. Specifically, the i-th row of the value matrix V represents the value vector of the i-th element in the input sequence, which will be used to calculate the contribution of other elements in the input sequence to the i-th element.
[0077] The input of the attention layer is represented by a vector matrix X, and the query matrix Q, key matrix K, and value matrix V can be calculated using linear matrix WQ, WK, and WV. Specifically, the vector matrix X is multiplied by the linear matrix WQ to obtain the query matrix Q; the vector matrix X is multiplied by the linear matrix WK to obtain the key matrix K; the vector matrix X is multiplied by the linear matrix WV to obtain the value matrix V. It should be noted that each row of the vector matrix X, the query matrix Q, the key matrix K, and the value matrix V corresponds to an element in the input sequence. Among them, WQ, WK, and WV are appropriate parameters learned during the training process of the target model.
[0078] For the Encoder, the input of the first Block is the vector matrix X composed of the representation vectors x of each element in the input sequence. The input of each subsequent Block is the output of the previous Block. The output of the last Block is called the encoding information matrix C, which will be used in the Decoder later.
[0079] It should be noted that each attention layer in the Transformer can adopt a multi-head attention mechanism, called a multi-head attention layer. The multi-head attention mechanism fuses several identical attention calculations to make the attention calculation have a more powerful resolution ability. That is, the first layer of each Block in the Encoder can be a multi-head attention layer.
[0080] In an example, the Block of the Encoder may include: a multi-head attention layer (Multi-HeadAttention layer), a residual connection layer (Add&Norm layer), a feedforward layer (Feed Forward layer), and a residual connection layer in sequence. The residual connection layer consists of two parts, Add and Norm. For the residual connection layer connected after the multi-head attention layer, Add means: X+MultiHeadAttention(X), X represents the input of the multi-head attention layer, and MultiHeadAttention(X) represents the output of the multi-head attention layer. For the residual connection layer connected after the feedforward layer, Add means: X+Feed Forward(X), X represents the input of the feedforward layer, and Feed Forward(X) represents the output of the feedforward layer. Add is a residual connection, which is usually used to solve the problem of multi-layer network training. Norm refers to the normalization operation, which is used to speed up convergence. Multiple blocks of the above structure can be superimposed to form an Encoder.
[0081] The encoder is used to receive the vector matrix X corresponding to the input sequence and output the encoding information matrix C. The input received by the first Block in the encoder is the vector matrix X corresponding to the input sequence, and the input received by each subsequent Block is the output of the previous Block. The first Block in the encoder multiplies its received input (the vector matrix X corresponding to the input sequence) with the linear array matrix WQ1, WK1 and WV1 respectively to obtain the query matrix Q1, the key matrix K1 and the value matrix V1, and calculates its output based on the query matrix Q1, the key matrix K1 and the value matrix V1. The second Block in the encoder multiplies its received input (the output of the first Block) with the linear array matrix WQ2, WK2 and WV2 respectively to obtain the query matrix Q2, the key matrix K2 and the value matrix V2, and calculates its output based on the query matrix Q2, the key matrix K2 and the value matrix V2. By analogy, the output of the last Block in the encoder is the encoding information matrix C.
[0082] For the Block in the Decoder, the structure is similar to that of the Block in the Encoder, except that the Block in the Decoder contains two multi-head attention layers and a Softmax layer at the end. The first multi-head attention layer uses a mask operation. The second multi-head attention layer (also called the encoder-decoder attention layer) uses the encoded information matrix C output by the Encoder to calculate the key matrix K and the value matrix V, and uses the output of the previous Block to calculate the query matrix Q. The Softmax layer is used to calculate the probability of the output at the next moment. It should be noted that the first Block in the Decoder uses the output of the previous moment to calculate the query matrix Q, and uses the complete output sequence to calculate the key matrix K and the value matrix V. The subsequent calculation method is consistent with the previous description.
[0083] In an optional embodiment of the present invention, after the attention layer calculates the query matrix, the key matrix and the value matrix according to the received input, the method may further include:
[0084] quantizing the target data of the first bit width according to the second bit width and storing it in a preset area; the target data includes at least one of the query matrix, the key matrix and the value matrix;
[0085] The calculating the output of the attention layer according to the calculated query matrix, key matrix and value matrix may include:
[0086] After reading the target data of the second bit width from the preset area and converting it into the first bit width, the output of the attention layer is calculated based on the query matrix, key matrix and value matrix of the first bit width.
[0087] The target model can be a Transformer model. Each Block of the model can include an attention layer. Each Block needs to calculate the query matrix Q, key matrix K and value matrix V according to the received input. These query matrices Q, key matrices K and value matrices V may be used in subsequent calculations, so they need to be stored. In order not to exceed the storage space limit, the calculation request processed at one time needs to be based on the storage space. Therefore, the storage of the query matrix Q, key matrix K and value matrix V becomes a bottleneck for the entire service; in addition, if the historical query matrix Q, key matrix K and value matrix V need to be used in subsequent calculations, the historical query matrix Q, key matrix K and value matrix V need to be moved from the storage space to the computing unit, and the data handling capacity of the computing unit also becomes a bottleneck for the entire service.
[0088] To solve the above problems, in an embodiment of the present invention, the target data of the first bit width is quantized according to the second bit width and stored in a preset area, so that in subsequent calculations, the target data of the second bit width is read from the preset area and converted into the first bit width for calculation; the preset area is a specified storage space; the target data may include at least one of the query matrix Q, the key matrix K and the value matrix V. The first bit width is greater than the second bit width. Exemplarily, the first bit width is FP32, and the second bit width is INT8. When the target data stored in the preset area is needed to be used later, the target data is first read from the target area in INT8, and then the read INT8 target data is converted into FP32 for calculation. As a result, the storage space occupied by the target data stored in the preset area can be reduced to 1 / 4 of the original, and 4 times the original data can be transmitted. Under the condition of the same data bandwidth, the number of requests that the target model can process at one time is also increased to 4 times the original. Not only can the storage space occupied by the target data be reduced, but also the data handling capacity of the computing unit can be improved. It should be noted that the embodiment of the present invention does not limit the first bit width and the second bit width.
[0089] The embodiments of the present invention can provide the following two methods for calculating the attention layer.
[0090] Calculation method 1, calculating the output of the attention layer according to the calculated query matrix, key matrix and value matrix, may include:
[0091] Step S11, performing inner product calculation on each row in the query matrix and the key matrix to obtain a first intermediate result;
[0092] Step S12, performing e-th power calculation on each element in the first intermediate result to obtain a first transformation result;
[0093] Step S13, calculating a first proportion for each element in the first transformation result to obtain a second intermediate result;
[0094] Step S14: multiply each element in the second intermediate result by each row in the value matrix and then sum them up to obtain the output of the attention layer.
[0095] Specifically, for the attention layer in a certain Block of the target model, the above steps S11 to S14 can be executed to implement the calculation process of the attention layer.
[0096] Reference Figure 6 , which shows a flow chart of calculation method one.
[0097] In step S11, the first intermediate result is: S=Q×K, and each element in the first intermediate result S corresponds to a score s obtained by calculating the inner product of the query matrix Q and each row in the key matrix K.
[0098] In step S12, each element s in S is raised to the power of e, and s'=e is obtained. s The matrix composed of the results of calculating all the elements in S is called the first transformation result S'. The first element in the first transformation result S' is s 1 ', The second element is s 2 ', And so on.
[0099] In step S13, the first proportion of the i-th element is: p=s i ' / ∑s i The matrix composed of the first proportions of all elements in the first transformation result S' is called the second intermediate result P.
[0100] In step S14, the output of the attention layer is calculated as: O = ∑V × P.
[0101] It can be seen that calculation method 1 needs to generate a first intermediate result S and a second intermediate result P, and the sizes of the first intermediate result S and the second intermediate result P are both L Q ×L K , L Q represents the number of vectors in the query matrix Q, L K represents the number of vectors in the key matrix K, L Q ×L K It is difficult to store matrices of this size in the computing unit. Therefore, each time the first intermediate result S and the second intermediate result P need to be stored in a storage unit outside the computing unit. When the next calculation is performed, they are read from the storage unit to the computing unit for calculation. Figure 7, which shows a schematic diagram of the read operation of the attention layer in a calculation process in calculation method 1. Among them, matmul represents matrix multiplication, and softmax represents activation function processing. Figure 7 As shown, calculation method 1 needs to read Q and K from the storage unit to perform matmul calculation to generate a first intermediate result S and write it into the storage unit (corresponding to step S11), then read S from the storage unit to perform softmax activation function processing to generate a second intermediate result P and write it into the storage unit (corresponding to steps S12 and S13), and finally read V and P from the storage unit to obtain the final output O and write the output O into the storage unit (corresponding to step S14).
[0102] Calculation method 2, calculating the output of the attention layer according to the calculated query matrix, key matrix and value matrix, may include:
[0103] Step S21, performing inner product calculation on the query matrix and the i-th row in the key matrix to obtain the i-th score;
[0104] Step S22, performing e-th power calculation on the i-th score to obtain a second transformation result;
[0105] Step S23, adding the second transformation result to the accumulated value of the second transformation result of the previous round to obtain the accumulated value of the second transformation result of the current round; at the same time, multiplying the second transformation result by the i-th row in the value matrix to obtain a third transformation result, and adding the third transformation result to the accumulated value of the third transformation result of the previous round to obtain the accumulated value of the third transformation result of the current round;
[0106] Step S24, repeat the above steps S21 to S23 until the i-th row is the last row in the key matrix. After calculating the accumulated value of the third transformation result of the current round, calculate the ratio of the accumulated value of the third transformation result of the current round to the accumulated value of the second transformation result of the current round to obtain the output of the attention layer.
[0107] Calculation method 2 optimizes the calculation process of calculation method 1, so that the intermediate results generated in the calculation process are smaller and can be directly stored in the current calculation unit without the need for additional storage units.
[0108] Reference Figure 8 , which shows a flow chart of calculation method 2.
[0109] Specifically, in step S21, instead of calculating the inner product of each row in the query matrix Q and the key matrix K, the inner product of the query matrix Q and the i-th row in the key matrix K is calculated to obtain the i-th score S i , S i =Q×K i, i=1~k, k is the last row of the key matrix K.
[0110] In step S22, the i-th score is raised to the power of e to obtain a second transformation result:
[0111] In step S23, the second transformation result S i ' and the accumulated value S of the second transformation result of the previous round * Accumulate and get the accumulated value of the second transformation result of the current round: S * =S * +S i '. In the embodiment of the present invention, the second transformation result of the current round is accumulated to value S * It is called the third intermediate result.
[0112] At the same time, in step S23, the second transformation result S i ′ is multiplied by the i-th row in the value matrix V to obtain the third transformation result: P i =V i ×S i ′, and the third transformation result P i The accumulated value P of the third transformation result of the previous round * Accumulate and get the accumulated value of the third transformation result of the current round: P * =P * +P i In the embodiment of the present invention, the third transformation result of the current round is accumulated to value P * It is called the fourth intermediate result.
[0113] Steps S21 to S23 are executed repeatedly until the last row in the query matrix Q and the key matrix K is calculated, and the accumulated value of the third transformation result of the current round is obtained. At this time, the output of the attention layer can be calculated as: O = P * / S * .
[0114] It can be seen that Therefore, the output of the attention layer using calculation method 1 and calculation method 2 is the same.
[0115] In the calculation method 2, a third intermediate result S is generated during the calculation process. * and the fourth intermediate result P * The third intermediate result S * and the fourth intermediate result P * The size is 1×L V , 1×L V The size is much smaller than L Q ×L K Therefore, the third intermediate result S *and the fourth intermediate result P * It can be directly stored in the current calculation unit, and can be directly read from the calculation unit for calculation in the next step. Figure 7 As shown, the second calculation method only needs to read Q, K and V from the storage unit and write the output O to the storage unit, which reduces the process of reading and writing the first intermediate result S and the second intermediate result P. Therefore, compared with the first calculation method, the second calculation method can reduce the reading and writing consumption, thereby speeding up the calculation process.
[0116] In an example, assuming that the target model includes 6 blocks (Block1 to Block6), the electronic device includes 3 GPUs (GPU1 to GPU3), the target model is divided into 3 parts (part 1 to part 3), each part includes 2 blocks, such as part 1 includes Block1 and Block2, part 2 includes Block3 and Block4, and part 3 includes Block5 and Block6. Part 1 is deployed on GPU1, part 2 is deployed on GPU2, and part 3 is deployed on GPU3.
[0117] Assume that the first target task is a text generation task, and the input sequence of the first target task is "write a weekend team-building plan for 20 people in Hangzhou, starting from Beijing, and having fun", and the target model needs to be used to automatically generate text content based on the input sequence.
[0118] It should be noted that for text generation tasks, the self-attention calculation of the target model can include the following two stages: the first stage is called the prompt stage: used to parse user questions and generate the first output; the second stage is called the generate stage: used to generate new outputs word by word based on the generated output sequence. For example, in the prompt stage, the target model generates the first output as the word "Hang". In the generate stage, based on the generated output sequence "Hang", a new output is generated as the word "Zhou", and the generate stage is continued. For example, after the generate stage is executed for a period of time, the generated output sequence is "Hangzhou Weekend Team Building Plan: 1. Activity Background and Purpose". At the next moment, based on the generated output sequence "Hangzhou Weekend Team Building Plan: 1. Activity Background and Purpose", a new output is generated as the word "Biao", and the generate stage is continued.
[0119] Specifically, first, the input sequence "write a weekend team building plan for 20 people in Hangzhou, starting from Beijing, and having fun" needs to be converted into a vector matrix X, where X is composed of the representation vector x corresponding to each element in the input sequence, and x is the word embedding and position embedding of the element added or concatenated. This step can be performed by the embedding layer of the target model, or by other special conversion models, and the embodiment of the present invention does not limit it.
[0120] Input the input sequence of the first target task to the target model. Assuming that the input sequence has been converted into a vector matrix X, the vector matrix X first enters Block1 in Part 1 for calculation. Block1 is a Block in the Encoder. Block1 multiplies the received vector matrix X with the linear transformation matrices WQ1, WK1, and WV1 respectively to obtain the query matrix Q1, the key matrix K1, and the value matrix V1. At this time, the size of Q1 is [26, D], where D represents the vector dimension. The size of K1 is [26, D], and the size of V1 is [26, D]. Furthermore, at this time, Q1, K1, and V1 can be the same. Block1 executes steps S21 to S24, and based on the query matrix Q1, the key matrix K1, and the value matrix V1, calculates its output O1 and inputs it to Block2.
[0121] Block2 multiplies the received input (output O1 of Block1) with the linear matrix WQ2, WK2 and WV2 respectively to obtain the query matrix Q2, key matrix K2 and value matrix V2. Block2 executes steps S21 to S24 to calculate its output O2 based on the query matrix Q2, key matrix K2 and value matrix V2. Similarly, the output of the last Block of the Encoder (assuming it is Block3) is the encoding information matrix C corresponding to the input sequence "Write a weekend team building plan for 20 people in Hangzhou, starting from Beijing, and have fun".
[0122] Next, enter the decoding stage. The input to each block of the Decoder includes the encoded information matrix C output by the Encoder. Additionally, the input to the first block of the Decoder (assumed to be Block4) also includes the output of the previous moment and the complete output sequence. The calculation process of the blocks in the Decoder is similar to that in the Encoder, which will not be elaborated here. In the prompt stage, the last block of the Decoder (assumed to be Block6) generates the first output, which is the character "杭". Enter the generate stage. The input to Block4 includes the output "杭" of the previous moment and the complete output sequence "Write a weekend team-building plan for 20 people in Hangzhou, starting from Beijing, and having fun杭". Block4 multiplies the vector matrix corresponding to the output "杭" of the previous moment by the linear transformation matrix WQ4 to obtain the query matrix Q4, and the size of Q4 is [1, D]. Block4 multiplies the vector matrix corresponding to the complete output sequence "Write a weekend team-building plan for 20 people in Hangzhou, starting from Beijing, and having fun杭" by the linear transformation matrix WK4 to obtain the key matrix K4, and the size of K4 is [27, D]. The value matrix V4 is the same as K4, with a size of [27, D]. Block4 executes steps S21 to S24, and based on the query matrix Q4, the key matrix K4, and the value matrix V4, calculates its output O4 and inputs it to the next block (assumed to be Block5) for calculation. Block5 uses the output O4 of Block4 to calculate the query matrix Q5, and uses the encoded information matrix C output by the Encoder to calculate the key matrix K5 and the value matrix V5. Block5 executes steps S21 to S24, and based on the query matrix Q5, the key matrix K5, and the value matrix V5, calculates its output O5 and inputs it to the next block (assumed to be Block6) for calculation. The calculation process of Block6 is similar to that of Block5, and Block6 generates a new output, which is the character "州". Continue to execute the generate stage.
[0123] Assume that the generate stage is executed to a certain moment, and the generated output sequence is "Write a weekend team building plan for 20 people in Hangzhou, starting from Beijing, and having fun on the weekend team building plan in Hangzhou". At the next moment, the input of Block4 includes the output "square" of the previous moment and the complete output sequence "Write a weekend team building plan for 20 people in Hangzhou, starting from Beijing, and having fun on the weekend team building plan in Hangzhou". At this time, Block4 multiplies the vector matrix corresponding to the output "square" of the previous moment with the linear transformation matrix WQ4 to obtain the query matrix Q4, and the size of Q4 is [1,D]; and multiplies the vector matrix corresponding to the complete output sequence "Write a weekend team building plan for 20 people in Hangzhou, starting from Beijing, and having fun on the weekend team building plan in Hangzhou" with the linear transformation matrix WK4 to obtain the key matrix K4, and the size of K4 is [34,D]. The value matrix V4 is the same as K4, and the size is [34,D]. Continue to execute the generate stage until the generation is completed or the task is interrupted.
[0124] It should be noted that the linear transformation matrices WQ, WK, and WV used by each Block in the target model may be the same or different. The linear transformation matrices used by each Block may be model parameters trained during the training process of the target model.
[0125] In summary, the embodiment of the present invention divides the target model into multiple parts and deploys them on multiple different computing units to run independently, thereby realizing the internal parallelization of the target model. When the scale of the target model is large, the target model can be run on multiple computing units, and at the same time, the computing resource utilization rate of the computing unit can be greatly improved, and the efficiency of the target model in performing the target task can be improved. In addition, the embodiment of the present invention performs low-bit quantization on the target data, which can greatly reduce the storage resources required for the target data and increase the data transmission volume, thereby increasing the number of requests that the target model can process at one time. Furthermore, the embodiment of the present invention optimizes the calculation process of the attention layer, and the optimized calculation process can reduce the number of data reads and writes, save data bandwidth, and thus improve the calculation efficiency of the target model.
[0126] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0127] Reference Fig. 9, shows a structural block diagram of an embodiment of a model calculation device of the present invention, which is applied to an electronic device, wherein the electronic device includes at least two calculation units, and the device includes:
[0128] A partitioning and deployment module 901 is used to divide the target model into N parts according to the modules and deploy them on the computing unit, where N is less than or equal to the number of computing units included in the electronic device; wherein at least one part is deployed on one computing unit, and each part includes at least one module;
[0129] The model calculation module 902 is used to input the input sequence of the first target task into the target model, so that each part of the target model uses the calculation unit in which it is located to calculate the input it receives; wherein, the input received by the latter part includes the output of the previous part, and the input received by the first part includes the input sequence of the first target task.
[0130] Optionally, at least one module of the target model includes an attention layer, which calculates a query matrix, a key matrix and a value matrix based on the received input, and calculates the output of the attention layer based on the calculated query matrix, key matrix and value matrix; wherein the input received by the attention layer includes a vector matrix corresponding to the input sequence of the first target task, or the input received by the attention layer includes the output of a previous module.
[0131] Optionally, the device further comprises:
[0132] A quantization module, used for quantizing the target data of the first bit width according to the second bit width and storing it in a preset area; the target data includes at least one of the query matrix, the key matrix and the value matrix;
[0133] The attention layer is specifically used to read the target data of the second bit width from the preset area and convert it into the first bit width, and then calculate the output of the attention layer according to the query matrix, key matrix and value matrix of the first bit width.
[0134] Optionally, the attention layer is specifically used to: perform inner product calculations on the query matrix and each row in the key matrix to obtain a first intermediate result; perform e-th power calculations on each element in the first intermediate result to obtain a first transformation result; calculate a first proportion for each element in the first transformation result to obtain a second intermediate result; multiply each element in the second intermediate result by each row in the value matrix and sum the results to obtain the output of the attention layer.
[0135] Optionally, the attention layer is specifically used to: perform inner product calculation on the query matrix and the i-th row in the key matrix to obtain the i-th score; perform e-th power calculation on the i-th score to obtain a second transformation result; accumulate the second transformation result with the accumulated value of the second transformation result of the previous round to obtain the accumulated value of the second transformation result of the current round; multiply the second transformation result with the i-th row in the value matrix to obtain a third transformation result, and accumulate the third transformation result with the accumulated value of the third transformation result of the previous round to obtain the accumulated value of the third transformation result of the current round; repeat the above steps until the i-th row is the last row in the key matrix, and after calculating the accumulated value of the third transformation result of the current round, calculate the ratio of the accumulated value of the third transformation result of the current round to the accumulated value of the second transformation result of the current round to obtain the output of the attention layer.
[0136] Optionally, the device further comprises:
[0137] An input module is used to input an input sequence of a second target task into the target model at a target time, so that the part of the target model that has completed the calculation of the input sequence for the first target task starts to calculate the input sequence for the second target task; the target time is used to indicate that there is a part of the target model that has completed the calculation of the input sequence for the first target task, and there is a part that has not yet completed the calculation of the input sequence for the first target task.
[0138] Optionally, the target model includes M modules, M≥N, and the partitioning and deployment modules include:
[0139] A first deployment submodule is used for, if M is an integer multiple of N, dividing the M modules of the target model into N parts in equal proportion and deploying them on the N computing units, with one part being deployed on each computing unit, and each part including M / N modules; or
[0140] The second deployment submodule is used to divide the M modules of the target model into N parts and deploy them on the N computing units if M is not an integer multiple of N, with one part deployed on each computing unit, each of the N-1 parts containing M / N modules, and the Nth part containing M / N+M%N modules.
[0141] The target model of the embodiment of the present invention is divided into multiple parts and deployed on multiple different computing units to run independently, which can realize the internal parallelization of the target model. When the scale of the target model is large, the target model can be run on multiple computing units, and at the same time, the computing resource utilization rate of the computing unit can be greatly improved, and the efficiency of the target model in performing the target task can be improved. In addition, the embodiment of the present invention performs low-bit quantization on the target data, which can greatly reduce the storage resources required for the target data and increase the data transmission volume, thereby increasing the number of requests that the target model can process at one time. Furthermore, the embodiment of the present invention optimizes the calculation process of the attention layer, and the optimized calculation process can reduce the number of data reads and writes, save data bandwidth, and thus improve the calculation efficiency of the target model.
[0142] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0143] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0144] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0145] The embodiment of the present invention also provides a non-temporary computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a device (server or terminal), the device can execute the above Figure 1 The description of the variable magnification focusing method in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer program product or computer program embodiment involved in this application, please refer to the description of the method embodiment of this application.
[0146] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art not disclosed by the present invention. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present invention is indicated by the following claims.
[0147] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.
[0148] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
[0149] The above is a detailed introduction to a model calculation method and device provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A model calculation method, It is characterized in that Applied to an electronic device, the electronic device includes at least two computing units, and the method includes: Divide the target model into N parts according to the modules and deploy them on the computing unit, where N is less than or equal to the number of computing units included in the electronic device; wherein at least one part is deployed on one computing unit, and each part includes at least one module; Input the input sequence of the first target task into the target model so that each part of the target model uses the computing unit in which it is located to calculate the input it receives; wherein the input received by the latter part includes the output of the previous part, and the input received by the first part includes the input sequence of the first target task.
2. The method according to claim 1, It is characterized in that At least one module of the target model includes an attention layer, which calculates a query matrix, a key matrix and a value matrix based on the received input, and calculates the output of the attention layer based on the calculated query matrix, key matrix and value matrix; wherein the input received by the attention layer includes a vector matrix corresponding to an input sequence of the first target task, or the input received by the attention layer includes the output of a previous module.
3. The method according to claim 2, It is characterized in that After the attention layer calculates the query matrix, the key matrix, and the value matrix based on the received input, the method further includes: quantizing the target data of the first bit width according to the second bit width and storing it in a preset area; the target data includes at least one of the query matrix, the key matrix and the value matrix; The calculating the output of the attention layer according to the calculated query matrix, key matrix and value matrix includes: After reading the target data of the second bit width from the preset area and converting it into the first bit width, the output of the attention layer is calculated based on the query matrix, key matrix and value matrix of the first bit width.
4. The method according to claim 2, It is characterized in that The calculating the output of the attention layer according to the calculated query matrix, key matrix and value matrix includes: Perform inner product calculation on each row in the query matrix and the key matrix to obtain a first intermediate result; Performing e-th power calculation on each element in the first intermediate result to obtain a first transformation result; Calculate a first proportion for each element in the first transformation result to obtain a second intermediate result; Each element in the second intermediate result is multiplied by each row in the value matrix and then the sum is calculated to obtain the output of the attention layer.
5. The method according to claim 2, It is characterized in that The calculating the output of the attention layer according to the calculated query matrix, key matrix and value matrix includes: Performing inner product calculation on the query matrix and the i-th row in the key matrix to obtain the i-th score; Calculating the i-th score to the power of e to obtain a second transformation result; Accumulate the second transformation result and the accumulated value of the second transformation result of the previous round to obtain the accumulated value of the second transformation result of the current round; Multiplying the second transformation result by the i-th row in the value matrix to obtain a third transformation result, and accumulating the third transformation result with the accumulated value of the third transformation result of the previous round to obtain the accumulated value of the third transformation result of the current round; Repeat the above steps until the i-th row is the last row in the key matrix. After calculating the accumulated value of the third transformation result of the current round, calculate the ratio of the accumulated value of the third transformation result of the current round to the accumulated value of the second transformation result of the current round to obtain the output of the attention layer.
6. The method according to claim 1, It is characterized in that The method further comprises: At a target time, an input sequence of a second target task is input into the target model, so that the portion of the target model that has completed the calculation of the input sequence for the first target task starts to calculate the input sequence for the second target task; the target time is used to indicate that there is a portion of the target model that has completed the calculation of the input sequence for the first target task, and there is a portion that has not yet completed the calculation of the input sequence for the first target task.
7. The method according to claim 1, It is characterized in that The target model includes M modules, M≥N, and dividing the target model into N parts according to the modules and deploying them on the computing unit includes: If M is an integer multiple of N, the M modules of the target model are divided into N parts in equal proportion and deployed on the N computing units, with one part deployed on each computing unit, and each part containing M / N modules; or If M is not an integer multiple of N, the M modules of the target model are divided into N parts and deployed on the N computing units, with one part deployed on each computing unit. Each of the N-1 parts contains M / N modules, and the Nth part contains M / N+M%N modules.
8. A model calculation device, It is characterized in that Applied to an electronic device, the electronic device includes at least two computing units, and the device includes: A partitioning and deployment module, used to divide the target model into N parts according to the modules and deploy them on the computing unit, where N is less than or equal to the number of computing units included in the electronic device; wherein at least one part is deployed on one computing unit, and each part includes at least one module; A model calculation module is used to input the input sequence of the first target task into the target model, so that each part of the target model uses the calculation unit in which it is located to calculate the input it receives; wherein the input received by the latter part includes the output of the previous part, and the input received by the first part includes the input sequence of the first target task.
9. The device according to claim 8, It is characterized in that At least one module of the target model includes an attention layer, which calculates a query matrix, a key matrix and a value matrix based on the received input, and calculates the output of the attention layer based on the calculated query matrix, key matrix and value matrix; wherein the input received by the attention layer includes a vector matrix corresponding to an input sequence of the first target task, or the input received by the attention layer includes the output of a previous module.
10. The device according to claim 9, It is characterized in that The device also includes: A quantization module, used for quantizing the target data of the first bit width according to the second bit width and storing it in a preset area; the target data includes at least one of the query matrix, the key matrix and the value matrix; The attention layer is specifically used to read the target data of the second bit width from the preset area and convert it into the first bit width, and then calculate the output of the attention layer according to the query matrix, key matrix and value matrix of the first bit width.