Large model data processing method and device, equipment and storage medium
By performing pre-filling and decoding tasks of the large model in parallel, the problem of low GPU utilization in the existing technology is solved, more efficient resource allocation and memory access is achieved, and the inference performance of the large model is significantly improved.
Patent Information
- Application Number
- CN202510622881.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
The existing large-model inference architecture cannot achieve true parallel computing between pre-filling and decoding processing, resulting in low GPU utilization and limited performance.
Parallel execution of pre-filling and decoding tasks is achieved by generating an input feature matrix and determining thread blocks and operation types according to their processing types. This method allows pre-filling and decoding tasks to be processed simultaneously in the same SM cell, dynamically scheduling GPU resources.
It significantly improves the inference performance of the large model, improves the utilization rate of GPU resources, reduces the waiting time and redundant calculations, adapts to inputs of different lengths, and reduces the number of global memory accesses.
Smart Images

Figure CN120144322A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of large model inference optimization, and more specifically, to a data processing method, device, equipment, and storage medium for large models. Background Art
[0002] In large model inference, each request goes through two processing stages: the prefill stage (abbreviated as Prefill) and decoding (abbreviated as Decode). Among them, Prefill calculates attention and generates KV caches, and Decode autoregressively generates outputs. Among them, Prefill is computationally intensive, and Decode is memory intensive. How to efficiently schedule the two to improve the utilization rate of the Graphics Processing Unit (GPU) is a key challenge.
[0003] The currently common large model inference architecture uses the Chunked Prefill technology based on high-bandwidth memory to split prompts of different lengths into chunks of the same length for prefill, and then inserts the decoding requirements of other prefilled prompts.
[0004] However, the attention modules used in technologies such as Chunked Prefill can only be used for individual prefill or decoding calculations. When the GPU is executing, it still needs to process them sequentially in order, and still needs to wait until the current chunk is completely prefilled before starting decoding, and cannot achieve true parallelism, thus achieving performance improvement. Summary of the Invention
[0005] The purpose of this application is to provide a data processing method, device, equipment, and storage medium for large models to solve the problem of limited performance in the prior art in view of the above deficiencies in the prior art.
[0006] To achieve the above purpose, the technical solutions adopted in the embodiments of this application are as follows: In a first aspect, an embodiment of this application provides a data processing method for a large model, and the method includes: Generating at least one input feature matrix of the large model according to the prompts input by the user, and each of the input feature matrices includes at least one input feature sequence; Determine at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs according to the processing type of each input feature sequence in the input feature matrix, where the processing type includes: pre-filling processing type or decoding processing type, and the operation type includes: pre-filling operation type or decoding operation type; Allocate each input feature sequence to each thread block according to the operation type of each thread block and the processing type of each input feature sequence; Run each thread block in parallel to obtain the operation result corresponding to the input feature matrix; Based on the operation results corresponding to each of the input feature matrices, obtain the output result of the large model.
[0007] In a second aspect, another embodiment of the present application provides a data processing device for a large model, the device includes: A generation module, configured to generate at least one input feature matrix of the large model according to a prompt word input by a user, and each of the input feature matrices includes at least one input feature sequence; A determination module, configured to determine at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs according to the processing type of each input feature sequence in the input feature matrix, where the processing type includes: pre-filling processing type or decoding processing type, and the operation type includes: pre-filling operation type or decoding operation type; An allocation module, configured to allocate each input feature sequence to each thread block according to the operation type of each thread block and the processing type of each input feature sequence; A running module, configured to run each thread block in parallel to obtain the operation result corresponding to the input feature matrix; An output module, configured to obtain the output result of the large model based on the operation results corresponding to each of the input feature matrices.
[0008] In a third aspect, another embodiment of the present application provides an electronic device, including: a processor, a storage medium, and a bus, the storage medium stores machine-readable instructions executable by the processor, when the electronic device runs, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of any method described in the first aspect above.
[0009] In a fourth aspect, another embodiment of the present application provides a storage medium, on which a computer program is stored, and when the computer program is run by a processor, it performs the steps of any method described in the first aspect above.
[0010] The beneficial effects of this application are as follows: at least one input feature matrix of the large model is generated through the prompt words input by the user; and according to the processing types of the input feature sequences in the input feature matrix, at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs are determined; according to the operation type of each thread block and the processing type of each input feature sequence, each input feature sequence is allocated to each thread block; each thread block is run in parallel to obtain the operation result corresponding to the input feature matrix; based on the operation results corresponding to each input feature matrix, the output result of the large model is obtained, which can realize dynamic scheduling of resource allocation at the kernel function level of the GPU, enabling the same SM unit to process the prefill and decoding tasks simultaneously, enabling parallel execution of input feature sequences with different processing types in the input feature matrix, reducing the waiting time, improving the overall throughput, ensuring that the graphics processor resources are fully utilized, avoiding idle or repeated calculations, and thus significantly improving the inference performance of the large model. In addition, it can better adapt to inputs of different lengths, avoid additional padding operations caused by non-integer division of the sequence length, give full play to the capabilities of the graphics processor hardware, and enhance the adaptability to different application scenarios. At the same time, it can also reduce the number of global memory accesses and achieve efficient memory access. Brief Description of the Drawings
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a schematic diagram of a Chunked prefill technology in the prior art; Figure 2 It is a schematic diagram of an architecture of a graphics processor provided by an embodiment of this application; Figure 3 It is a schematic diagram of a structure of a large model provided by an embodiment of this application; Figure 4 It is a schematic flowchart of a data processing method of a large model provided by an embodiment of this application; Figure 5 It is a schematic diagram in a data processing method of a large model provided by an embodiment of this application; Figure 6 It is a schematic flowchart when determining at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs in the data processing method of a large model provided by an embodiment of this application; Figure 7 A flowchart for determining the number of first thread blocks with a prefill operation type and the number of second thread blocks with a decoding operation type in the data processing method of the large model provided by the embodiments of the present application; Figure 8 Another flowchart for determining at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs in the data processing method of the large model provided by the embodiments of the present application; Figure 9 A flowchart for determining the target basic computing unit to which each thread block belongs and the operation type of each thread block in the data processing method of the large model provided by the embodiments of the present application; Figure 10 A schematic diagram of a data processing device for a large model provided by the embodiments of the present application; Figure 11 A schematic diagram of the structure of an electronic device provided by the embodiments of the present application. Detailed implementation manners
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the present application. In addition, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in the present application illustrate the operations implemented according to some embodiments of the present application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and the steps without logical context relationships may be reversed or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the present application.
[0014] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the protection scope of the present application.
[0015] It should be noted that the term "including" will be used in the embodiments of the present application to indicate the existence of the features stated hereinafter, but does not exclude the addition of other features.
[0016] The currently common large model inference architecture uses the Chunked Prefill technology based on high-bandwidth memory to split prompts of different lengths into chunks of the same length for prefill, and then inserts the decoding requirements of other prefilled prompts. Among them, prefill refers to the process in which the model sees the input text and generates intermediate results (such as Key and Value caches), and decode refers to the process in which the model gradually generates the output based on the previously generated results.
[0017] However, although the attention module used in technologies such as Chunked Prefill can already split inputs of different lengths into small chunks to optimize calculations, it can only be used for individual prefill or decode calculations. When the GPU is executing, it still needs to wait for the current chunk to be completely prefilled before starting to decode, that is, it must complete the prefill before starting to decode, and true parallelism cannot be achieved. When the context length increases, such as in applications like RAG or Agent, it is difficult to achieve performance improvement.
[0018] Exemplarily, Figure 1 is a schematic diagram of a Chunked prefill technology in the prior art. Referring to Figure 1 as shown, Chunked prefill cannot calculate the prefill chunk and decode simultaneously in one SM unit. Since the length of the input feature sequence seq is usually not divisible by the total number of threads, the remaining seq must be padded and participate in a new operation again.
[0019] Exemplarily, assume a GPU with 2 Streaming Multiprocessors (SMs), and each SM unit can execute at most two threads. When inputting a sequence seq with a prefill length of 5 and a decode length of 3, then according to the process of chunked prefill, it needs to be calculated three times to obtain the final result.
[0020] This also means that in the prior art, if the length of the input sequence is not divisible by the number of threads, the remaining part needs to be padded, which will result in additional computational overhead. In addition, if the lengths of the input texts are uneven (for example, some Prompts are very long and some are very short), it will lead to uneven allocation of GPU resources, and some computing units in the GPU will be idle. In addition, in the decoding stage, the model needs to frequently load data from the global memory of the GPU, increasing the latency.
[0021] Based on the above problems, the embodiments of the present application propose a data processing method for large models. Through the prompt words input by the user, at least one input feature matrix of the large model is generated; and according to the processing types of the input feature sequences in the input feature matrix, at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs are determined; according to the operation type of each thread block and the processing type of each input feature sequence, each input feature sequence is allocated to each thread block; each thread block is run in parallel to obtain the operation result corresponding to the input feature matrix; based on the operation results corresponding to each input feature matrix, the output result of the large model is obtained, so that the input feature sequences with different processing types in the input feature matrix can be executed in parallel, reducing the waiting time, improving the overall throughput, ensuring that the graphics processor resources are fully utilized, avoiding idle or repeated calculations, and thus significantly improving the inference performance of the large model. In addition, it can better adapt to inputs of different lengths, avoid additional padding operations caused by non-integer division of sequence lengths, and give full play to the capabilities of the graphics processor hardware.
[0022] First, the relevant architecture involved in the data processing method for large models provided by the embodiments of the present application will be described.
[0023] Figure 2 is a schematic diagram of an architecture of a graphics processor provided by the embodiments of the present application. Referring to Figure 2 as shown, the graphics processor provided by the embodiments of the present application includes multiple basic computing units (Streaming Multiprocessor, abbreviated as SM), global memory (abbreviated as Global mem), and secondary cache (abbreviated as L2 cache).
[0024] Among them, the global memory is the main memory of the GPU, and all warps and thread blocks can access it. The secondary cache is used to store data copies recently accessed from the global memory, as well as prefetch content of data to be accessed soon.
[0025] Among them, one basic computing unit serves as an independent computing module, and each basic computing unit includes multiple integer and floating-point arithmetic units (abbreviated as CUDA cores), multiple matrix arithmetic units (abbreviated as Tensor cores), as well as L1 cache and on-chip shared memory.
[0026] Among them, multiple thread blocks can be executed simultaneously on one basic computing unit. Exemplarily, the basic computing unit can dynamically pull thread blocks from the GPU global scheduling queue until the basic computing unit reaches the thread block capacity limit or resource shortage.
[0027] Among them, a thread block includes multiple warps. The warps within the same thread block can share data and cooperate to complete tasks.
[0028] Exemplarily, a warp usually contains 32 threads and is the basic granularity of GPU scheduling. The GPU scheduler will schedule the warp to the CUDA cores of the SM for execution.
[0029] Exemplarily, each thread has independent registers for storing local variables and intermediate results.
[0030] Exemplarily, during the operation process in the large model inference, data will first be loaded from global memory to shared memory or L1 cache, then passed to the registers, and then the CUDA cores or Tensor cores are used for operation to obtain the operation result.
[0031] Exemplarily, the parameters of the large model (such as the weight matrix) are usually stored in global memory. During the inference process of the large model data processing method provided in this application, part of the weights can be loaded into the L2 cache or shared memory through the Tile Technology to reduce the overhead of frequently accessing global memory.
[0032] Exemplarily, the KV cache (Key-Value Cache) is the core data structure in the autoregressive decoding stage and is used to store the historical context information of the generated sequence. The content of the KV cache can first be loaded from global memory to the L2 cache and then further migrated to shared memory or registers to accelerate the calculation of the attention mechanism.
[0033] Exemplarily, the intermediate results of the attention mechanism (such as the calculation results of Query, Key, and Value) can be stored in shared memory for sharing by the threads within the same thread block.
[0034] Exemplarily, the data of the input sequence and output sequence are stored in global memory and are loaded into shared memory or registers as needed during the inference process.
[0035] It can be understood that in the prefill stage of large model inference, the SM unit is responsible for calculating the Query, Key, and Value of the input sequence and generating the initial KV cache. In the decoding stage, the SM unit generates the output sequence word by word, calculates the Query for each time step, and interacts with the historical KV cache. During the large model inference process, by executing the large model data processing method provided in the embodiments of this application, the GPU can perform prefill and decoding calculations simultaneously in a single SM unit, which not only improves the utilization rate of hardware resources but also reduces redundant calculations and memory access latency, thus significantly improving the inference efficiency of the large language model.
[0036] Optionally, the large model can be a large model based on the Transformer architecture. Hereinafter, the structure of the large model involved in the data processing method of the large model provided by the embodiments of the present application will be exemplarily described.
[0037] Figure 3 FIG. is a schematic structural diagram of a large model provided by an embodiment of the present application. Refer to Figure 3 As shown, the large model includes a text embedding layer (abbreviated as Text Embedding) and N processing layers connected in sequence, and finally obtains a text prediction result (Text Prediction), where N is a positive integer.
[0038] Among them, Text Embedding is the input layer of the model, which converts the input text (for example: words or tokens) into a high-dimensional vector representation. Each word is mapped to a vector of a fixed dimension. Specifically, the positional information of the word can be retained through positional encoding (Positional Encoding).
[0039] Among them, each processing layer includes two normalization layers (abbreviated as Layer Norm), a pre-projection layer (abbreviated as Preprojection), a pre-fill attention mechanism (abbreviated as Prefill Attention), a decoding attention mechanism (abbreviated as decodeAttention), and a feed-forward network (abbreviated as Feed Forward).
[0040] Exemplarily, the input text is first converted into an embedding vector, normalized and pre-projected through the normalization layer and the pre-projection layer, then enters the attention mechanism (pre-fill or decoding), and after the attention output, it passes through the feed-forward network again, and is repeatedly processed in multiple layers. Finally, the probability distribution of the predicted next token is output to obtain the final text prediction result (Text Prediction).
[0041] Hereinafter, the data processing method of the large model provided by the embodiments of the present application will be described in detail with reference to multiple embodiments.
[0042] Figure 4 FIG. is a schematic flowchart of a data processing method of a large model provided by an embodiment of the present application. Refer to Figure 4 As shown, the execution subject of this method can be any electronic device with processing capabilities. This method includes: S401. Generate at least one input feature matrix of the large model according to the prompt word input by the user.
[0043] It can be understood that large models can perform reasoning based on the prompt words input by users and generate natural language answers. During the reasoning process of large models, referring to Figure 3 as shown, the input of each processing layer of the large model will be dynamically updated during the processing. The normalization layer and pre-projection layer of each processing layer can process the input of this processing layer into an input feature matrix for the operation of the attention mechanism.
[0044] Optionally, at the beginning of the reasoning process, the large model can first input the prompt words input by the user into the text embedding layer to convert the prompt words input by the user into a Token sequence that the large model can understand, and add special markers to the Token sequence to identify the task type or sequence boundary. At the same time, truncate the Token sequence to the maximum supported length of the model, or pad the Token sequence to a fixed length, and use a special Padding marker to represent the padded part.
[0045] Optionally, after obtaining the Token sequence, the Token sequence can be input into the first processing layer, and the normalization layer and pre-projection layer of the first processing layer can convert the token sequence into a vector and perform positional encoding to obtain the input feature matrix of the first processing layer.
[0046] Optionally, the input feature matrix of the first processing layer can be input into the attention mechanism of the first processing layer of the large model for processing, and after performing feed-forward processing on the result of the attention mechanism operation, it is passed layer by layer to obtain the input feature matrices of other processing layers of the large model. Among them, the number of input feature matrices of the large model can be the same as the number of processing layers of the large model.
[0047] That is to say, referring to Figure 3 as shown, the input feature matrix can be the output of the pre-projection layer of a processing layer, that is, the set of inputs of the pre-padding attention mechanism or the decoding attention mechanism.
[0048] Exemplarily, taking any one input feature matrix as an example, this input feature matrix can be obtained through normalization and pre-projection processing of the input of the corresponding processing layer, and the input of the corresponding processing layer can be obtained through residual connection of the output of the previous processing layer.
[0049] Among them, each input feature matrix includes at least one input feature sequence.
[0050] Exemplarily, the matrix shape of the input feature matrix can be expressed as: [sequence length, feature dimension]. An input feature matrix can be, for example, a sentence, and multiple input feature sequences in this input feature matrix can be, for example, the vectors of multiple tokens in this sentence.
[0051] S402. Determine at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs according to the processing types of the input feature sequences in the input feature matrix.
[0052] It can be understood that taking the inference process of any processing layer in the large model inference process as an example, in the inference process of this processing layer, the input feature sequences in the input feature matrix can perform operations of the prefill attention mechanism or the decoding attention mechanism in this processing layer.
[0053] In this process, the number of thread blocks required for each input feature sequence to perform the attention mechanism operation, the operation type of each thread block, and the target basic computing unit to which each thread block belongs during execution can be determined first, so that each input feature sequence can be executed in parallel in each target computing unit.
[0054] Among them, the processing types include: prefill processing type or decoding processing type, and the operation types include: prefill operation type or decoding operation type.
[0055] Exemplarily, the prefill processing type means that the input feature sequence needs to be processed by the prefill attention mechanism, the prefill operation type means performing the operation of the prefill attention mechanism, the decoding processing type means that the input feature sequence needs to be processed by the decoding attention mechanism, and the decoding operation type means performing the operation of the decoding attention mechanism.
[0056] Exemplarily, according to the different processing types of the input feature sequences, it can be determined whether the input feature sequence needs to perform the operation of the prefill attention mechanism or the decoding attention mechanism, so that the number and operation type of each thread block performing the attention mechanism operation can be determined, and the target basic computing unit to which each thread block belongs can be dynamically determined by the GPU hardware scheduler.
[0057] Exemplarily, Figure 5 is a schematic diagram of a data processing method for a large model provided by an embodiment of the present application. Refer to Figure 5As shown, taking the inference process of any processing layer in the large model inference process as an example, during the inference process of this processing layer, assume that the input feature matrix at this time is an 8*12 matrix, and among the first 5 input feature sequences in the input feature matrix at this time, the processing type is the pre-fill processing type, and the processing type of the last 3 input feature sequences is the decoding processing type. By executing this step, it can be determined that 8 thread blocks are required to execute the operation of the attention mechanism of the input special matrix, and it can be determined that among the 8 thread blocks, the operation types of 5 thread blocks are the pre-fill processing type, and the operation types of 3 thread blocks are the decoding operation type, and it can be determined that the target basic computing units to which these 8 thread blocks belong are SM0 and SM1 respectively.
[0058] Optionally, the above S402 can be deployed in each processing layer of the large model in the form of a kernel function during specific implementation. For example, it can be deployed in the pre-projection layer of each processing layer in the form of a kernel function, and then it can also be deployed at the end of the pre-projection layer of each processing layer.
[0059] Exemplarily, when the attention mechanism is multi-head attention (Multi-Head Self-Attention, abbreviated as MHA), the input data can be divided into multiple "heads" for independent calculation, and finally stitched together, that is, when calculating the blocks, it is divided by a single head, and finally multiplied by the number of heads (num_head).
[0060] S403. According to the operation types of each thread block and the processing types of each input feature sequence, allocate each input feature sequence to each thread block.
[0061] Optionally, continue to refer to Figure 5 As shown, after determining each thread block and the operation type of each thread block, each input feature sequence can be input and matched according to the processing type of each input feature sequence and the operation type of each thread block, so as to allocate each input feature sequence to the corresponding each thread block.
[0062] S404. Run each thread block in parallel to obtain the operation result corresponding to the input feature matrix.
[0063] Optionally, after each input feature sequence is allocated to each thread block, each thread block can be run in parallel, so that each input feature sequence can execute the operation of the pre-fill attention mechanism or the decoding attention mechanism in its respective thread block in each target basic computing unit at the same time, thereby saving the time required for repeated calculation due to the length of the input feature sequence not being divisible, and improving the operation efficiency.
[0064] Exemplarily, the operation result corresponding to the input feature matrix can be calculated by calling flashattention.
[0065] Among them, the operation result corresponding to the input feature matrix can be understood as the set of sub-operation results obtained after each input feature sequence in the input feature matrix parallelly executes the operation of the pre-padding attention mechanism or the operation of the decoding attention mechanism during the inference process of this processing layer.
[0066] S405. Obtain the output result of the large model based on the operation results corresponding to each input feature matrix.
[0067] Optionally, taking the inference process of any processing layer during the inference process of the large model as an example, after obtaining the operation result corresponding to the input feature matrix of this processing layer, subsequent pre-projection processing, normalization, and feed-forward processing can be continued on the operation result corresponding to the input feature matrix in this processing layer, and the output result of this processing layer can be obtained. Then, based on the output result of this processing layer, combined with each subsequent processing layer of this processing layer, the operation results corresponding to each input feature matrix are processed to obtain the output result of the large model.
[0068] In this embodiment, at least one input feature matrix of the large model is generated through the prompt words input by the user; and according to the processing types of each input feature sequence in the input feature matrix, at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs are determined; according to the operation type of each thread block and the processing type of each input feature sequence, each input feature sequence is allocated to each thread block; each thread block is run in parallel to obtain the operation result corresponding to the input feature matrix; based on the operation results corresponding to each input feature matrix, the output result of the large model is obtained, which can realize the dynamic scheduling of resource allocation at the kernel function level of the GPU, enable the same SM unit to process the pre-padding and decoding tasks simultaneously, enable the input feature sequences with different processing types in the input feature matrix to be executed in parallel, reduce the waiting time, improve the overall throughput, ensure that the graphics processor resources are fully utilized, avoid idle or repeated calculations, and thus significantly improve the inference performance of the large model. In addition, it can better adapt to inputs of different lengths, avoid additional padding operations caused by non-integer division of the sequence length, give full play to the capabilities of the graphics processor hardware, enhance the adaptability to different application scenarios. At the same time, it can also reduce the number of global memory accesses and achieve efficient memory access.
[0069] In a possible implementation manner, Figure 6 is a schematic flowchart of a process for determining at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs in the data processing method of the large model provided by the embodiments of the present application. Refer to Figure 6 As shown, in the above S402, it includes: S601. Determine the number of first thread blocks with a pre-filling operation type and the number of second thread blocks with a decoding operation type according to the processing types of the input feature sequences in the input feature matrix.
[0070] Optionally, the number of first thread blocks with a pre-filling operation type and the number of second thread blocks with a decoding operation type can be calculated according to the processing types of the input feature sequences in the input feature matrix.
[0071] Exemplarily, the number of input feature sequences with a pre-filling processing type in the input feature matrix can be determined, and according to the number of input feature sequences with a pre-filling processing type, combined with the number of threads in each preset thread block, the number of first thread blocks with a pre-filling operation type can be determined.
[0072] Exemplarily, the number of input feature sequences with a decoding processing type in the input feature matrix can be determined, and according to the number of input feature sequences with a decoding processing type, combined with the number of threads in each preset thread block, the number of second thread blocks with a decoding operation type can be determined.
[0073] Exemplarily, continue to refer to Figure 5 As shown, taking the inference process of any processing layer in the large model inference process as an example, in the inference process of this processing layer, assume that the input feature matrix is an 8*12 matrix at this time, and among the input feature sequences in the input feature matrix at this time, the processing types of the first 5 input feature sequences are pre-filling processing types, and the processing types of the last 3 input feature sequences are decoding processing types. By executing this step, the number of first thread blocks with a pre-filling operation type can be determined to be 5, and the number of second thread blocks with a decoding operation type can be determined to be 3.
[0074] S602. Determine at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs according to the number of first thread blocks and the number of second thread blocks.
[0075] Optionally, after determining the number of first thread blocks and the number of second thread blocks, the thread blocks required for the input feature matrix during the attention mechanism operation can be determined according to the number of first thread blocks and the number of second thread blocks, that is, at least one thread block corresponding to the input feature matrix, and the operation type of each thread block can be determined.
[0076] Optionally, after determining the thread blocks and the operation type of the thread blocks, the target basic computing unit to which each thread block belongs can be dynamically determined by the GPU hardware scheduler.
[0077] Exemplarily, continue to refer toFigure 5 As shown in the figure, at this time, the processing types of the first 5 input feature sequences in the input feature matrix are pre-filling processing types, and the processing types of the last 3 input feature sequences are decoding processing types. Then, it can be determined that at least one thread block corresponding to the input feature matrix is 5 thread blocks with a pre-filling operation type and 3 thread blocks with a decoding operation type, and the target basic computing units to which each thread block belongs can be dynamically determined by the GPU hardware scheduler as SM0 and SM1.
[0078] By the processing types of the input feature sequences in the input feature matrix, determining the number of the first thread blocks with a pre-filling operation type and the number of the second thread blocks with a decoding type, and based on the number of the first thread blocks and the number of the second thread blocks, determining at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs can improve the utilization rate of each basic computing unit in the GPU, and can also reduce the resource waste caused by sequence filling, thereby significantly improving the inference performance of the large model.
[0079] In a possible implementation manner, Figure 7 FIG. is a schematic flowchart of a process for determining the number of the first thread blocks with a pre-filling operation type and the number of the second thread blocks with a decoding type in the data processing method of the large model provided by the embodiment of the present application. Referring to Figure 7 As shown in the figure, in S601 above, according to the processing types of the input feature sequences in the input feature matrix, determining the number of the first thread blocks with a pre-filling operation type and the number of the second thread blocks with a decoding type includes: S701. Determine the number of input feature sequences with a pre-filling type in the input feature matrix, and calculate the number of the first thread blocks with a pre-filling operation type according to the number of input feature sequences with a pre-filling type.
[0080] Optionally, the number of input feature sequences with a pre-filling type in the input feature matrix can be determined first, and the number of the first thread blocks with a pre-filling operation type can be calculated according to the number of input feature sequences with a pre-filling type, the number of threads of each preset thread block, and the number of heads in the multi-head attention mechanism.
[0081] Exemplarily, the number of the first thread blocks with a pre-filling operation type prefill_blocks can be calculated with reference to the following formula (1): (1) where prefill_blocks is the number of the first thread blocks with a pre-filling operation type, is the number of heads in the multi-head attention mechanism, prefill_blockM is the number of threads in a thread block with a prefill operation type as the preset operation type, and prefill_seqlen is the number of input feature sequences of the prefill type for processing.
[0082] S702. Determine the number of input feature sequences of the decoding type in the input feature matrix, and calculate the number of second thread blocks of the decoding operation type based on the number of input feature sequences of the decoding type.
[0083] Optionally, the number of input feature sequences of the decoding type in the input feature matrix can be determined first, and the number of second thread blocks of the decoding operation type can be calculated based on the number of input feature sequences of the decoding type, the number of threads in each preset thread block, and the number of heads in the multi-head attention mechanism.
[0084] Exemplarily, the number of second thread blocks decode_blocks of the decoding operation type can be calculated with reference to the following formula (2): (2) where decode_blocks is the number of second thread blocks of the decoding operation type, is the number of heads in the multi-head attention mechanism, decode_blockM is the number of threads in a thread block with a decoding operation type as the preset operation type, and decode_seqlen is the number of input feature sequences of the decoding type.
[0085] In a possible implementation manner, Figure 8 This is another flowchart diagram for determining at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs in the data processing method of the large model provided by the embodiments of this application. Refer to Figure 8 As shown, in the above S602, it includes: S801. Determine at least one thread block corresponding to the input feature matrix and the ratio of the prefill processing type to the decoding processing type in the input feature matrix according to the number of first thread blocks and the number of second thread blocks.
[0086] Optionally, at least one thread block corresponding to the input feature matrix can be determined according to the number of first thread blocks and the number of second thread blocks.
[0087] In one example, continue to refer to Figure 5As shown, taking the inference process of any processing layer in the large model inference process as an example, in the inference process of this processing layer, through the above S601, it can be determined that the number of the first thread blocks with the operation type of prefill operation type is 5, and the number of the second thread blocks with the operation type of decoding type is 3. Then, the number of the first thread blocks and the number of the second thread blocks are summed up to determine that the number of thread blocks corresponding to the input feature matrix is 8, and 8 thread blocks can be obtained by starting the kernel and specifying the dimension of the grid. That is, the number of the first thread blocks and the number of the second thread blocks can be summed up, and each thread block corresponding to the input feature matrix can be obtained through the kernel startup configuration.
[0088] In another example, taking the inference process of any processing layer in the large model inference process as an example, in the inference process of this processing layer, it is assumed that the number of the first thread blocks with the operation type of prefill operation type is determined to be 10, and the number of the second thread blocks with the operation type of decoding type is 6. Then, the number of the first thread blocks and the number of the second thread blocks can be summed up to determine that the number of thread blocks corresponding to the input feature matrix is 16, and 16 thread blocks can be obtained by starting the kernel and specifying the dimension of the grid.
[0089] Optionally, the greatest common divisor can be obtained by using the Euclidean algorithm for the number of the first thread blocks and the number of the second thread blocks, and the simplest fraction can be obtained by reduction to get the ratio (prefill_blocks_min: decode_blocks_min) of the prefill processing type to the decoding processing type in the input feature matrix.
[0090] S802. Calculate the indication value corresponding to the input feature matrix according to the ratio of the prefill processing type to the decoding processing type in the input feature matrix.
[0091] Among them, the indication value is used to indicate the sum of the two values in the ratio value of the two operation types in the input feature matrix.
[0092] Optionally, the minimum value prefill_blocks_min of the prefill processing type and the minimum value decode_blocks_min of the decoding processing type can be determined according to the ratio of the prefill processing type to the decoding processing type in the input feature matrix, and the indication value corresponding to the input feature matrix can be calculated.
[0093] Exemplarily, the sum of the minimum value prefill_blocks_min of the prefill processing type and the minimum value decode_blocks_min of the decoding processing type can be calculated as the indication value tag corresponding to the input feature matrix.
[0094] S803. Determine the target basic computing unit to which each thread block belongs and the operation type of each thread block according to the indicated value corresponding to the input feature matrix.
[0095] Optionally, after obtaining the indicated value corresponding to the input feature matrix, the target basic computing unit to which each thread block belongs and the operation type of each thread block can be determined by combining the indicated value tag corresponding to the input feature matrix with the basic computing units allocated by the GPU hardware scheduler.
[0096] In a possible implementation manner, Figure 9 is a schematic flowchart when determining the target basic computing unit to which each thread block belongs and the operation type of each thread block in the data processing method of the large model provided by the embodiments of the present application. Refer to Figure 9 as shown, in the above S803, it includes: S901. Traverse each thread block. For the currently traversed thread block, use the currently allocated basic computing unit as the target basic computing unit to which the current thread block belongs.
[0097] It can be understood that after obtaining multiple thread blocks, the target basic computing unit to which each thread block belongs and the operation type can be determined in sequence.
[0098] Optionally, traverse the obtained multiple thread blocks. For the currently traversed thread block, use the currently allocated basic computing unit as the target basic computing unit to which the current thread block belongs.
[0099] Among them, the currently allocated basic computing unit can be obtained through the GPU hardware scheduler. That is to say, the process of allocating thread blocks to basic computing units can dynamically allocate thread blocks to basic computing units through the GPU hardware scheduler, thereby improving the operation efficiency and avoiding resource waste.
[0100] S902. Obtain the number of executed thread blocks of the current basic computing unit, and determine the operation type of the current thread block according to the indicated value corresponding to the input feature matrix and the number of executed thread blocks.
[0101] Optionally, after determining the current basic computing unit to which the current thread block belongs, the number of executed thread blocks count of the current basic computing unit can be obtained.
[0102] Among them, the number of executed thread blocks is used to indicate how many tasks have been executed on the current basic computing unit.
[0103] Optionally, determine the operation type of the current thread block according to the indicated value corresponding to the input feature matrix, the number of executed thread blocks, and a preset threshold.
[0104] Exemplarily, perform a modulo operation on the executed thread block count and the indication value tag, and compare the result op of the modulo operation with the minimum value prefill_blocks_min of the prefill processing type. If the result op of the modulo operation is less than the minimum value prefill_blocks_min of the prefill processing type, determine that the operation type of the current thread block is the prefill operation type. If the result of the modulo operation is greater than the minimum value prefill_blocks_min of the prefill processing type, determine that the operation type of the current thread block is the decoding operation type.
[0105] Exemplarily, it can be implemented with reference to the following formula (3): (3) Wherein, is the result of the modulo operation, is the executed thread block count, is the indication value.
[0106] Determine the operation type of the current thread block through the indication value corresponding to the input feature matrix and the executed thread block count, so that the operation type of the thread block can be carried out according to a specified ratio, which can ensure that the task distribution on each basic computing unit meets the expectations, achieve dynamic balance of tasks. At the same time, it can also make full use of each basic computing unit, improve the hardware utilization rate and resource utilization rate, reduce the memory access overhead, and help to quickly obtain the operation result in an efficient parallel computing scenario.
[0107] In a possible implementation manner, in the above S403, allocating each input feature sequence to each thread block according to the operation type of each thread block and the processing type of each input feature sequence includes: Traverse the first input feature sequence of the prefill processing type and the second input feature sequence of the decoding processing type in the input feature matrix respectively.
[0108] Optionally, the first input feature sequence of the prefill processing type and the second input feature sequence of the decoding processing type in the input feature matrix can be traversed respectively.
[0109] For the currently traversed first input feature sequence, determine whether there is a thread block of the prefill operation type in the current target basic computing unit. If so, allocate the current first input feature sequence to the thread block of the prefill operation type. Otherwise, use the next target basic computing unit of the current target basic computing unit as the new current target basic computing unit and perform thread block allocation.
[0110] It can be understood that after determining at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs, the first input feature sequence of the pre-filling processing type in the input feature matrix can be allocated to the thread blocks with the pre-filling operation type in a traversal manner.
[0111] Optionally, for the currently traversed first input feature sequence, check whether there is a thread block with the pre-filling operation type in the current target basic computing unit. If so, allocate the current first input feature sequence to the thread block with the pre-filling operation type; otherwise, use the next target basic operation unit of the current target basic computing unit as the new current target basic computing unit and perform thread block allocation.
[0112] Among them, the next target basic operation unit of the current target basic computing unit can be understood as any target basic computing unit other than the current target basic computing unit.
[0113] For the currently traversed second input feature sequence, determine whether there is a thread block with the decoding operation type in the current target basic computing unit. If so, allocate the current second input feature sequence to the thread block with the decoding operation type; otherwise, use the next target basic operation unit of the current target basic computing unit as the new current target basic computing unit and perform thread block allocation.
[0114] It can be understood that after determining at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs, the second input feature sequence of the decoding processing type in the input feature matrix can be allocated to the thread blocks with the decoding operation type in a traversal manner.
[0115] Optionally, for the currently traversed second input feature sequence, check whether there is a thread block with the decoding operation type in the current target basic computing unit. If so, allocate the current second input feature sequence to the thread block with the decoding operation type; otherwise, use the next target basic operation unit of the current target basic computing unit as the new current target basic computing unit and perform thread block allocation.
[0116] Among them, the next target basic operation unit of the current target basic computing unit can be understood as any target basic computing unit other than the current target basic computing unit.
[0117] In a possible implementation manner, the above S404 runs each thread block in parallel to obtain the operation result corresponding to the input feature matrix, including: When the current thread block is running, read the model weight information from the shared storage space corresponding to the target basic computing unit to which the current thread block belongs, and perform operations according to the model weight information to obtain the operation result of the input feature sequence corresponding to the current thread block.
[0118] Optionally, when the current thread block is running, the model weight information can be read from the shared storage space corresponding to the target basic computing unit to which the current thread block belongs, and operations are performed according to the model weight information to obtain the operation result of the input feature sequence corresponding to the current thread block.
[0119] Among them, the model weight information includes: pre-filled weight values and decoding weight values.
[0120] Among them, referring to Figure 2 As shown, the shared storage space corresponding to the target basic computing unit can be understood as the L1 cache and on-chip shared memory in the SM unit.
[0121] Exemplarily, before the inference process of the large model starts, the frequently accessed model weights can be copied from the global memory to the on-chip shared memory of the SM, so that the thread blocks running in the target basic computing unit can read the model weight information from the on-chip shared memory and perform matrix operations, thereby reducing the shared memory access latency and improving the performance of attention calculation.
[0122] Based on the same inventive concept, an apparatus for processing data of a large model corresponding to the method for processing data of a large model is further provided in an embodiment of the present application. Since the principle of solving problems by the apparatus in the embodiment of the present application is similar to the method for processing data of the large model in the above embodiment of the present application, the implementation of the apparatus can refer to the implementation of the method, and the repeated parts will not be described again.
[0123] Referring to Figure 10 As shown, Figure 10 is a schematic diagram of an apparatus for processing data of a large model provided in an embodiment of the present application. The apparatus includes: a generation module 1001, a determination module 1002, an allocation module 1003, a running module 1004, and an output module 1005; The generation module 1001 is configured to generate at least one input feature matrix of the large model according to the prompt words input by the user, and each input feature matrix includes at least one input feature sequence; The determination module 1002 is configured to determine at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs according to the processing type of each input feature sequence in the input feature matrix, where the processing type includes: pre-fill processing type or decoding processing type, and the operation type includes: pre-fill operation type or decoding operation type; The allocation module 1003 is used to allocate each input feature sequence to each thread block according to the operation type of each thread block and the processing type of each input feature sequence; The running module 1004 is used to run each thread block in parallel to obtain the operation result corresponding to the input feature matrix; The output module 1005 is used to obtain the output result of the large model based on the operation results corresponding to each input feature matrix.
[0124] Optionally, the determination module 1002 is specifically used for: According to the processing type of each input feature sequence in the input feature matrix, determine the number of first thread blocks with the operation type of pre-fill operation type and the number of second thread blocks with the operation type of decoding type; According to the number of first thread blocks and the number of second thread blocks, determine at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs.
[0125] Optionally, the determination module 1002 is specifically used for: Determine the number of input feature sequences with the pre-fill type in the input feature matrix, and calculate the number of first thread blocks with the operation type of pre-fill operation type according to the number of input feature sequences with the pre-fill type; Determine the number of input feature sequences with the decoding type in the input feature matrix, and calculate the number of second thread blocks with the operation type of decoding operation type according to the number of input feature sequences with the decoding type.
[0126] Optionally, the determination module 1002 is specifically used for: According to the number of first thread blocks and the number of second thread blocks, determine at least one thread block corresponding to the input feature matrix and the ratio of the pre-fill processing type to the decoding processing type in the input feature matrix; According to the ratio of the pre-fill processing type to the decoding processing type in the input feature matrix, calculate the indication value corresponding to the input feature matrix, and the indication value is used to indicate the sum of two values in the ratio values of the two operation types in the input feature matrix; According to the indication value corresponding to the input feature matrix, determine the target basic computing unit to which each thread block belongs and the operation type of each thread block.
[0127] Optionally, the determination module 1002 is specifically used for Traverse each thread block, and for the currently traversed thread block, use the currently allocated basic computing unit as the target basic computing unit to which the current thread block belongs; Obtain the number of executed thread blocks of the current basic computing unit, and determine the operation type of the current thread block according to the indicated value corresponding to the input feature matrix and the number of executed thread blocks.
[0128] Optionally, the allocation module 1003 is specifically used for Traverse the first input feature sequence of the pre-filling processing type and the second input feature sequence of the decoding processing type in the input feature matrix respectively; For the currently traversed first input feature sequence, determine whether there is a thread block of the pre-filling operation type in the current target basic computing unit. If so, allocate the current first input feature sequence to the thread block of the pre-filling operation type. Otherwise, use the next target basic operation unit of the current target basic computing unit as the new current target basic computing unit and perform thread block allocation; For the currently traversed second input feature sequence, determine whether there is a thread block of the decoding operation type in the current target basic computing unit. If so, allocate the current second input feature sequence to the thread block of the decoding operation type. Otherwise, use the next target basic operation unit of the current target basic computing unit as the new current target basic computing unit and perform thread block allocation.
[0129] Optionally, the running module 1004 is specifically used for: When the current thread block is running, read the model weight information from the shared storage space corresponding to the target basic computing unit to which the current thread block belongs, and perform operations according to the model weight information to obtain the operation result of the input feature sequence corresponding to the current thread block, where the model weight information includes: pre-filling weight value and decoding weight value.
[0130] The description of the processing flow of each module in the device and the interaction flow between modules can refer to the relevant descriptions in the above method embodiments, and will not be elaborated here.
[0131] The embodiment of the present application also provides an electronic device, such as Figure 11 shown Figure 11 is a schematic structural diagram of the electronic device provided by the embodiment of the present application, including: a processor 1101, a memory 1102. Optionally, a bus 1103 may also be included. The memory 1102 stores machine-readable instructions executable by the processor 1101 (for example, Figure 10 the execution instructions corresponding to the generation module 1001, the determination module 1002, the allocation module 1003, the running module 1004, and the output module 1005 in the device), when the electronic device runs, the processor 1101 communicates with the memory 1102 through the bus 1103, and when the machine-readable instructions are executed by the processor 1101, the steps of the data processing method of the above large model are executed.
[0132] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the data processing method of the above-mentioned large model.
[0133] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the method embodiments, which will not be elaborated in this application. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical, or other form.
[0134] In addition, in each embodiment of this application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0135] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application.
Claims
1. A data processing method for a large model, characterized in that: The method comprises: Generate at least one input feature matrix of the large model according to the prompt word input by the user, each of the input feature matrices including at least one input feature sequence; Determine, according to the processing type of each input feature sequence in the input feature matrix, at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs, wherein the processing type includes: a pre-filling processing type or a decoding processing type, and the operation type includes: a pre-filling operation type or a decoding operation type; Allocate each input feature sequence to each thread block according to the operation type of each thread block and the processing type of each input feature sequence; Running each thread block in parallel to obtain a calculation result corresponding to the input feature matrix; Based on the calculation results corresponding to each of the input feature matrices, the output result of the large model is obtained.
2. The data processing method of the large model according to claim 1, characterized in that: The step of determining, according to the processing type of each input feature sequence in the input feature matrix, at least one thread block corresponding to the input feature matrix, the operation type of each thread block, and the target basic computing unit to which each thread block belongs comprises: Determining, according to the processing type of each input feature sequence in the input feature matrix, the number of first thread blocks whose operation type is a pre-filling operation type and the number of second thread blocks whose operation type is a decoding type; At least one thread block corresponding to the input feature matrix, a calculation type of each thread block, and a target basic computing unit to which each thread block belongs are determined according to the first thread block number and the second thread block number.
3. The data processing method of the large model according to claim 2 is characterized in that: The determining, according to the processing type of each input feature sequence in the input feature matrix, the number of first thread blocks whose operation type is a pre-filling operation type and the number of second thread blocks whose operation type is a decoding type comprises: Determine the number of input feature sequences whose processing type is a pre-filling type in the input feature matrix, and calculate the number of first thread blocks whose operation type is a pre-filling operation type according to the number of input feature sequences whose processing type is a pre-filling type; The number of input feature sequences whose processing type is decoding type in the input feature matrix is determined, and the number of second thread blocks whose operation type is decoding operation type is calculated according to the number of input feature sequences whose processing type is decoding type.
4. The data processing method of the large model according to claim 2 is characterized in that: The step of determining, according to the first thread block number and the second thread block number, at least one thread block corresponding to the input feature matrix, a calculation type of each thread block, and a target basic computing unit to which each thread block belongs includes: Determine, according to the first thread block number and the second thread block number, at least one thread block corresponding to the input feature matrix and a ratio of the pre-filled processing type to the decoded processing type in the input feature matrix; According to the ratio of the pre-filled processing type to the decoded processing type in the input feature matrix, an indication value corresponding to the input feature matrix is calculated, and the indication value is used to indicate the sum of two values in the ratio values of the two operation types in the input feature matrix; According to the indication value corresponding to the input feature matrix, the target basic computing unit to which each thread block belongs and the operation type of each thread block are determined.
5. The data processing method of the large model according to claim 4 is characterized in that: The step of determining the target basic computing unit to which each thread block belongs and the operation type of each thread block according to the indication value corresponding to the input feature matrix includes: Traversing each of the thread blocks, and for the traversed current thread block, using the current basic computing unit allocated as the target basic computing unit to which the current thread block belongs; The number of executed thread blocks of the current basic computing unit is obtained, and the operation type of the current thread block is determined according to the indication value corresponding to the input feature matrix and the number of executed thread blocks.
6. The data processing method of the large model according to claim 1, characterized in that: The allocating each input feature sequence to each thread block according to the operation type of each thread block and the processing type of each input feature sequence includes: traversing respectively a first input feature sequence of a pre-filled processing type and a second input feature sequence of a decoding processing type in the input feature matrix; For the traversed current first input feature sequence, determine whether there is a thread block of the pre-filled operation type in the current target basic computing unit, if so, allocate the current first input feature sequence to the thread block of the pre-filled operation type, otherwise, use the next target basic computing unit of the current target basic computing unit as the new current target basic computing unit and allocate the thread block; For the traversed current second input feature sequence, determine whether there is a thread block of the decoding operation type in the current target basic computing unit. If so, assign the current second input feature sequence to the thread block of the decoding operation type. Otherwise, take the next target basic computing unit of the current target basic computing unit as the new current target basic computing unit and perform thread block assignment.
7. The data processing method of the large model according to claim 1 is characterized in that: The parallel execution of the thread blocks to obtain the operation result corresponding to the input feature matrix includes: When the current thread block is running, the model weight information is read from the shared storage space corresponding to the target basic computing unit to which the current thread block belongs, and calculations are performed according to the model weight information to obtain the calculation results of the input feature sequence corresponding to the current thread block, wherein the model weight information includes: pre-filled weight values and decoded weight values.
8. A data processing device for a large model, characterized in that: The device comprises: A generating module, used to generate at least one input feature matrix of the large model according to the prompt word input by the user, each of the input feature matrices including at least one input feature sequence; A determination module, configured to determine, according to a processing type of each input feature sequence in the input feature matrix, at least one thread block corresponding to the input feature matrix, an operation type of each thread block, and a target basic computing unit to which each thread block belongs, wherein the processing type includes: a pre-filling processing type or a decoding processing type, and the operation type includes: a pre-filling operation type or a decoding operation type; An allocation module, used for allocating each input feature sequence to each thread block according to the operation type of each thread block and the processing type of each input feature sequence; A running module, used for running each thread block in parallel to obtain a calculation result corresponding to the input feature matrix; The output module is used to obtain the output result of the large model based on the calculation results corresponding to each of the input feature matrices.
9. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the data processing method of the large model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the large model data processing method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Data processing method and device, equipment, storage medium and program product
CN118863067A
Inference method and system, electronic equipment and storage medium
CN119831033A
Response information generation method and device, medium and computer program product
CN119884332A
Acceleration method of matrix multiplication, electronic equipment and storage medium
CN119988811A
Semantic understanding model training method and apparatus, computer device, and storage medium
WO2021169288A1
Cited By
Data processing method and device, electronic equipment and storage medium
CN120610829A