Data Processing Method and Apparatus, Storage Medium, and Electronic Device Based on Large Model
By storing key-value information in a continuous video memory area in the pre-filling stage of large-scale model separated inference, the problems of transmission delay and low data processing efficiency are solved, and more efficient data processing is achieved.
Patent Information
- Application Number
- CN202411846922.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-13
AI Technical Summary
When pre-filling of large model separated inference, the kv cache output from each layer of the model is stored in a continuous video memory area, but when transmitted to the decoding part, only one continuous video memory space can be transmitted at a time, resulting in transmission delay and low data processing efficiency.
The description information is processed by the first graphics processor that pre-filled the node, outputs the key-value information and stores it in the target video memory area, ensuring that the key-value information in the same key-value logical block identification is stored in continuous video memory addresses.
The number of transmission times of kv cache is reduced, the transmission efficiency of kv cache is improved, and thus the data processing efficiency of the large model is improved.
Smart Images

Figure CN119338010B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, it relates to a data processing method, device, storage medium, and electronic device based on a large model. Background Art
[0002] For generative large model inference services, reusing the kv cache (key-value cache) as much as possible to reduce the required computing resources is the core to improve the overall throughput. For a split inference framework, prefill and decoding are separated so that each can run independently. An important process in the split architecture is to transfer the reusable kv cache generated in the prefill stage to the selected decode. This process needs to be as fast as possible, otherwise it will affect the response speed of the inference service.
[0003] In the prior art, to facilitate the forward propagation calculation of the model layer by layer, the physical cache structure of the kv cache is designed such that each layer of the model uses a continuous video memory area. For example, the data structure of the kv cache is: , where each list element represents the continuous kv cache of a single layer. This design brings the problem of inter-block discreteness. Since the send API and recv API of the collective communication library (used for communication between multiple GPUs) can only transfer a single kv cache with continuous video memory space each time they are called, the number of API calls required for kv cache transfer is directly related to the number of discrete memory blocks. This will result in significant transfer latency. For example, the number of discrete memory blocks that a block of kv cache may have is the number of layers (64 layers in a 32B model) multiplied by 2 (for k or v). Transmitting one block requires 128 calls to the API atomic capability. This discreteness has an adverse impact on the overall computing performance.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide a data processing method, device, storage medium, and electronic device based on a large model, so as to at least solve the technical problem that when pre-filling in the split inference of a large model, storing the kv output by each layer of the model in a continuous video memory area, and only being able to transfer one kv with a continuous video memory space each time when the kv needs to be transferred to the decoding part for decoding, resulting in transfer latency and thus relatively low data processing efficiency.
[0006] According to one aspect of the embodiments of the present application, a data processing method based on a large model is provided, including: processing the description information input by the target object through a first inference model in a first graphics processor corresponding to a prefill node to obtain key-value information and a first reply character output by the first inference model in the first graphics processor; storing the key-value information belonging to the same key-value logical block identifier in a target video memory area in the first graphics processor based on the key-value logical block identifier corresponding to the key-value information, where the storage addresses of the key-value information belonging to the same key-value logical block identifier in the target video memory area are consecutive; processing the key-value information and the first reply character in the target video memory area through a second inference model in a second graphics processor corresponding to a decoding node to obtain target reply information corresponding to the description information.
[0007] Further, before processing the description information through the first inference model in the first graphics processor corresponding to the prefill node to obtain the key-value information and the first reply character output by the first inference model in the first graphics processor, the method further includes: determining whether the weight parameters corresponding to the target inference model need to be split; if the weight parameters corresponding to the target inference model need to be split, determining a first quantity of the first graphics processors and a second quantity of the second graphics processors required based on the description information; splitting the weight parameters of the target inference model to obtain a first quantity of first inference models, and determining the first graphics processor based on the first quantity of first inference models; splitting the weight parameters of the target inference model to obtain a second quantity of second inference models, and determining the second graphics processor based on the second quantity of second inference models.
[0008] Further, before processing the key-value information and the first reply character in the target video memory area through the second inference model in the second graphics processor corresponding to the decoding node to obtain the target reply information corresponding to the description information, the method further includes: determining whether the quantity of the first graphics processors is the same as the quantity of the second graphics processors; if the quantity of the first graphics processors is the same as the quantity of the second graphics processors, sending the key-value information in the target video memory area to a second target graphics processor through the first graphics processor, where the second target graphics processor is the second graphics processor corresponding to the decoding node that needs to decode the key-value information stored in the first graphics processor.
[0009] Further, sending the key value information in the target video memory area to the second graphics processor by the first graphics processor includes: determining the second target graphics processor by the first graphics processor based on the mapping relationship between the first graphics processor and the second graphics processor; determining the ID of the first target key value information required by the second target graphics processor in the target video memory area; reading the first target key value information from the target video memory area by the first graphics processor based on the ID of the first target key value information; and sending the first target key value information to the second target graphics processor by the first graphics processor.
[0010] Further, after sending the first target key value information to the second target graphics processor by the first graphics processor, the method further includes: receiving the key value information by the second target graphics processor and determining the video memory address for storing the first target key value information in the second target graphics processor; and writing the first target key value information to the video memory address by the second target graphics processor.
[0011] Further, after determining whether the number of the first graphics processors is the same as the number of the second graphics processors, the method further includes: if the number of the first graphics processors is not the same as the number of the second graphics processors, determining the ID of the second target key value information to be sent in the first graphics processor; reading the second target key value information from the target video memory area by a parallel copy operator based on the ID of the second target key value information, and writing the second target key value information into a preset send buffer in parallel; and sending the third target key value information in the send buffer to the second image processor in the decoding node by a preset send function.
[0012] Further, after sending the third target key value information in the send buffer to the multiple decoding nodes by a preset send function, the method further includes: receiving the third target key value information by a preset receive function and writing the third target key value information into a preset receive buffer; determining the fourth target key value information required by the second graphics processor corresponding to the decoding node; splitting the key value information in the receive buffer by a parallel copy operator to obtain the fourth target key value information, and writing the fourth target key value information into the video memory area in the second graphics processor corresponding to the decoding node.
[0013] Further, before processing the description information through the first inference model in the first graphics processor corresponding to the pre-fill node, the method further includes: for the first graphics processor, dividing the description information into key-value logic blocks according to a preset key-value block size to obtain the key-value logic block identifiers of the characters in the description information; and determining the key-value logic block identifier corresponding to the key-value information to be generated according to the key-value logic block identifiers of the characters in the description information.
[0014] According to another aspect of the embodiments of the present application, there is also provided a data processing device based on a large model, including: a first processing unit, configured to process description information through a first inference model in a first graphics processor corresponding to a pre-fill node to obtain key-value information and a first reply character output by the first inference model in the first graphics processor, where the first inference model in the first graphics processor corresponding to the pre-fill node is determined by a target inference model; a storage unit, configured to store the key-value information belonging to the same key-value logic block identifier into a target video memory area in the first graphics processor based on the key-value logic block identifier corresponding to the key-value information, where the storage addresses of the key-value information belonging to the same key-value logic block identifier in the target video memory area are consecutive; and a second processing unit, configured to process the key-value information and the first reply character in the target video memory area through a second inference model in a second graphics processor corresponding to a decoding node to obtain a target reply information corresponding to the description information, where the second inference model in the second graphics processor corresponding to the decoding node is determined by the target inference model.
[0015] Further, the device further includes: a first judgment unit, configured to judge whether the weight parameters corresponding to the target inference model need to be split before processing the description information through the first inference model in the first graphics processor corresponding to the pre-fill node to obtain the key-value information and the first reply character output by the first inference model in the first graphics processor; a first determination unit, configured to, if the weight parameters corresponding to the target inference model need to be split, determine a first quantity of the first graphics processors and a second quantity of the second graphics processors required based on the description information; a first splitting unit, configured to split the weight parameters of the target inference model to obtain a first quantity of first inference models, and determine the first graphics processors based on the first quantity of first inference models; and a second splitting unit, configured to split the weight parameters of the target inference model to obtain a second quantity of second inference models, and determine the second graphics processors based on the second quantity of second inference models.
[0016] Further, the device further includes: a second determination unit, configured to determine whether the number of the first graphics processors is the same as the number of the second graphics processors before processing the key-value information and the first reply character in the target video memory area through a second inference model in the second graphics processor corresponding to the decoding node to obtain the target reply information corresponding to the description information; a first sending unit, configured to, if the number of the first graphics processors is the same as the number of the second graphics processors, send the key-value information in the target video memory area to a second target graphics processor through the first graphics processor, where the second target graphics processor is the second graphics processor corresponding to the decoding node that needs to perform decoding processing on the key-value information stored in the first graphics processor.
[0017] Further, the sending unit includes: a first determination module, configured to determine the second target graphics processor through the first graphics processor based on the mapping relationship between the first graphics processor and the second graphics processor; a second determination module, configured to determine the ID of the first target key-value information in the target video memory area required by the second target graphics processor; a reading module, configured to read the first target key-value information from the target video memory area through the first graphics processor based on the ID of the first target key-value information; a sending module, configured to send the first target key-value information to the second target graphics processor through the first graphics processor.
[0018] Further, the device further includes: a first receiving unit, configured to, after sending the first target key-value information to the second target graphics processor through the first graphics processor, receive the key-value information through the second target graphics processor and determine the video memory address for storing the first target key-value information in the second target graphics processor; a writing unit, configured to write the first target key-value information into the video memory address through the second target graphics processor.
[0019] Further, the device further includes: a second determination unit, configured to, after determining whether the number of the first graphics processors is the same as the number of the second graphics processors, if the number of the first graphics processors is not the same as the number of the second graphics processors, determine the ID of the second target key-value information to be sent in the first graphics processor; a reading unit, configured to read the second target key-value information from the target video memory area based on the ID of the second target key-value information through a parallelized copy operator and parallelly write the second target key-value information into a preset sending buffer; a second sending unit, configured to send the third target key-value information in the sending buffer to the second graphics processor corresponding to the decoding node through a preset sending function.
[0020] Further, the device further includes: a second receiving unit, configured to, after sending the third target key-value information in the sending buffer to the second graphics processor corresponding to the decoding node through a preset sending function, receive the third target key-value information through a preset receiving function, and write the third target key-value information into a preset receiving buffer; a third determining unit, configured to determine fourth target key-value information required by the second graphics processor corresponding to the decoding node; a third splitting unit, configured to split the key-value information in the receiving buffer through a parallelized replication operator to obtain the fourth target key-value information, and write the fourth target key-value information into the video memory area of the second graphics processor corresponding to the decoding node.
[0021] Further, the device further includes: a partitioning unit, configured to, before processing the description information through a first inference model in a first graphics processor corresponding to a pre-filling node, perform key-value logical block partitioning on the description information according to a preset key-value block size for the first graphics processor, so as to obtain key-value logical block identifiers of characters in the description information; a fourth determining unit, configured to determine a key-value logical block identifier corresponding to the key-value information to be generated according to the key-value logical block identifiers of characters in the description information.
[0022] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the data processing method based on a large model described in any one of the above.
[0023] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including a memory storing an executable program; a processor, configured to run the program, and when the program runs, it executes the data processing method based on a large model described in any one of the above.
[0024] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, where the computer program product includes a stored computer program, and when the computer program runs on a processor, it implements the data processing method based on a large model described in any one of the above.
[0025] In the embodiment of the present application, the description information is processed by using the first inference model in the first graphics processor corresponding to the pre-filled node to obtain the key-value information and the first reply character output by the first inference model in the first graphics processor; based on the key-value logic block identifier corresponding to the key-value information, the key-value information belonging to the same key-value logic block identifier is stored in the target video memory area in the first graphics processor, wherein the storage addresses of the key-value information belonging to the same key-value logic block identifier in the target video memory area are continuous; by using the second inference model in the second graphics processor corresponding to the decoding node to process the key-value information and the first reply character in the target video memory area to obtain the target reply information corresponding to the description information, by storing the key-value information belonging to the same key-value logic block identifier in the continuous target video memory area in the first graphics processor, the purpose of reducing the number of kv transmissions is achieved, thereby achieving the technical effect of improving the data processing efficiency of the inference model, and further solving the technical problem that when pre-filling in the split inference of the large model, the kv output by each layer of the model is stored in a continuous video memory area, and when it is necessary to transmit the kv to the decoding part for decoding, only one continuous video memory space of kv can be transmitted each time, resulting in transmission delay and thus relatively low data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0027] Figure 1 is a hardware structure block diagram of a computer terminal provided in Embodiment 1 of the present application;
[0028] Figure 2 is a flowchart of a data processing method based on a large model provided in Embodiment 1 of the present application Figure 1 ;
[0029] Figure 3 is a flowchart of a data processing method based on a large model provided in Embodiment 1 of the present application Figure 2 ;
[0030] Figure 4 is a flowchart of a data processing method based on a large model provided in Embodiment 1 of the present application Figure 3 ;
[0031] Figure 5 is a schematic diagram of a data processing device based on a large model provided in Embodiment 2 of the present application.
[0032] Figure 6 is a block diagram of the structure of an electronic device provided in Embodiment 3 of the present application. Detailed implementation manners
[0033] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] First, some nouns or terms that appear during the description of the embodiments of this application are applicable to the following explanations.
[0036] kv block: The attention key and value cache for storing a fixed number of tokens.
[0037] num_blocks: The number of kv blocks.
[0038] num_layers: The number of model layers.
[0039] block_size: The size of the kv block.
[0040] num_kv_heads: The number of heads of the kv cache.
[0041] head_size: The size of the head of the kv cache.
[0042] cache_dim: The tensor dimension of the kv cache actually stored in a single kvblock;
[0043] .
[0044] Tensor Parallelism (TP): A technique for distributed training and inference. In Tensor Parallelism, the parameters of the model (such as weight matrices) are divided into multiple small pieces and distributed to different GPUs, with each GPU responsible for processing a different part.
[0045] GPU rank: Usually refers to the unique identifier (ID or serial number) of each GPU in a multi-GPU system.
[0046] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant region, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0047] Example 1
[0048] According to the embodiments of the present application, a data processing method based on a large model is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0049] The method embodiments provided in the first embodiment of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method based on a large model is shown. As Figure 1 shown, the computer terminal (or mobile device) 10 may include a set of processors 102 (the set of processors 102 may include, but is not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA, and the set of processors 102 may include a set of processors, Figure 1 where 102a, 102b,..., 102n are used to illustrate), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB, Universal Serial Bus) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include moreFigure 1 more or fewer components as shown, or having a configuration different from that Figure 1 shown.
[0050] It should be noted that one or more of the above-mentioned processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0051] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the large model-based data processing method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned large model-based data processing method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include memories remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0052] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0053] The display can be a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0054] Under the above operating environment, the present application provides a large model-based data processing method as Figure 2 shown. Figure 2is the process of the data processing method based on the large model according to Embodiment 1 of the present application Figure 1 .
[0055] Step S201, process the description information input by the target object through the first inference model in the first graphics processor corresponding to the pre-fill node, and obtain the key-value information and the first reply character output by the first inference model in the first graphics processor.
[0056] Step S202, based on the key-value logic block identifier corresponding to the key-value information, store the key-value information belonging to the same key-value logic block identifier in the target video memory area in the first graphics processor, where the storage addresses of the key-value information belonging to the same key-value logic block identifier in the target video memory area are continuous.
[0057] Optionally, the split inference framework splits the pre-fill and decoding of the inference model and processes them separately by different GPU (i.e., graphics processor) nodes. The specific inference process is as Figure 3 shown. Receive the description information input by the target object (for example, the question information to be replied) through the front-end interface, and send the description information to the GPU node corresponding to the pre-fill. In the pre-fill stage, convert the received description information to be replied into kv and store it in its own video memory area, and output the first reply character, that is, the first token. Send the kv and the first taken to the GPU node corresponding to the decoding for decoding until the final reply information is obtained.
[0058] In the prior art, in order to facilitate the forward propagation calculation of the model layer by layer, the physical cache structure of the kv cache is designed such that each layer of the model uses a continuous video memory area. For example, the data structure of the kv cache is as follows: , and each list element represents a single-layer continuous kv cache. This design brings the problem of inter-block discreteness. Since the send API and recv API of the collective communication library (the collective communication library is used for communication between multiple GPUs) can only transfer a single continuous kv cache in the video memory space each time they are called, the number of API calls required for kv cache transmission is directly related to the number of discrete memory blocks, which will cause significant transmission delays. For example, the number of discrete memory blocks that a block of kv cache may have is the number of layers (64 layers in a 32B model) multiplied by 2 (for k or v), and 128 API atomic capability calls are required to transfer one block.
[0059] To solve the above problems, in the data processing method based on a large model provided in the first embodiment of the present application, the description information is processed by the first inference model in the first graphics processor corresponding to the prefill node, and the key-value information and the first reply character output by the first inference model in the first graphics processor corresponding to the prefill node are obtained.
[0060] It should be noted that the first inference model in the first graphics processor corresponding to the above prefill node is determined by the target inference model. The target inference model is a large model, which refers to a machine learning model with large-scale parameters and complex computational structures, usually constructed by a deep neural network.
[0061] After obtaining the above kv cache (i.e., the above key-value information), the cache structure of the kv cache is transformed. For the first graphics processor corresponding to the prefill node, the key-value information in the kv cache output by the first graphics processor that belongs to the same key-value logical block identifier is stored in the target video memory area in the first graphics processor. It should be noted that the storage addresses of the key-value information belonging to the same key-value logical block identifier in the target video memory area are continuous, that is, the physical addresses of the key-value information belonging to the same key-value logical block identifier are continuous. The kv cache cache structure is transformed into a tensor with a continuous space, as shown below.
[0062]
[0063] Compared with the prior art
[0064] This improvement can effectively reduce the number of network API calls. Because the kv cache of one block is at the same physical address and there is no need to read the kv cache of one block from each layer, therefore, the number of times each kv block calls the network API can be reduced by "the number of layers multiplied by 2 (for k or v)".
[0065] Step S203, the key-value information and the first reply character in the target video memory area are processed by the second inference model in the second graphics processor corresponding to the decoding node to obtain the target reply information corresponding to the description information, where the second inference model in the second graphics processor corresponding to the decoding node is determined by the target inference model.
[0066] Optionally, after obtaining the corresponding reusable kv cache and first token, the key-value information and the first reply character in the target video memory area are processed by the second inference model in the second graphics processor corresponding to the decoding node, and then the target reply information corresponding to the description information is obtained.
[0067] It should be noted that the number of first graphics processors corresponding to the pre-fill nodes can be one or more, and the number of second graphics processors corresponding to the decoding nodes can be one or more. When the number of first graphics processors is more than one, for any one of the first graphics processors, only the kv cache output by itself is stored.
[0068] In summary, after the key-value information output by the first inference model in the first graphics processor corresponding to the pre-fill node, the key-value information belonging to the same key-value logical block identifier in the key-value information output by the first graphics processor is stored in the video memory area with continuous addresses. When it is necessary to transfer the kv in a key-value logical block to the decoding part for decoding, there is no need to repeatedly read the kv from the video memory area many times, which can effectively reduce the number of kv transfers and improve the kv transfer efficiency, thereby achieving the effect of improving the data processing efficiency.
[0069] To improve the parallel processing efficiency, in the data processing method based on the large model provided in the first embodiment of the present application, before processing the description information through the first inference model in the first graphics processor corresponding to the pre-fill node to obtain the key-value information and the first reply character output by the first inference model in the first graphics processor corresponding to the pre-fill node, the method further includes: determining whether the weight parameters corresponding to the target inference model need to be split; if the weight parameters corresponding to the target inference model need to be split, then based on the description information, determining the first quantity of the first graphics processors required and the second quantity of the second graphics processors required; splitting the weight parameters of the target inference model to obtain the first quantity of first inference models, and determining the first graphics processors based on the first quantity of first inference models; splitting the weight parameters of the target inference model to obtain the second quantity of second inference models, and determining the second graphics processors based on the second quantity of second inference models.
[0070] Optionally, to improve the pre-fill and decoding speeds, multiple first graphics processors can be set at the pre-fill nodes, and multiple second graphics processors can be set at the decoding nodes. It can be judged according to actual needs whether the weight parameters corresponding to the target inference model need to be split, that is, whether it is necessary to perform parallel pre-fill through multiple first graphics processors and whether it is necessary to perform parallel decoding processing through multiple second graphics processors.
[0071] If the weight parameters corresponding to the target inference model need to be split, determine the first quantity of the required first graphics processors and the second quantity of the required second graphics processors according to the type of the description information. For example, the long text summarization Q&A task is a typical long input and short output, and the content creation and generation task is a short input and long output. For the task type of long input and short output, the number of pre-fill nodes can be set to be larger, and the number of decoding nodes can be set to be smaller. For the task type of short input and long output, the number of pre-fill nodes can be set to be smaller, and the number of decoding nodes can be set to be larger.
[0072] After obtaining the above first quantity and second quantity, the weight parameters of the target inference model can be split by Tensor Parallelism (TP) to obtain the first inference models of the first quantity, and then the first graphics processors can be obtained based on the first inference models of the first quantity. And split the weight parameters of the target inference model by Tensor Parallelism (TP) to obtain the second inference models of the second quantity, and determine the second graphics processors according to the second inference models of the second quantity.
[0073] By splitting the weight parameters of the target inference model, parallel pre-filling and parallel decoding can be achieved, improving the data processing efficiency of the target inference model.
[0074] It should be noted that if the weight parameters of the target inference model do not need to be split, the first inference model in the first graphics processor corresponding to a pre-fill node is the target inference model, and the second inference model in the second graphics processor corresponding to a decoding node is also the target inference model. And the number of GPUs corresponding to a pre-fill node is one, and the number of GPUs corresponding to a decoding node is also one.
[0075] To improve the data processing efficiency, in the data processing method based on a large model provided in the first embodiment of this application, before the second inference model in the second graphics processor corresponding to the decoding node processes the key-value information and the first reply character in the target video memory area to obtain the target reply information corresponding to the description information, the method further includes: determining whether the number of the first graphics processors is the same as the number of the second graphics processors; if the number of the first graphics processors is the same as the number of the second graphics processors, send the key-value information in the target video memory area to the second target graphics processor through the first graphics processor, where the second target graphics processor is the second image processor corresponding to the decoding node that decodes the key-value information stored in the first graphics processor.
[0076] Optionally, since the last layer of the original KV cache changes when the inference service enables TP, as follows.
[0077] From
[0078] change to
[0079]
[0080] This is because TP will split according to the dimension of the head and calculate the corresponding kv cache on the corresponding prefill. The TP size refers to the number of tensor parallelisms, which is also the first number or the second number mentioned above. When transmitting the kv cache, if the number of tensor parallelisms for prefill and decode is inconsistent, when different GPUs write data to the kv cache of the decoding node, due to the dimension mismatch, they cannot write directly through a single send and receive. Therefore, to improve the transmission efficiency, it is possible to first determine whether the number of the first graphics processors is the same as the number of the second graphics processors, that is, whether the number of tensor parallelisms between prefill and decode is the same.
[0081] If the number of tensor parallelisms between prefill and decode is consistent, that is, the TP size is the same, at this time, it is ensured that the kv cache structures of all nodes in the cluster are consistent, then the key-value information in the target video memory area is directly sent from the first graphics processor to the second target graphics processor. It should be noted that the second target graphics processor is the second target graphics processor corresponding to the decoding node that needs to decode the key-value information stored in the first graphics processor. The second target graphics processor may be one or more.
[0082] By judging whether the number of the first graphics processors is the same as the number of the second graphics processors, it can be ensured that the dimensions of prefill and decode are consistent when transmitting the kv cache, and the correct transmission and reception of the kv cache are guaranteed.
[0083] It should be noted that if TP is not enabled for the inference service (that is, the weight parameters of the target inference model do not need to be split), then the number of GPUs corresponding to one prefill node is one, and the number of GPUs corresponding to one decoding node is also one. At this time, the number of tensor parallelisms between prefill and decode is consistent. Therefore, the key-value information in the target video memory area is also directly sent from the first graphics processor to the second graphics processor.
[0084] If the number of tensor parallelisms between prefill and decode is the same, in the data processing method based on a large model provided in the first embodiment of this application, sending the key-value information in the target video memory area to the second target graphics processor by the first graphics processor includes: determining the second target graphics processor by the first graphics processor based on the mapping relationship between the first graphics processor and the second graphics processor; determining the ID of the first target key-value information in the target video memory area required by the second target graphics processor; reading the first target key-value information from the target video memory area by the first graphics processor based on the ID of the first target key-value information; and sending the first target key-value information to the second target graphics processor by the first graphics processor.
[0085] After sending the first target key-value information to the second target graphics processor by the first graphics processor, receiving the key-value information by the second target graphics processor and determining the video memory address for storing the first target key-value information in the second target graphics processor; and writing the first target key-value information to the video memory address by the second target graphics processor.
[0086] Optionally, before data processing is performed through the prefill node and the decode node, if weight splitting is performed, the mapping relationship between the first graphics processor and the second graphics processor will be determined according to the split weight parameters, that is, the kv cache meta-information list is determined. The kv cache meta-information list is divided into a send list send_meta and a receive list recv_meta. The send list send_meta and the receive list recv_meta include the mapping relationship of Gpurank between sending and receiving and the index information of the kv block (that is, the ID of the kv block to be sent or the ID of the kv block to be received).
[0087] Therefore, when sending the key-value information in the target video memory area to the second target graphics processor by the first graphics processor, the following steps are included: determining the second target graphics processor and the ID of the first target key-value information in the target video memory area required by the second target graphics processor by the first graphics processor based on the mapping relationship between the first graphics processor and the second graphics processor, that is, the sending node (that is, the above-mentioned first graphics processor) specifies the block id index to be sent and the current number of shards and the total number of shards (for handling the sending and receiving when data is split into small pieces due to TP). It should be noted that since TP may cause the kv cache of a block to be split into small pieces, when transmitting, it is necessary to indicate the current number of shards being transmitted and the total number of shards corresponding to the kv cache of the block.
[0088] Then, the first graphics processor executes the corresponding send API to send the kv cache corresponding to the address of the block index stored in itself to the GPU of the receiving rank (i.e., the second target graphics processor mentioned above).
[0089] For the receiving node (i.e., the second target graphics processor mentioned above), it executes the corresponding recv API to receive the tensor from the sending rank GPU (i.e., the first graphics processor mentioned above), and directly writes it to the address corresponding to the block index, that is, the second target graphics processor receives the key-value information and determines the video memory address for storing the first target key-value information in the second target graphics processor; the first target key-value information is written to the video memory address by the second target graphics processor.
[0090] The accurate transmission of the kv cache between the prefill node and the decoding node can be achieved through the mapping relationship between the prefill node and the decoding node.
[0091] If the number of the first graphics processors is different from the number of the second graphics processors, in the data processing method based on large models provided in Embodiment 1 of this application, after determining whether the number of the first graphics processors is the same as the number of the second graphics processors, the method further includes: determining the ID of the second target key-value information to be sent in the first graphics processor; reading the second target key-value information from the target video memory area based on the ID of the second target key-value information through the parallelized copy operator, and writing the second target key-value information into the preset sending buffer in parallel; sending the third target key-value information in the sending buffer to the second graphics processor corresponding to the decoding node through the preset sending function.
[0092] After sending the third target key-value information in the sending buffer to the second graphics processor corresponding to the decoding node through the preset sending function, receiving the third target key-value information through the preset receiving function, and writing the third target key-value information into the preset receiving buffer; determining the fourth target key-value information required by the second graphics processor corresponding to the decoding node; splitting the key-value information in the receiving buffer through the parallelized copy operator to obtain the fourth target key-value information, and writing the fourth target key-value information into the video memory area of the second graphics processor corresponding to the decoding node.
[0093] Optionally, if the number of parallel prefilling and decoding tensors is inconsistent, when different first GPUs write the kv cache to the second GPU in the decoding node, due to dimension mismatch, they cannot be directly written in a single send and receive operation, which will cause the kv cache from prefill to decode to be restricted by discrete memory during transmission. Therefore, to avoid the above problems, when the number of parallel prefilling and decoding tensors is inconsistent, the ID of the second target key-value information to be sent in the first GPU is determined through the send list send_meta, and then the kv cache of the scattered kv blocks is written into the send buffer in the order of send_meta through the parallelized copy operator. It should be noted that the size of the send buffer reserved for the kv cache is determined according to the model size and the gpu video memory resources.
[0094] After completing the parallelized writing, perform the nccl.send atomic operation (i.e., the above-mentioned preset send function) to send the content in the send buffer to the GPU of the receiving rank.
[0095] The receiving node (i.e., the above-mentioned second GPU) performs the nccl.recv atomic operation (i.e., the above-mentioned preset receive function) to receive the kv cache of the sending rank GPU device and write the received kv cache (i.e., the above-mentioned third target key-value information) into the preset receive buffer. After writing, the fourth target key-value information required by the second GPU corresponding to the decoding node is determined according to the kv cache meta information list (recv_meta), and then the continuous kv cache in the receive buffer is split into the video memory area of the corresponding second GPU through the Triton operator (i.e., the above-mentioned parallelized copy operator). It should be noted that if the size of the kv cache tensor to be sent exceeds the size of the reserved send and receive buffer, multiple write + send + receive operations are performed.
[0096] Through the above steps, the low-latency kv cache transmission with consistent or inconsistent numbers of parallel prefilling / decoding tensors can be effectively solved, the transmission efficiency can be improved, and thus the effect of improving the data processing efficiency can be achieved.
[0097] In the data processing method based on a large model provided in the first embodiment of the present application, before processing the description information through the first inference model in the first GPU corresponding to the prefill node, the method further includes: for the first GPU, performing key-value logical block division on the description information according to a preset key-value block size to obtain the key-value logical block identifier of the characters in the description information; determining the key-value logical block identifier corresponding to the key-value information to be generated according to the key-value logical block identifier of the characters in the description information.
[0098] Optionally, before processing the description information through the first inference model in the first graphics processor corresponding to the pre-filled node, the description information is divided into key-value logical blocks according to a preset key-value block size (i.e., the size of the kv block), that is, the key-value logical block identifier corresponding to the token in the description information is determined, and then the key-value logical block identifier corresponding to the key-value information to be generated corresponding to the token is determined, so that the kv cache belonging to the same kv block can be stored in consecutive physical addresses according to the key-value logical block identifier in the subsequent process, that is, the storage addresses of the key-value information in the same key-value logical block identifier are consecutive in the target video memory area.
[0099] It should be noted that the size of the kv block refers to the number of kvs stored in a kv block, and the key-value logical block refers to the logical address corresponding to the kvs belonging to the same kv block. In large model inference, the kv cache corresponds to logical addresses and physical addresses. When storing in the target video memory area in the first graphics processor, it corresponds to the physical address, and the logical address is the address used when operating on the data. Before generating the kv, the logical address corresponding to the kv will be divided first, and then the corresponding physical address will be determined.
[0100] It should be noted that since the cache structure of the kv cache has changed, the corresponding operation operators for reading and calculating the kv cache also need to be changed accordingly. The operation operators use thread blocks (Thread Block) and threads (Thread) to process data in parallel. Each thread calculates the address of the data to be accessed in the global video memory through the thread index (Thread Index) combined with the offset and stride. The thread index is calculated based on the thread ID and the data layout. When the dimension of the data structure changes, the way of calculating this global index should also change accordingly in theory. In addition, in the operation operator, the offset and stride are used to calculate the position of the element. When the dimension or layout of the data structure changes, the formulas for calculating the offset and stride of each element can be adjusted accordingly.
[0101] In an optional embodiment, the transmission of the kv cache can be implemented through a flowchart as shown in Figure 4 and the transmission of the kv cache can be realized through the flowchart as shown in Figure 4As shown, the kv cache cache structure is adjusted to aggregate the kv caches of one kv block into a continuous block, and the operation operators are adjusted according to the adjustment of the kv cache cache structure. Obtain the list of kv cache meta-information for sending and receiving, namely send_meta and recv_meta. Determine whether the number of prefill and decode tensor parallelisms is the same. If the number of prefill and decode tensor parallelisms is the same, the kv block of the prefill node can be directly transmitted through the inplace method (that is, directly operate on the original storage location without establishing a send and receive buffer for the kv cache). If the number of prefill and decode tensor parallelisms is different, establish a send and receive buffer for the kv cache, and then optimize the data reading and writing process through parallel processing. The Triton parallel copy operator writes the scattered kv caches of each block into the send buffer in the order in send_meta. The receiving node receives the kv cache from the sending rank GPU device and directly writes it into the receive buffer. The Triton parallel copy operator disassembles the continuous kv cache in the send and receive buffer and writes it back to the address of the corresponding block index.
[0102] In the data processing method based on a large model provided in the first embodiment of this application, the description information is processed by the first inference model in the first graphics processor corresponding to the prefill node to obtain the key-value information and the first response character output by the first inference model in the first graphics processor; based on the key-value logic block identifier corresponding to the key-value information, the key-value information belonging to the same key-value logic block identifier is stored in the target video memory area in the first graphics processor, where the storage addresses of the key-value information belonging to the same key-value logic block identifier are consecutive in the target video memory area; the key-value information and the first response character in the target video memory area are processed by the second inference model in the second graphics processor corresponding to the decoding node to obtain the target response information corresponding to the description information, which solves the problem in the related art that when pre-filling in the split inference of the large model, the kv output by each layer of the model is stored in a continuous video memory area, and when it is necessary to transfer the kv to the decoding part for decoding, only the kv in one continuous video memory space can be transferred each time, resulting in a delay in transmission and thus a relatively low data processing efficiency. In this solution, after the key-value information output by the first inference model in the first graphics processor corresponding to the prefill node, for the first graphics processor, the key-value information belonging to the same key-value logic block identifier in the key-value information output by the first graphics processor is stored in the video memory area with consecutive addresses. When it is necessary to transfer the kv in a key-value logic block to the decoding part for decoding, it is not necessary to repeatedly read the kv in the video memory area many times, which can effectively reduce the number of transmissions of the kv, improve the transmission efficiency of the kv, and thus achieve the effect of improving the data processing efficiency.
[0103] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0104] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of this application.
[0105] Example 2
[0106] According to an embodiment of the present application, there is also provided a large model-based data processing apparatus for implementing the above-mentioned large model-based data processing method, as Figure 5 shown. The apparatus includes: a first processing unit 501, a storage unit 502, and a second processing unit 503.
[0107] The first processing unit 501 is configured to process the description information input by the target object through a first inference model in a first graphics processor corresponding to a pre-fill node, and obtain key-value information and a first reply character output by the first inference model in the first graphics processor;
[0108] The storage unit 502 is configured to store the key-value information belonging to the same key-value logical block identifier into a target video memory area in the first graphics processor based on the key-value logical block identifier corresponding to the key-value information, where the storage addresses of the key-value information belonging to the same key-value logical block identifier in the target video memory area are continuous;
[0109] The second processing unit 503 is configured to process the key-value information and the first reply character in the target video memory area through a second inference model in a second graphics processor corresponding to a decoding node, and obtain target reply information corresponding to the description information.
[0110] In the data processing device based on a large model provided in the second embodiment of the present application, the first processing unit 501 processes the description information through the first inference model in the first graphics processor corresponding to the pre-fill node to obtain the key-value information and the first reply character output by the first inference model in the first graphics processor; the storage unit 502 stores the key-value information belonging to the same key-value logical block identifier in the target video memory area in the first graphics processor based on the key-value logical block identifier corresponding to the key-value information, wherein the storage addresses of the key-value information belonging to the same key-value logical block identifier in the target video memory area are consecutive; the second processing unit 503 processes the key-value information and the first reply character in the target video memory area through the second inference model in the second graphics processor corresponding to the decoding node to obtain the target reply information corresponding to the description information, which solves the problem in the related art that when pre-filling in the split inference of the large model, the kv output by each layer of the model is stored in a continuous video memory area, and when it is necessary to transmit the kv to the decoding part for decoding, only one continuous video memory space of kv can be transmitted each time, resulting in a delay in transmission and thus a relatively low data processing efficiency. In this solution, after the key-value information output by the first inference model in the first graphics processor corresponding to the pre-fill node, for the first graphics processor, the key-value information belonging to the same key-value logical block identifier in the key-value information output by the first graphics processor is stored in the video memory area with consecutive addresses. When it is necessary to transmit the kv in a key-value logical block to the decoding part for decoding, there is no need to repeatedly read the kv in the video memory area multiple times, which can effectively reduce the transmission times of the kv, improve the transmission efficiency of the kv, and thus achieve the effect of improving the data processing efficiency.
[0111] Optionally, in the data processing device based on a large model provided in the second embodiment of the present application, the device further includes: a first judgment unit, configured to judge whether the weight parameters corresponding to the target inference model need to be split before the first inference model in the first graphics processor corresponding to the pre-fill node processes the description information to obtain the key-value information and the first reply character output by the first inference model in the first graphics processor; a first determination unit, configured to determine the first quantity of the first graphics processors and the second quantity of the second graphics processors required based on the description information if the weight parameters corresponding to the target inference model need to be split; a first splitting unit, configured to split the weight parameters of the target inference model to obtain the first quantity of the first inference models, and determine the first graphics processors based on the first quantity of the first inference models; a second splitting unit, configured to split the weight parameters of the target inference model to obtain the second quantity of the second inference models, and determine the second graphics processors based on the second quantity of the second inference models.
[0112] Optionally, in the data processing device based on a large model provided in the second embodiment of this application, the device further includes: a second determination unit, configured to determine whether the number of first graphics processors is the same as the number of second graphics processors before processing the key-value information and the first reply character in the target video memory area through a second inference model in the second graphics processor corresponding to the decoding node to obtain the target reply information corresponding to the description information; a first sending unit, configured to, if the number of first graphics processors is the same as the number of second graphics processors, send the key-value information in the target video memory area to a second target graphics processor through the first graphics processor, where the second target graphics processor is the second graphics processor corresponding to the decoding node that needs to perform decoding processing on the key-value information stored in the first graphics processor.
[0113] Optionally, in the data processing device based on a large model provided in the second embodiment of this application, the sending unit includes: a first determination module, configured to determine a second target graphics processor through the first graphics processor based on the mapping relationship between the first graphics processor and the second graphics processor; a second determination module, configured to determine the ID of the first target key-value information required by the second target graphics processor in the target video memory area; a reading module, configured to read the first target key-value information from the target video memory area through the first graphics processor based on the ID of the first target key-value information; a sending module, configured to send the first target key-value information to the second target graphics processor through the first graphics processor.
[0114] Optionally, in the data processing device based on a large model provided in the second embodiment of this application, the device further includes: a first receiving unit, configured to, after sending the first target key-value information to the second target graphics processor through the first graphics processor, receive the key-value information through the second target graphics processor and determine the video memory address for storing the first target key-value information in the second target graphics processor; a writing unit, configured to write the first target key-value information to the video memory address through the second target graphics processor.
[0115] Optionally, in the data processing device based on a large model provided in the second embodiment of this application, the device further includes: a second determination unit, configured to, after determining whether the number of first graphics processors is the same as the number of second graphics processors, if the number of first graphics processors is not the same as the number of second graphics processors, determine the ID of the second target key-value information to be sent in the first graphics processor; a reading unit, configured to read the second target key-value information from the target video memory area based on the ID of the second target key-value information through a parallelized copy operator and write the second target key-value information into a preset sending buffer in parallel; a second sending unit, configured to send the third target key-value information in the sending buffer to the second graphics processor corresponding to the decoding node through a preset sending function.
[0116] Optionally, in the data processing device based on a large model provided in the second embodiment of this application, the device further includes: a second receiving unit, configured to, after sending the third target key-value information in the sending buffer to the second graphics processor corresponding to the decoding node through a preset sending function, receive the third target key-value information through a preset receiving function, and write the third target key-value information into a preset receiving buffer; a third determining unit, configured to determine the fourth target key-value information required by the second graphics processor corresponding to the decoding node; a third splitting unit, configured to split the key-value information in the receiving buffer through a parallelized replication operator to obtain the fourth target key-value information, and write the fourth target key-value information into the video memory area of the second graphics processor corresponding to the decoding node.
[0117] Optionally, in the data processing device based on a large model provided in the second embodiment of this application, the device further includes: a partitioning unit, configured to, before processing the description information through the first inference model in the first graphics processor corresponding to the pre-filling node, perform key-value logical block partitioning on the description information according to a preset key-value block size for the first graphics processor, so as to obtain the key-value logical block identifiers of the characters in the description information; a fourth determining unit, configured to determine the key-value logical block identifier corresponding to the key-value information to be generated according to the key-value logical block identifiers of the characters in the description information.
[0118] It should be noted here that the above first processing unit 501, storage unit 502, and second processing unit 503 correspond to steps S201 to S203 in the first embodiment. The functions of the three units and the corresponding steps are the same in terms of the implemented examples and application scenarios, but are not limited to the content disclosed in the above first embodiment. It should be noted that the above units, as part of the device, can run in the computer terminal 10 provided in the first embodiment.
[0119] It should be noted that the preferred implementation schemes involved in the above embodiments of this application are the same as the schemes, application scenarios, and implementation processes provided in the first embodiment, but are not limited to the schemes provided in the first embodiment.
[0120] Embodiment 3
[0121] An embodiment of this application may provide an electronic device, and the electronic device may be any one of the electronic devices in a group of electronic devices. Optionally, in this embodiment, the above electronic device may also be replaced with a terminal device such as a mobile terminal.
[0122] Optionally, in this embodiment, the above electronic device may be located in at least one of multiple network devices in a computer network.
[0123] In this embodiment, the above electronic device may execute the program code of the following steps in the data processing method based on the large model: process the description information through the first inference model in the first graphics processor corresponding to the pre-fill node to obtain the key-value information and the first reply character output by the first inference model in the first graphics processor; based on the key-value logic block identifier corresponding to the key-value information, store the key-value information belonging to the same key-value logic block identifier in the target video memory area in the first graphics processor, where the storage addresses of the key-value information belonging to the same key-value logic block identifier in the target video memory area are continuous; process the key-value information and the first reply character in the target video memory area through the second inference model in the second graphics processor corresponding to the decoding node to obtain the target reply information corresponding to the description information.
[0124] The above electronic device may execute the program code of the following steps in the data processing method based on the large model: before processing the description information through the first inference model in the first graphics processor corresponding to the pre-fill node to obtain the key-value information and the first reply character output by the first inference model in the first graphics processor corresponding to the pre-fill node, the method further includes: determining whether the weight parameters corresponding to the target inference model need to be split; if the weight parameters corresponding to the target inference model need to be split, determining the first quantity of the first graphics processors and the second quantity of the second graphics processors required based on the description information; splitting the weight parameters of the target inference model to obtain the first quantity of the first inference models, and determining the first graphics processors based on the first quantity of the first inference models; splitting the weight parameters of the target inference model to obtain the second quantity of the second inference models, and determining the second graphics processors based on the second quantity of the second inference models.
[0125] The above electronic device may execute the program code of the following steps in the data processing method based on the large model: before processing the key-value information and the first reply character in the target video memory area through the second inference model in the second graphics processor corresponding to the decoding node to obtain the target reply information corresponding to the description information, the method further includes: determining whether the quantity of the first graphics processors is the same as the quantity of the second graphics processors; if the quantity of the first graphics processors is the same as the quantity of the second graphics processors, sending the key-value information in the target video memory area to the second target graphics processor through the first graphics processor, where the second target graphics processor is the second graphics processor corresponding to the decoding node that needs to decode the key-value information stored in the first graphics processor.
[0126] The above electronic device can execute the program code for the following steps in the data processing method based on a large model: Sending key-value information in the target video memory area to a second target graphics processor through a first graphics processor includes: determining the second target graphics processor through the first graphics processor based on the mapping relationship between the first graphics processor and the second graphics processor; determining the ID of the first target key-value information required by the second target graphics processor in the target video memory area; reading the first target key-value information from the target video memory area through the first graphics processor based on the ID of the first target key-value information; and sending the first target key-value information to the second target graphics processor through the first graphics processor.
[0127] The above electronic device can execute the program code for the following steps in the data processing method based on a large model: After sending the first target key-value information to the second target graphics processor through the first graphics processor, the method further includes: receiving the key-value information through the second target graphics processor and determining the video memory address for storing the first target key-value information in the second target graphics processor; and writing the first target key-value information to the video memory address through the second target graphics processor.
[0128] The above electronic device can execute the program code for the following steps in the data processing method based on a large model: After determining whether the number of first graphics processors is the same as the number of second graphics processors, the method further includes: if the number of first graphics processors is not the same as the number of second graphics processors, determining the ID of the second target key-value information to be sent in the first graphics processor; reading the second target key-value information from the target video memory area based on the ID of the second target key-value information through a parallelized copy operator and writing the second target key-value information into a preset send buffer in parallel; and sending the third target key-value information in the send buffer to the second graphics processor corresponding to the decoding node through a preset send function.
[0129] The above electronic device can execute the program code for the following steps in the data processing method based on a large model: After sending the third target key-value information in the send buffer to multiple decoding nodes through a preset send function, the method further includes: receiving the third target key-value information through a preset receive function and writing the third target key-value information into a preset receive buffer; determining the fourth target key-value information required by the second graphics processor corresponding to the decoding node; splitting the key-value information in the receive buffer through a parallelized copy operator to obtain the fourth target key-value information and writing the fourth target key-value information into the video memory area of the second graphics processor corresponding to the decoding node.
[0130] The above electronic device may execute program code for the following steps in the large model-based data processing method: Before processing the description information through the first inference model in the first graphics processor corresponding to the pre-fill node, the method further includes: for the first graphics processor, dividing the description information into key-value logic blocks according to a preset key-value block size to obtain the key-value logic block identifiers of the characters in the description information; determining the key-value logic block identifier corresponding to the key-value information to be generated according to the key-value logic block identifiers of the characters in the description information.
[0131] Optionally, Figure 6 is a structural block diagram of a computer terminal according to an embodiment of the present application. As Figure 6 shown, the electronic device 60 may include: one or more ( Figure 6 only one is shown in the figure) processors 602, a memory 604. The electronic device 60 may further include a storage controller for controlling and managing the memory 604; the electronic device 60 may further include a peripheral interface for connecting a radio frequency module, an audio module, a display screen, etc.
[0132] Among them, the memory may be used to store software programs and modules, such as program instructions / modules corresponding to the large model-based data processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned large model-based data processing method. The memory may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided relative to the processor, and these remote memories may be connected to the electronic device 60 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0133] The processor may call the information and application programs stored in the memory through a transmission device to execute the following steps: processing the description information through the first inference model in the first graphics processor corresponding to the pre-fill node to obtain the key-value information and the first reply character output by the first inference model in the first graphics processor; storing the key-value information belonging to the same key-value logic block identifier in the target video memory area in the first graphics processor based on the key-value logic block identifier corresponding to the key-value information, wherein the key-value information belonging to the same key-value logic block identifier is stored continuously in the target video memory area; processing the key-value information and the first reply character in the target video memory area through the second inference model in the second graphics processor corresponding to the decoding node to obtain the target reply information corresponding to the description information.
[0134] Optionally, the above-mentioned processor may also execute the program code of the following steps: Before processing the description information through the first inference model in the first graphics processor corresponding to the pre-fill node to obtain the key-value information and the first response character output in the first inference model in the first graphics processor corresponding to the pre-fill node, the method further includes: determining whether the weight parameters corresponding to the target inference model need to be split; if the weight parameters corresponding to the target inference model need to be split, determining, based on the description information, a first quantity of the first graphics processors required and a second quantity of the second graphics processors required; splitting the weight parameters of the target inference model to obtain a first quantity of first inference models, and determining the first graphics processors based on the first quantity of first inference models; splitting the weight parameters of the target inference model to obtain a second quantity of second inference models, and determining the second graphics processors based on the second quantity of second inference models.
[0135] Optionally, the above-mentioned processor may also execute the program code of the following steps: Before processing the key-value information and the first response character in the target video memory area through the second inference model in the second graphics processor corresponding to the decoding node to obtain the target response information corresponding to the description information, the method further includes: determining whether the quantity of the first graphics processors is the same as the quantity of the second graphics processors; if the quantity of the first graphics processors is the same as the quantity of the second graphics processors, sending, by the first graphics processor, the key-value information in the target video memory area to a second target graphics processor, where the second target graphics processor is the second graphics processor corresponding to the decoding node that needs to perform decoding processing on the key-value information stored in the first graphics processor.
[0136] Optionally, the above-mentioned processor may also execute the program code of the following steps: Sending, by the first graphics processor, the key-value information in the target video memory area to the second target graphics processor includes: determining, by the first graphics processor, the second target graphics processor based on the mapping relationship between the first graphics processor and the second graphics processor; determining the ID of the first target key-value information required by the second target graphics processor in the target video memory area; reading, by the first graphics processor, the first target key-value information from the target video memory area based on the ID of the first target key-value information; and sending, by the first graphics processor, the first target key-value information to the second target graphics processor.
[0137] Optionally, the above-mentioned processor may also execute the program code of the following steps: After sending, by the first graphics processor, the first target key-value information to the second target graphics processor, the method further includes: receiving, by the second target graphics processor, the key-value information and determining the video memory address for storing the first target key-value information in the second target graphics processor; and writing, by the second target graphics processor, the first target key-value information to the video memory address.
[0138] Optionally, the above-mentioned processor may also execute the program code of the following steps: After determining whether the number of first graphics processors is the same as the number of second graphics processors, the method further includes: If the number of first graphics processors is different from the number of second graphics processors, determine the ID of the second target key value information to be sent in the first graphics processor; Read the second target key value information from the target video memory area based on the ID of the second target key value information through a parallelized copy operator, and write the second target key value information into a preset sending buffer in parallel; Send the third target key value information in the sending buffer to the second image processor corresponding to the decoding node through a preset sending function.
[0139] Optionally, the above-mentioned processor may also execute the program code of the following steps: After sending the third target key value information in the sending buffer to multiple decoding nodes through a preset sending function, the method further includes: Receive the third target key value information through a preset receiving function, and write the third target key value information into a preset receiving buffer; Determine the fourth target key value information required by the second graphics processor corresponding to the decoding node; Split the key value information in the receiving buffer through a parallelized copy operator to obtain the fourth target key value information, and write the fourth target key value information into the video memory area of the second graphics processor corresponding to the decoding node.
[0140] Optionally, the above-mentioned processor may also execute the program code of the following steps: Before processing the description information through the first inference model in the first graphics processor corresponding to the pre-fill node, the method further includes: For the first graphics processor, perform key value logic block division on the description information according to a preset key value block size to obtain the key value logic block identifier of the characters in the description information; Determine the key value logic block identifier corresponding to the key value information to be generated based on the key value logic block identifier of the characters in the description information.
[0141] Those of ordinary skill in the art can understand that Figure 6 The structure shown is only for illustration, and the electronic device 60 may also be a smart phone (such as an Android phone, an IOS phone, etc.), a tablet computer, a handheld computer, and a mobile Internet device (Mobile Internet Devices, MID), a PAD and other terminal devices. Figure 6 It does not limit the structure of the above-mentioned electronic device. For example, the electronic device 60 may further include more or fewer components than those shown in Figure 6 (such as a network interface, a display device, etc.), or have a different configuration from that shown in Figure 6 shown.
[0142] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing the relevant hardware of the terminal device. This program can be stored in a computer-readable storage medium, which can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc.
[0143] Embodiment 4
[0144] An embodiment of the present application also provides a computer-readable storage medium. Optionally, in this embodiment, the above computer program product can be used to store the program code executed by the data processing method based on the large model provided in the first embodiment.
[0145] Embodiment 5
[0146] An embodiment of the present application also provides a computer program product. Optionally, in this embodiment, the above computer program product can be used to store the program code executed by the data processing method based on the large model provided in the first embodiment.
[0147] Optionally, in this embodiment, the above computer program product can be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0148] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0149] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0150] In the several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the units or modules can be in electrical or other forms.
[0151] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0152] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0153] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: USB flash drive, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disc and other various media that can store program codes.
[0154] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A data processing method based on a large model, characterized in that: include: Processing the description information input by the target object through the first reasoning model in the first graphics processor corresponding to the pre-filled node to obtain the key value information and the first reply character output by the first reasoning model in the first graphics processor; Based on the key-value logical block identifier corresponding to the key-value information, the key-value information belonging to the same key-value logical block identifier is stored in a target video memory area in the first graphics processor, wherein the storage addresses of the key-value information belonging to the same key-value logical block identifier in the target video memory area are continuous; Processing the key information and the first reply character in the target video memory area through a second inference model in a second graphics processor corresponding to the decoding node to obtain target reply information corresponding to the description information; The first reasoning model and the second reasoning model are determined by a target reasoning model, and the target reasoning model is a large model.
2. The method according to claim 1, characterized in that Before the description information is processed by the first reasoning model in the first graphics processor corresponding to the pre-filled node to obtain the key value information and the first reply character output by the first reasoning model in the first graphics processor, the method further includes: Determine whether the weight parameters corresponding to the target inference model need to be split; If the weight parameters corresponding to the target reasoning model need to be split, determining a first number of first graphics processors required and a second number of second graphics processors required based on the description information; Splitting the weight parameters of the target inference model to obtain a first number of first inference models, and determining the first graphics processor based on the first number of first inference models; The weight parameters of the target inference model are split to obtain a second number of second inference models, and the second graphics processor is determined based on the second number of second inference models.
3. The method according to claim 1, characterized in that Before processing the key value information and the first reply character in the target video memory area by a second inference model in a second graphics processor corresponding to the decoding node to obtain target reply information corresponding to the description information, the method further includes: Determining whether the number of the first graphics processors is the same as the number of the second graphics processors; If the number of the first graphics processors is the same as the number of the second graphics processors, the key value information in the target video memory area is sent to the second target graphics processor via the first graphics processor, wherein the second target graphics processor is the second graphics processor corresponding to the decoding node that needs to perform decoding processing on the key value information stored in the first graphics processor.
4. The method according to claim 3, characterized in that Sending the key value information in the target video memory area to the second target graphics processor through the first graphics processor includes: Determining the second target graphics processor by the first graphics processor based on a mapping relationship between the first graphics processor and the second graphics processor; Determine the ID of the first target key value information in the target video memory area required by the second target graphics processor; Reading the first target key value information from the target video memory area based on the ID of the first target key value information by the first graphics processor; The first target key value information is sent to the second target graphics processor through the first graphics processor.
5. The method according to claim 4, characterized in that After sending the first target key value information to the second target graphics processor through the first graphics processor, the method further includes: receiving the key value information through the second target graphics processor and determining a video memory address in the second target graphics processor for storing the first target key value information; The first target key value information is written into the video memory address through the second target graphics processor.
6. The method according to claim 3, characterized in that After determining whether the number of the first graphics processors is the same as the number of the second graphics processors, the method further includes: If the number of the first graphics processors is different from the number of the second graphics processors, determining the ID of the second target key-value information to be sent in the first graphics processor; Reading the second target key value information from the target video memory area based on the ID of the second target key value information through a parallelized replication operator, and writing the second target key value information into a preset sending buffer area in parallel; The third target key value information in the sending buffer is sent to the second graphics processor corresponding to the decoding node through a preset sending function.
7. The method according to claim 6, characterized in that After sending the third target key value information in the sending buffer to the second graphics processor corresponding to the decoding node through a preset sending function, the method further includes: Receiving the third target key value information through a preset receiving function, and writing the third target key value information into a preset receiving buffer area; Determine fourth target key value information required by a second graphics processor corresponding to the decoding node; The key value information in the receiving buffer area is split by a parallelized copy operator to obtain the fourth target key value information, and the fourth target key value information is written into a graphics memory area in the second graphics processor corresponding to the decoding node.
8. The method according to claim 1, characterized in that: Before processing the description information by using a first inference model in a first graphics processor corresponding to the pre-filled node, the method further includes: For the first graphics processor, dividing the description information into key value logic blocks according to a preset key value block size to obtain key value logic block identifiers of characters in the description information; According to the key value logical block identifier of the character in the description information, the key value logical block identifier corresponding to the key value information to be generated is determined.
9. A data processing device based on a large model, characterized in that: include: A first processing unit, configured to process the description information inputted by the target object through a first reasoning model in a first graphics processor corresponding to the pre-filled node, and obtain key value information and a first reply character outputted by the first reasoning model in the first graphics processor; a storage unit, configured to store the key-value information belonging to the same key-value logical block identifier into a target video memory area in the first graphics processor based on the key-value logical block identifier corresponding to the key-value information, wherein the storage addresses of the key-value information belonging to the same key-value logical block identifier in the target video memory area are continuous; A second processing unit is used to process the key value information and the first reply character in the target video memory area through a second inference model in a second graphics processor corresponding to the decoding node to obtain target reply information corresponding to the description information; The first reasoning model and the second reasoning model are determined by a target reasoning model, and the target reasoning model is a large model.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the storage medium is located is controlled to execute the large model-based data processing method according to any one of claims 1 to 8.
11. An electronic device, characterized in that: include: A memory storing an executable program; A processor is used to run the program, wherein the program, when running, executes the large model-based data processing method described in any one of claims 1 to 8.
12. A computer program product, characterized in that The computer program product comprises a stored computer program, and when the computer program is executed by a processor, the large model-based data processing method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Video memory management method and device of large language model, electronic equipment and storage medium
CN118094037A
Large model reasoning acceleration method and system combining machine learning and speculation sampling
CN118657220A