Accelerator card deployment method, device, equipment, storage medium and program product
By obtaining the storage usage of the target model and building multiple accelerator card topology architectures, the hardware resource mismatch problem caused by the fixed configuration of the accelerator card topology structure is solved, and the optimized configuration of accelerator card resources and the improvement of model performance are achieved.
Patent Information
- Application Number
- CN202510897024.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In the existing technology, the topology structure of the accelerator card is a fixed configuration, which leads to a mismatch between the hardware resource configuration of the model and the computing requirements, resulting in performance degradation, increased latency, decreased throughput, and the model being unable to operate normally.
By obtaining the storage usage of the target model, determining the required number of accelerator cards, and building multiple accelerator card topology architectures, we select the topology architecture that meets the preset conditions to deploy the accelerator cards, ensuring that hardware resources match computing requirements.
The storage capacity of the accelerator card is matched with the storage usage of the model, avoiding resource waste or shortage, ensuring that the model performance meets the expected requirements, reducing latency and improving throughput.
Smart Images

Figure CN120407200B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an accelerator card deployment method, apparatus, device, storage medium, and program product. Background Art
[0002] Accelerator cards (such as graphics cards and SmartNet cards) are the core hardware supporting model operations. With the development of technologies like Large Language Models (LLMs), the number of model parameters is increasing rapidly. It's common for a single accelerator card to be unable to load all the model parameters. Therefore, a single model needs to be split across multiple accelerator cards for execution.
[0003] Currently, some technologies use a fixed configuration of the number of accelerator cards and the topology between them in the servers used to run models. This is not a rational approach and can lead to a mismatch between the model's hardware resources and computing requirements, which can degrade model performance, causing increased latency, decreased throughput, and even model failures. Summary of the Invention
[0004] The present application provides an accelerator card deployment method, an accelerator card deployment device, an electronic device, a computer non-volatile readable storage medium, and a computer program product to at least solve the problem of mismatch between the hardware resource configuration of the model and the computing requirements in the related art.
[0005] This application provides an accelerator card deployment method, including:
[0006] Get the storage usage of the target model during operation;
[0007] Determining the number of accelerator cards required to run the target model based on the storage occupancy and the storage capacity of each accelerator card;
[0008] Constructing multiple accelerator card topology architectures based on the number of accelerator cards, wherein different accelerator card topology architectures have different connection modes between the accelerator cards;
[0009] Determine model performance metrics when running the target model using various accelerator card topology architectures;
[0010] If there is a first accelerator card topology architecture whose model performance indicators meet preset conditions, the number of servers used to run the target model is determined based on the first accelerator card topology architecture, and accelerator cards are deployed in the servers used to run the target model.
[0011] The present application also provides an accelerator card deployment device, comprising:
[0012] A storage occupancy acquisition module is used to obtain the storage occupancy during the operation of the target model;
[0013] an accelerator card quantity determination module, configured to determine the number of accelerator cards required to run the target model based on the storage occupancy and the storage capacity of each accelerator card;
[0014] A topology construction module is used to construct multiple accelerator card topology architectures according to the number of accelerator cards, wherein the connection methods between the accelerator cards are different in different accelerator card topology architectures;
[0015] A calculation module, configured to determine a model performance indicator when running the target model using each accelerator card topology architecture;
[0016] An accelerator card deployment module is used to determine the number of servers used to run the target model and deploy accelerator cards in the servers used to run the target model based on a first accelerator card topology architecture if there is a first accelerator card topology architecture whose model performance indicators meet preset conditions.
[0017] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of the above-mentioned accelerator card deployment method when executing the computer program.
[0018] The present application also provides a computer non-volatile readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned accelerator card deployment method are implemented.
[0019] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned accelerator card deployment method when executed by a processor.
[0020] In the technical solutions of some embodiments of the present application, on the one hand, the number of accelerator cards required to run the target model is determined based on the storage occupancy during the operation of the target model. In this way, the sum of the storage capacity of the accelerator cards can be guaranteed to match the storage occupancy of the target model, thereby avoiding the problem of waste or insufficient accelerator card resources. On the other hand, based on the number of accelerator cards, multiple accelerator card topology architectures are constructed, and based on the model performance indicators of each accelerator card topology architecture when running the target model, the first accelerator card topology architecture whose model performance indicators meet the preset conditions is selected as the architecture reference for deploying the accelerator card. In this way, when the target model is run in the accelerator card, the performance of the target model can be guaranteed to meet the expected performance requirements. Based on the above two aspects, the hardware resource configuration of the target model can be matched with the computing requirements, thereby solving the problem of mismatch between the hardware resource configuration of the model and the computing requirements in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 This is a diagram of the architecture of some servers used to run large language models;
[0023] Figure 2 A schematic diagram of one of the accelerator card topologies;
[0024] Figure 3 A schematic diagram of another accelerator card topology;
[0025] Figure 4 A schematic diagram of another accelerator card topology;
[0026] Figure 5 A schematic diagram of another accelerator card topology;
[0027] Figure 6 A schematic diagram of a process for deploying an accelerator card according to some embodiments of the present application;
[0028] Figure 7 A schematic diagram of a module of an accelerator card deployment device provided in some embodiments of the present application;
[0029] Figure 8 A schematic diagram of a module of an electronic device provided for some embodiments of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0031] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or precedence.
[0032] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0033] See also Figure 1 , which is a schematic diagram of the architecture of some servers used to run large language models. Figure 1 In this example, a server includes a central processing unit (CPU) and multiple accelerator cards. These accelerator cards are connected to the CPU, and each accelerator card includes a storage unit and a core computing unit. The accelerator cards are connected via a communication bus. This communication bus can be a PCIe (Peripheral Component Interconnect Express) bus, an NVLink bus, or other similar bus.
[0034] When running a large language model on a server, if the number of parameters exceeds the storage capacity of a single accelerator card, the CPU can split the model across multiple accelerator cards. For ease of understanding, the following two examples of model splitting are provided: 1) and 2).
[0035] 1) Each accelerator card runs one or more neural network layers of a large language model, i.e., pipeline parallelism.
[0036] For example, assuming the large language model includes 10 neural network layers, the model can be split as follows: save the parameters of the 1st to 3rd neural network layers of the large language model to storage unit S1 of accelerator card G1, save the parameters of the 4th to 8th neural network layers of the large language model to storage unit S2 of accelerator card G2, and save the parameters of the 9th to 10th neural network layers of the large language model to storage unit S3 of accelerator card G3. In this way, the core computing unit C1 in accelerator card 1 can load the parameters of the 1st to 3rd neural network layers of the large language model from storage unit S1 and run the 1st to 3rd neural network layers of the large language model; the core computing unit C2 in accelerator card 2 can load the parameters of the 4th to 8th neural network layers of the large language model from storage unit S2 and run the 4th to 8th neural network layers of the large language model; and the core computing unit C3 in accelerator card 3 can load the parameters of the 9th to 10th neural network layers of the large language model from storage unit S3 and run the 9th to 10th neural network layers of the large language model.
[0037] When using a large language model for inference, the CPU can preprocess the text sequence to be inferred (such as a question that the large language model needs to answer) to map the text sequence into the feature vector required by the large language model (i.e., perform embedding processing on the text sequence). After preprocessing, the CPU can save the feature vector of the text sequence to storage unit S1 of accelerator card G1. Accelerator card G1's core computing unit C1 reads the feature vector and model parameters from storage unit S1, calculates the outputs of the first through third neural network layers, and saves the output of the third neural network layer (i.e., the intermediate processing result) to storage unit S2 of accelerator card G2 via the inter-accelerator communication bus. Accelerator card G2's core computing unit C2 reads the model parameters and the output of the third neural network layer from storage unit S2, continues to calculate the outputs of the fourth through eighth neural network layers, and saves the output of the eighth neural network layer to storage unit S3 of accelerator card G3 via the inter-accelerator communication bus. Following a similar principle as accelerator card G2, accelerator card G3 calculates the outputs of the ninth and tenth neural network layers. The output of the 10th neural network layer can be considered the inference result (i.e., the answer to the question). The core computing unit C3 in the accelerator card G3 returns the inference result to the central processing unit, thus completing the inference of the large language model.
[0038] 2) Divide the parameters of the same neural network layer into different accelerator cards for execution, i.e. tensor parallelism.
[0039] For example, assuming that the large language model includes 10 neural network layers, the model can be split as follows: the parameters of the first neural network layer of the large language model are divided into subsets 11 and 12, and the parameters in subset 11 are saved to the storage unit S11 of the accelerator card G11, and the parameters in subset 12 are saved to the storage unit S12 of the accelerator card G12. The parameters of the second neural network layer of the large language model are divided into subsets 21, 22, and 23, and the parameters in subset 21 are saved to the storage unit S21 of the accelerator card G21, the parameters in subset 22 are saved to the storage unit S22 of the accelerator card G22, and the parameters in subset 23 are saved to the storage unit S23 of the accelerator card G23. And so on.
[0040] When using a large language model for inference, the central processing unit can save the feature vector of the text sequence to be inferred to the accelerator card running the first neural network layer, that is, save the feature vector to the storage unit S11 of the accelerator card G11 and the storage unit S12 of the accelerator card G12. The core computing unit C11 of the accelerator card G11 reads the subset 11 of the model parameters and the feature vector from the storage unit S11, and obtains the calculation result Y11 after calculation, and the core computing unit C12 of the accelerator card G12 reads the subset 12 of the model parameters and the feature vector from the storage unit S12, and obtains the calculation result Y12 after calculation. Based on the communication bus between the accelerator cards, the accelerator card G11 can synchronize the calculation result Y11 to the accelerator card G12, and the accelerator card G12 can synchronize the calculation result Y12 to the accelerator card G11. The accelerator card G11 combines the calculation results Y11 and Y12 to obtain the complete calculation results of the first neural network layer. Similarly, the accelerator card G12 combines the calculation results Y11 and Y12 to obtain the complete calculation results of the first neural network layer.
[0041] For ease of understanding, the following is explained with examples. For example, assuming that the parameters of the first neural network layer include a weight matrix M, the weight matrix M can be split into sub-matrices M1 and M2. Among them, the sub-matrix M1 can be regarded as a subset 11 of the above-mentioned model parameters, and the sub-matrix M2 can be regarded as a subset 12 of the above-mentioned model parameters. The sub-matrix M1 can be saved in the storage unit S11 of the accelerator card G11, and the sub-matrix M2 can be saved in the storage unit S12 of the accelerator card G12. The core computing unit C11 of the accelerator card G11 can read the sub-matrix M1 and the eigenvector from the storage unit S11, and perform calculations based on the sub-matrix M1 and the eigenvector to obtain the calculation result Y11, and the core computing unit C12 of the accelerator card G12 can read the sub-matrix M2 and the eigenvector from the storage unit S12, and perform calculations based on the sub-matrix M2 and the eigenvector to obtain the calculation result Y12. Based on the communication bus, the accelerator cards G11 and G12 synchronize and merge the calculation results Y11 and Y12 to obtain the complete calculation results of the weight matrix M and the eigenvector, which can be used as the output of the first neural network layer.
[0042] Furthermore, core computing units C11 and C12 can save the output of the first neural network layer to the accelerator card running the second neural network layer, namely, to storage unit S21 of accelerator card G21, storage unit S22 of accelerator card G22, and storage unit S23 of accelerator card G23. Similar to the principle of the first neural network layer, core computing unit C21 of accelerator card G21 can read subset 21 of model parameters and the output of the first neural network layer from storage unit S21 to obtain calculation result Y21. Core computing unit C22 of accelerator card G22 can read subset 22 of model parameters and the output of the first neural network layer from storage unit S22 to obtain calculation result Y22. Core computing unit C23 of accelerator card G23 can read subset 23 of model parameters and the output of the first neural network layer from storage unit S23 to obtain calculation result Y23. Based on the communication bus, accelerator cards G21, G22, and G23 can synchronize their calculation results with each other to obtain the complete output of the second neural network layer. Among them, when the parameters of the same neural network layer are divided into three or more accelerator cards, the method of merging the calculation results in segments can be used to obtain the complete output of the corresponding neural network layer. For example, the calculation results between the accelerator card G21 and the accelerator card G22 are synchronized to obtain the intermediate calculation result Y1, the calculation results between the accelerator card G21 and the accelerator card G23 are synchronized to obtain the intermediate calculation result Y2, and the intermediate calculation results Y1 and Y2 are synchronized between the accelerator card G21 and the accelerator card G23 to obtain the complete output Y of the second neural network layer. Compared with merging all calculation results in the same accelerator card, this method of merging calculation results in segments can reduce the communication pressure between accelerator cards. For example, if the accelerator cards G21 and G22 both send the calculation results to the accelerator card G23 for merging, then the accelerator card G23 may have a communication bottleneck, thereby reducing the efficiency of merging the calculation results.
[0043] Based on the above description, it's clear that while splitting a large language model across multiple accelerator cards can solve the problem of a single accelerator card being unable to load all model parameters, it also introduces communication overhead and synchronization issues. This means that the performance of the large language model is affected by both the accelerator card's computation speed and the communication and synchronization time between accelerator cards. Model performance metrics may include, but are not limited to, latency and throughput of the large language model. Latency refers to the time it takes for the large language model to output a complete inference result after inputting the text sequence to be inferred. Throughput refers to the number of tokens output per second by the large language model. Generally, the more accelerator cards there are, the fewer computational tasks each accelerator card can perform, which can increase computation speed and shorten computation time. However, with more accelerator cards, communication and synchronization time also increases. Therefore, selecting the right number of accelerator cards is a key factor in reducing model latency and improving model throughput. Furthermore, given the same number of accelerator cards, different accelerator card topologies will result in different communication and synchronization times. Therefore, selecting the right accelerator card topology is also a key factor in reducing model latency and improving model throughput. For ease of understanding, the following example illustrates this. Assume there are 8 accelerator cards. Figures 2 to 4 , which is a schematic diagram of different topological structures composed of 8 accelerator cards.
[0044] Figure 2 In the example, eight accelerator cards are deployed in the same server A. Server A includes a central processor A1 and a switch device A2. The eight accelerator cards are connected to the central processor A1 through the switch device A2, and the accelerator cards are connected to each other through a communication bus. Figure 2 In the accelerator card topology shown, CPU A1 can split the large language model and store different model parameters on different accelerator cards. When performing inference tasks on the large language model, CPU A1 preprocesses the text sequence to be inferred, obtains the feature vector of the text sequence, and stores the feature vector on the accelerator card running the first neural network layer. Accelerator cards can execute inference tasks in a pipelined and parallel manner, communicating and synchronizing data via a communication bus.
[0045] Figure 3In the example, 8 accelerator cards are deployed in the same server B, and the 8 accelerator cards are divided into two groups, each group includes 4 accelerator cards. Server B includes central processing units B11 and B21 and switching devices B12 and B22. One group of accelerator cards is connected to the central processing unit B11 through the switching device B12, and the other group of accelerator cards is connected to the central processing unit B21 through the switching device B22. The two groups of accelerator cards are not directly connected through a communication bus. The accelerator cards in the same group are connected through a communication bus. The central processing units B11 and B21 are connected through a communication bus. Based on Figure 3 In the accelerator card topology shown, the central processors B11 and B21 can split the large language model and save different model parameters in different accelerator cards. For example, the central processor B11 can split the parameters of the 1st to 6th neural network layers of the large language model into 4 subsets 11, 12, 13, and 14, and save subset 11 to the first connected accelerator card, and save subset 12 to the second connected accelerator card, and so on. Similarly, the central processor B21 can split the parameters of the 7th to 14th neural network layers of the large language model into 4 subsets, and save the parameters in these 4 subsets to the 4 connected accelerator cards. When the large language model performs inference tasks, the central processor B11 or the central processor B12 can preprocess the text sequence to be inferred, obtain the feature vector of the text sequence, and save the feature vector in the accelerator card running the first neural network layer. The accelerator cards can perform inference tasks in a pipeline parallel or tensor parallel manner. During the execution of inference tasks, accelerator cards in the same group can communicate and synchronize data through the communication bus, and accelerator cards in different groups can communicate and synchronize data through the communication bus between the central processors B11 and B21.
[0046] Figure 4 and Figure 3 They are basically similar, with the main difference being that the two groups of accelerator cards are connected to the same central processor C11 through switching devices C12 and C22. The switching devices C12 and C22 are connected via a communication bus, and the central processors C11 and C21 are connected via a communication bus. The central processor C11 is used to split the large language model and save the split model parameters to the two groups of accelerator cards. When the large language model performs inference tasks, the central processor C11 can preprocess the text sequence to be inferred, obtain the feature vector of the text sequence, and save the feature vector in the accelerator card running the first neural network layer. The accelerator cards can perform inference tasks in a pipeline parallel or tensor parallel manner. During the execution of the inference task, the accelerator cards in the same group can communicate and synchronize data through the communication bus, and the accelerator cards in different groups can communicate and synchronize data through the communication bus between the switching devices C12 and C22.
[0047] Figure 5 and Figure 3 Basically similar, the main difference is that, of the two groups of accelerator cards obtained, one group of accelerator cards is deployed in server E, and the other group of accelerator cards is deployed in server F. Server E includes a central processing unit E11 and a switching device E12, and the accelerator card in server E is connected to the central processing unit E11 through the switching device E12. Server F includes a central processing unit F11 and a switching device F12, and the accelerator card in server F is connected to the central processing unit F11 through the switching device F12. The central processing units E11 and F11 are connected via an inter-server communication bus. When executing inference tasks in a large language, accelerator cards in the same group communicate and synchronize data via the intra-group communication bus, and accelerator cards in different groups communicate and synchronize data via the inter-server communication bus.
[0048] It is understandable that due to the differences between the topologies, the above Figures 2 to 5 When running the same large language model using the accelerator card topology shown, the latency and throughput can vary. Therefore, a reasonable accelerator card topology is also a key factor in reducing model latency and improving model throughput.
[0049] Currently, when some server vendors sell servers for model operation, the number of accelerator cards and the topology of the accelerator cards in the server are fixed. For example, the number of accelerator cards is 10, and they are all based on Figure 2 The topology shown in the figure is used for deployment. Since the parameter quantity and structure of different large language models may be different, if the number of accelerator cards and the accelerator card topology are fixed, the hardware resource configuration of the large language model may not match the computing requirements, which may lead to performance degradation of the large language model, increased latency, decreased throughput, and model failure. Figure 2 The topology shown is illustrated by taking the deployment of 10 accelerator cards in the server as an example.
[0050] For example, suppose customer T1 needs to purchase a server for running the large language model D1. The large language model D1 has a large number of parameters and actually requires 12 accelerator cards. However, since the server is only configured with 10 accelerator cards or has only 10 accelerator card slots, customer T1 may not be able to run the large language model D1 normally based on the purchased server.
[0051] For example, suppose customer T2 needs to purchase a server for running the large language model D2. The large language model D2 has a relatively small number of parameters and actually requires five accelerator cards. However, since 10 accelerator cards are actually deployed in the server, there is a possibility of wasting accelerator card resources.
[0052] For example, suppose customer T3 needs to purchase a server for running the large language model D3, and the parameter volume of the large language model D3 matches 10 accelerator cards, but these 10 accelerator cards need to be configured as described above. Figure 3 Only when the topology shown is deployed can the large language model D3 achieve optimal performance. In this case, the performance of the large language model D3 cannot be optimized based on the purchased server.
[0053] In view of this, the present application provides an accelerator card deployment method, which can obtain the hardware resource configuration that is compatible with the computing requirements of the large language model (i.e., the number of accelerator cards in the server, the topology between the accelerator cards) based on the relevant information of the large language model to be run on the server before the server leaves the factory, and configure the server according to the hardware resource configuration that is compatible with the computing requirements of the large language model. In this way, after the server is sold to the customer, the customer can achieve better performance of the large language model based on the purchased server. The accelerator card deployment method can be applied to electronic devices. Electronic devices may include but are not limited to tablet computers, laptop computers, desktop computers, etc. In conjunction with reference Figure 6 , which is a flow chart of the acceleration card deployment method provided in some embodiments of the present application. Figure 6 The accelerator card deployment method includes the following steps:
[0054] Step S601: Obtain the storage usage during the operation of the target model.
[0055] In this embodiment, the target model is the large language model that needs to be run. Figure 1 From the relevant description, we can know that during the operation of the target model, the model parameters of the target model and the intermediate operation results (such as activation values, gradients, etc.) generated during the operation of the target model need to occupy the storage space of the accelerator card. Therefore, the parameter quantity, model structure, inference logic, etc. of the target model can be analyzed to obtain the storage usage during the operation of the target model.
[0056] In the subsequent embodiments of this application, taking the model based on the self-attention mechanism as an example, it is explained how to obtain the storage occupancy during the operation of the target model, which will not be repeated here.
[0057] Step S602: Determine the number of accelerator cards required to run the target model based on the storage usage and the storage capacity of each accelerator card.
[0058] Specifically, the number of accelerator cards required to run the target model can be determined based on the principle that the sum of the storage capacities of the multiple accelerator cards is greater than or equal to the total memory usage. The storage capacities of the multiple accelerator cards can be the same or different. For example, if the target model's memory usage is 50GB, and each accelerator card has a storage capacity of 10GB, then the number of accelerator cards required to run the target model must be at least five.
[0059] Step S603: construct multiple acceleration card topology architectures based on the number of acceleration cards. In different acceleration card topology architectures, the connection methods between the acceleration cards are different.
[0060] Specifically, the multiple accelerator card topologies constructed include at least two of the following topologies:
[0061] 1) All accelerator cards are connected to the same central processor, and the accelerator cards are connected via the first communication bus (i.e. Figure 2 topology shown).
[0062] 2) The accelerator cards are divided into at least two groups. The accelerator cards of different groups are connected to different central processing units of the same server. The central processing units connected to the accelerator cards of different groups are connected via the second communication bus. The accelerator cards of the same group are connected via the third communication bus. The accelerator cards of different groups are connected via the second communication bus (i.e. Figure 3 topology shown).
[0063] 3) The accelerator cards are divided into at least two groups. The accelerator cards of different groups are connected to the same central processor of the same server. The accelerator cards of the same group are connected via the fourth communication bus, and the accelerator cards of different groups are connected via the fifth communication bus (i.e. Figure 4 topology shown).
[0064] 4) The accelerator cards are divided into at least two groups. Accelerator cards of different groups are set in different servers. The servers are connected through the sixth communication bus (i.e. Figure 5 topology shown).
[0065] For example, if the number of accelerator cards required to run the target model is 6, you can follow Figure 2 The architecture shown in the figure is used to construct an accelerator card topology architecture E1 including 6 accelerator cards. At the same time, the 6 accelerator cards can be divided into two groups, each group including 3 accelerator cards, and Figure 3 The architecture shown in the figure is used to construct the accelerator card topology E2, and Figure 4 The architecture shown in the figure is used to build the accelerator card topology E3, and Figure 5 The architecture shown is used to construct the accelerator card topology architecture E4.
[0066] Of course, it is understandable that the accelerator card topology architecture can be more than just Figures 2 to 5 The topology shown. In actual applications, the accelerator card topology can be constructed according to actual needs. For example, assuming that the number of accelerator cards required to run the target model is 12, the accelerator cards can be divided into 3 groups, the first group includes 2 accelerator cards, the second group includes 4 accelerator cards, and the third group includes 6 accelerator cards. Among them, the accelerator cards of the first group are deployed in server S1, and all the accelerator cards in the first group are connected to the same central processing unit in server S1. The accelerator cards of the second group are deployed in server S2, and all the accelerator cards in the second group are connected to the same central processing unit in server S2. The accelerator cards of the third group are deployed in server S3, and the accelerator cards in the third group are divided into two subgroups, the accelerator cards in the first subgroup are connected to the central processing unit U1 in server S3, and the accelerator cards in the second subgroup are connected to the central processing unit U2 in server S3. This application does not limit the specific form of the accelerator card topology.
[0067] Step S604: determining the model performance indicators when the target model is run using each accelerator card topology architecture.
[0068] In this embodiment, the model performance indicator may include at least one of the target model's latency and throughput. In other embodiments, the model performance indicator may also include other indicators besides latency and throughput, such as the target model's hardware resource utilization and queue delay (i.e., the length of time a request waits for execution). The following description uses the example of the model performance indicator including the target model's latency and throughput.
[0069] based on Figures 2 to 5 It can be seen that since the connection methods between accelerator cards are different in different accelerator card topology architectures, even if the number of accelerator cards is the same, the communication and synchronization durations in different accelerator card topologies may be different. Furthermore, the delay duration and throughput when running the same target model in different accelerator card topologies may be different.
[0070] Specifically, when calculating the latency and throughput of each accelerator card topology running the target model, the calculation can be based on the bandwidth of each accelerator card topology, the accelerator card computing power, the amount of data to be transmitted or loaded, etc. In subsequent embodiments of this application, using a model based on the self-attention mechanism as an example, we explain how to calculate the latency and throughput corresponding to each accelerator card topology, which will not be repeated here.
[0071] Step S605: If there is a first accelerator card topology architecture whose model performance indicators meet the preset conditions, the number of servers used to run the target model is determined based on the first accelerator card topology architecture, and accelerator cards are deployed in the servers used to run the target model.
[0072] Specifically, the model performance indicators include the target model's latency and throughput as an example. The preset conditions refer to the expected latency and throughput that the target model needs to achieve. For any accelerator card topology, if the latency corresponding to that accelerator card topology is lower than the expected latency and the throughput is higher than the expected throughput, that accelerator card topology can be considered the first accelerator card topology architecture that meets the preset conditions.
[0073] For example, if the expected latency is 30 milliseconds and the expected throughput is 20 tokens / second, the latency and throughput corresponding to the accelerator card topologies E1 to E4 are as follows:
[0074] Accelerator card topology E1: latency 20 milliseconds, throughput 25 tokens / second;
[0075] Accelerator card topology E2: 25 milliseconds latency, 17 tokens / second throughput;
[0076] Accelerator card topology E3: 35 milliseconds latency, 10 tokens / second throughput;
[0077] Accelerator card topology E4: latency is 40 milliseconds and throughput is 5 tokens / second.
[0078] In the above example, since only the latency and throughput corresponding to accelerator card topology E1 meet the preset conditions, the number of servers used to run the target model and the deployment of accelerator cards within these servers can be determined based on accelerator card topology E1. For example, since there is only one server in accelerator card topology E1, the number of servers running the target model can be determined as 1, and the connection relationships between accelerator cards within this single server can be configured based on accelerator card topology E1.
[0079] In summary, in the technical solutions of some embodiments of the present application, on the one hand, the number of accelerator cards required to run the target model is determined based on the storage occupancy during the operation of the target model. In this way, the sum of the storage capacity of the accelerator cards can be guaranteed to match the storage occupancy of the target model, thereby avoiding the problem of waste or insufficient accelerator card resources. On the other hand, based on the number of accelerator cards, multiple accelerator card topology architectures are constructed, and based on the model performance indicators of each accelerator card topology architecture when running the target model, the first accelerator card topology architecture whose model performance indicators meet the preset conditions is selected as the architecture reference for deploying the accelerator card. In this way, when the target model is run in the accelerator card, the performance of the target model can be guaranteed to meet the expected performance requirements. Based on the above two aspects, the hardware resource configuration of the target model can be matched with the computing requirements, thereby solving the problem of mismatch between the hardware resource configuration of the model and the computing requirements in the related technology.
[0080] In some embodiments, if the first accelerator card topology does not exist, the method of the present application may further include:
[0081] Increase the number of accelerator cards required to run the target model;
[0082] Rebuild multiple different accelerator card topologies based on the increased number of accelerator cards;
[0083] Determine the model performance indicators when running the target model using each of the reconstructed accelerator card topology architectures.
[0084] If, among the reconstructed multiple accelerator card topology architectures, there is a second accelerator card topology architecture whose model performance indicators meet the preset conditions, then according to the second accelerator card topology architecture, the number of servers used to run the target model is determined, and the accelerator cards are deployed in the servers used to run the target model.
[0085] Specifically, when increasing the number of accelerator cards required to run the target model, the number of accelerator cards determined in step S602 is used as the baseline number, and the number of newly added accelerator cards must be greater than the baseline number. For example, if the number of accelerator cards determined in step S602 is 10, then using 10 as the baseline number, the number of accelerator cards is increased to 12, and multiple different accelerator card topologies are rebuilt according to the specifications of 12 accelerator cards.
[0086] During the operation of the target model, the communication and synchronization time is usually much shorter than the calculation time. Therefore, although the communication and synchronization time between accelerator cards will increase after increasing the number of accelerator cards, the calculation time can be greatly shortened by executing computing tasks in parallel with more accelerator cards, which can reduce the latency of the target model and improve the throughput of the target model.
[0087] In addition, since the number of accelerator cards determined in step S602 can already meet the storage space requirements during the operation of the target model, the accelerator card topology structure is reconstructed based on the increased number of accelerator cards, and the sum of the storage capacity of the accelerator cards can still meet the storage space requirements of the target model.
[0088] In the above embodiment, by searching for the second accelerator card topology architecture by adding the number of accelerator cards, the number of accelerator cards and the accelerator card topology structure can be explored to ensure that the final accelerator card topology structure can meet the performance requirements and storage space requirements of the target model, while reducing the waste of hardware resources.
[0089] In some embodiments, the method of the present application further comprises:
[0090] If there are multiple first accelerator card topology architectures or multiple second accelerator card topology architectures whose model performance indicators meet the preset conditions, then among the accelerator card topology architectures that meet the preset conditions, respectively determine the cost corresponding to each accelerator card topology architecture, where the cost includes at least one of a hardware procurement cost and a hardware operating cost;
[0091] Determine the number of servers used to run the target model and deploy accelerator cards in the servers used to run the target model based on the accelerator card topology with the lowest cost.
[0092] Specifically, hardware procurement costs include, but are not limited to, the purchase costs of servers and accelerator cards. Hardware operating costs refer to the expenses incurred during the operation of the accelerator card or server after deploying the accelerator card in the server according to the specific accelerator card topology, such as accelerator card power consumption, cooling costs, and maintenance costs. Different accelerator card architectures can have different operating costs.
[0093] In the above embodiment, among multiple accelerator card topology architectures whose delay duration and throughput meet preset conditions, the number of servers and the accelerator card deployment method are determined according to the accelerator card topology architecture with the lowest cost, so that the operating cost of the target model can be minimized while the performance requirements of the target model meet actual needs.
[0094] In some embodiments, considering that different types of accelerator cards have their own advantages, when constructing multiple accelerator card topology architectures in step S603, multiple different types of accelerator cards can also be used to construct the accelerator card topology architecture. For example, when a graphics processing unit (GPU) is used as an accelerator card, it can have the characteristics of low latency; when a tensor processing unit (TPU) is used as an accelerator card, it can have the characteristics of high throughput; when a data processing unit (DPU) is used as an accelerator card, it can reduce the waiting time of the graphics processor and the tensor processor by offloading tasks such as communication and storage management. Therefore, in the same accelerator card topology architecture, different types of accelerator cards such as graphics processors, tensor processors, and data processors can be used at the same time. For example, a neural network layer that needs to perform data preprocessing is deployed in the graphics processor, and a neural network layer that needs to update model parameters or perform high-throughput inference is deployed in the tensor processor.
[0095] The following uses a self-attention-based target model as an example to explain how to determine the memory usage, latency, and throughput of the target model during its operation.
[0096] Specifically, in a target model based on the self-attention mechanism, the target model architecture can be categorized into the traditional Transformer architecture and the MoE-Transformer architecture. In the traditional Transformer architecture, the target model may include a self-attention layer (usually multiple), a fully connected feed-forward network (FFN), a normalization layer, and an embedding layer. The parameters of the self-attention layer may include the Q matrix, the K matrix, the V matrix, and the output projection matrix; the parameters of the FFN may include the dimension increase and reduction matrices; the parameters of the normalization layer may include the scaling factor and offset term; and the parameters of the embedding layer may include the position embedding matrix and tokens. The parameters of these layers serve as the model parameters of the target model. In the MoE-Transformer architecture, the FFN is split into multiple independent network layers, also known as expert layers. After model training, different expert layers are assigned to handle tasks in specific domains. For example, expert layer S1 is assigned to handle computational tasks in the code domain, while expert layer S2 is assigned to handle computational tasks in the mathematics domain. Tokens generated during inference are routed to one or more of these expert layers for processing. Typically, the expert layer to which the token is routed is called the activated expert layer, and the expert layer to which the token is not routed is called the inactivated expert layer. In addition, similar to the traditional Transformer architecture, the MoE-Transformer architecture can also include neural network layers such as attention layers, normalization layers, and embedding layers, and these neural network layers can be called non-expert layers.
[0097] Furthermore, for target models based on traditional architectures and MoE-Transformer architectures, both models generate key-value pairs for the text sequence to be inferred based on the Q, K, and V matrices of each self-attention layer during inference. The key-value pairs calculated by the front-end self-attention layer are stored in the accelerator card's storage unit. The back-end self-attention layer reads the key-value pairs calculated by the front-end self-attention layer from the storage unit and performs subsequent calculations based on the read key-value pairs. Therefore, during the operation of the target model, storage space for these key-value pairs must be pre-planned. The storage usage corresponding to these key-value pairs is included in the target model's storage usage. Furthermore, during inference, each layer of both models outputs intermediate results, including activation values. The storage usage corresponding to these activation values is also included in the target model's storage usage.
[0098] Based on the above description, in some embodiments, obtaining the storage usage of the target model during operation in step S601 may include:
[0099] If the target model does not include an expert layer (i.e., a traditional Transformer architecture), obtain the first storage occupancy corresponding to the model parameters of the target model, the second storage occupancy corresponding to the key-value pairs generated during the operation of the target model, and the third storage occupancy corresponding to the activation values generated during the operation of the target model;
[0100] The storage occupancy during the operation of the target model is determined based on the first storage occupancy, the second storage occupancy, the third storage occupancy and the maximum number of concurrency that the target model needs to support.
[0101] Specifically, the storage capacity required by the target model can be determined based on expression (1): .
[0102] (1)
[0103] in, represents the first storage occupancy corresponding to the model parameters of the target model, Indicates the second storage usage corresponding to the key-value pairs generated during the operation of the target model. It represents the third storage usage corresponding to the activation value generated during the operation of the target model, and b represents the maximum number of concurrent connections that the target model needs to support.
[0104] In some embodiments, obtaining the first storage occupancy corresponding to the model parameters of the target model may include:
[0105] Obtain the parameter quantity of the target model and the quantitative accuracy of each model parameter;
[0106] A first storage occupancy is determined based on the parameter quantity and the quantization accuracy of the model parameter.
[0107] Specifically, quantization precision is used to represent the number of bytes required for a single model parameter. Quantization precision can include, but is not limited to, FP32, FP16, and INT8. When the quantization precision is FP32, a single model parameter requires 4 bytes; when the quantization precision is FP16, a single model parameter requires 2 bytes; and when the quantization precision is INT8, a single model parameter requires 1 byte. Typically, the number of parameters and quantization precision of the target model can be provided by the designer of the target model.
[0108] After obtaining the parameter quantity and quantization accuracy of the target model, the first storage occupancy corresponding to the model parameters of the target model can be determined based on expression (2): .
[0109] (2)
[0110] Where N represents the number of parameters of the target model, and q represents the quantization accuracy of the model parameters.
[0111] Furthermore, the obtaining of the second storage occupancy and the third storage occupancy may include:
[0112] Get the number of hidden layers of the target model, the first hidden dimension of a single hidden layer, the maximum sequence length to be input to the target model, and the maximum number of tokens to be output by the target model;
[0113] A second memory footprint and a third memory footprint are determined based on the number of hidden layers, the first hidden dimension, the maximum sequence length, and the maximum number of tokens.
[0114] Specifically, the second storage occupancy corresponding to the key-value pairs generated during the operation of the target model can be determined based on expression (3): .
[0115] (3)
[0116] Among them, b represents the maximum number of concurrent connections that the target model needs to support. Indicates the maximum sequence length that needs to be input into the target model, and n indicates the maximum number of tokens that the target model needs to output. represents the number of hidden layers of the target model, represents the first hidden dimension of a single hidden layer, and q represents the quantization accuracy of the model parameters.
[0117] And, the third storage occupancy corresponding to the activation value generated during the operation of the target model can be determined based on expression (4): .
[0118] (4)
[0119] The description of the relevant parameters in expression (4) can be found in expression (3) and will not be repeated here.
[0120] In the above calculation process, the maximum sequence length, maximum number of sequences, maximum number of tokens, and maximum number of concurrent connections supported by the target model are used to ensure that the final storage usage meets the actual needs of the target model and avoid problems such as insufficient storage space required by the target model.
[0121] In some embodiments, obtaining the storage usage of the target model during operation in step S601 may include:
[0122] If the target model does not include an expert layer (i.e., a traditional Transformer architecture), obtain the first storage occupancy corresponding to the model parameters of the target model, the activation value redundancy coefficient, and the key-value pair redundancy coefficient of the target model;
[0123] A storage occupancy is determined based on the first storage occupancy, the activation value redundancy coefficient, and the key-value pair redundancy coefficient.
[0124] Specifically, the activation value redundancy coefficient and key-value pair redundancy coefficient of the target model can be determined based on experience. After obtaining the first storage occupancy corresponding to the model parameters of the target model, the activation value redundancy coefficient and key-value pair redundancy coefficient of the target model, the storage occupancy during the operation of the target model can be determined based on expression (5). .
[0125] (5)
[0126] in, represents the first storage occupancy corresponding to the model parameters of the target model, K1 represents the activation value redundancy coefficient of the target model, and K2 represents the key-value pair redundancy coefficient of the target model. The calculation formula of the first storage occupancy can be found in the above expression (2) and will not be repeated here. and key-value pair redundancy coefficient It can be between 0.2 and 0.5.
[0127] The above expressions (1) and (5) can be seen as different ways of calculating the target model's storage usage under the traditional Transformer architecture. Compared to expression (1), the calculation method in expression (5) does not need to calculate the second storage usage corresponding to the key-value pairs generated during the operation of the target model, nor does it need to calculate the third storage usage corresponding to the activation values generated during the operation of the target model, thereby greatly simplifying the calculation logic.
[0128] In some embodiments, obtaining the storage usage of the target model during operation in step S601 may include:
[0129] If the target model includes an expert layer and a non-expert layer (i.e., a MoE-Transformer architecture), obtain a first memory occupancy corresponding to the model parameters of the target model, a fourth memory occupancy corresponding to the parameters of the activated expert layer, a fifth memory occupancy corresponding to the key-value pairs and activation values of the expert layer, and a sixth memory occupancy corresponding to the key-value pairs and activation values of the non-expert layer;
[0130] Based on the first storage occupancy, the fourth storage occupancy, the fifth storage occupancy, the sixth storage occupancy and the maximum number of concurrency that the target model needs to support, the storage occupancy during the operation of the target model is determined.
[0131] Specifically, the storage footprint required by the target model of the MoE-Transformer architecture can be calculated based on expression (6): .
[0132]
[0133] in, Indicates the first storage occupancy corresponding to the model parameters of the target model. Indicates the fourth memory occupancy corresponding to the parameters of the activated expert layer. represents the sixth storage occupancy corresponding to the key-value pairs and activation values of the non-expert layer, where is the storage usage corresponding to the key-value pairs of the non-expert layer, is the storage occupancy corresponding to the activation value of the non-expert layer. The fifth storage occupancy corresponding to the key-value pairs and activation values of the expert layer is represented by is the storage usage corresponding to the key-value pairs of the expert layer, is the storage usage corresponding to the activation value of the expert layer. b represents the maximum number of concurrent connections that the target model needs to support.
[0134] In some embodiments, obtaining the fourth storage occupancy includes:
[0135] The maximum number of expert layers supported by the target model to be activated and the maximum number of parameters included in a single expert layer are obtained, and a fourth storage occupancy is determined according to the maximum number of expert layers and the maximum number of parameters.
[0136] Specifically, the maximum number of expert layers supported by the target model and the maximum number of parameters in a single expert layer can be provided by the target model designer. Multiplying the maximum number of expert layers by the maximum number of parameters in a single expert layer yields the number of parameters in the activated expert layer. Furthermore, similar to Expression (2) above, the fourth memory usage can be determined based on the number of parameters in the activated expert layer and the quantization accuracy of the target model.
[0137] In some embodiments, obtaining the fifth storage occupancy may include:
[0138] Obtain the total number of expert layers of the target model, the compression dimension of a single expert layer, the maximum number of expert layers allowed to be activated, the maximum sequence length required to be input to the target model, and the maximum number of tokens required to be output by the target model;
[0139] A fifth storage occupancy is determined based on the total number of expert layers, the compression dimension, the maximum number of expert layers allowed to be activated, the maximum sequence length, and the maximum number of tokens.
[0140] Specifically, the storage usage corresponding to the key-value pairs of the expert layer can be determined based on expression (7): , and the storage occupancy corresponding to the activation value of the expert layer is determined based on expression (8) , and and Add them together and the result is the fifth storage occupancy.
[0141]
[0142]
[0143] In expressions (7) and (8), the parameters b, s, n, and q can be found in the above descriptions and will not be repeated here. represents the total number of expert layers of the target model, Indicates the maximum number of expert layers allowed to be activated, Represents the compressed dimension of a single expert layer.
[0144] In some embodiments, obtaining the sixth storage occupancy may include:
[0145] Get the number of non-expert layers of the target model, the second hidden dimension of a single non-expert layer, the maximum sequence length to be input to the target model, and the maximum number of tokens to be output by the target model;
[0146] A sixth memory footprint is determined based on the number of non-expert layers, the second hidden dimension, the maximum sequence length, and the maximum number of tokens.
[0147] Specifically, the storage usage corresponding to the key-value pairs of the non-expert layer can be determined based on expression (9): , and the storage occupancy corresponding to the activation value of the non-expert layer is determined based on expression (10) , and and Add them together and the result is the sixth storage occupancy.
[0148]
[0149]
[0150] In expressions (9) and (10), the parameters b, s, n, and q can be found in the above descriptions and are not described here in detail. represents the total number of non-expert layers, Represents the second hidden dimension of a single non-expert layer.
[0151] In some embodiments, when the model performance indicator includes the delay duration of the target model, determining the model performance indicator when running the target model using various accelerator card topology architectures in step S604 may include:
[0152] For any target accelerator card topology architecture among multiple accelerator card topology architectures, obtain the first token duration and average token duration when the target accelerator card topology architecture runs the target model, where the first token duration represents the time taken from the target model receiving the sequence to be inferred to the target model outputting the first token, and the average token duration represents the average time taken for the target model to output each other token after the target model outputs the first token;
[0153] Determine the latency of the target accelerator card topology when running the target model based on the first token duration, average token duration, and the maximum number of tokens that the target model needs to output.
[0154] Specifically, the calculation process of the above delay time can be expressed by expression (11).
[0155] (11)
[0156] in, Indicates the delay time of the target model, Indicates the duration of the first token. represents the average token duration, Indicates the number of tokens generated.
[0157] In some embodiments, obtaining the first token duration when the target accelerator card topology architecture runs the target model may include:
[0158] Obtain the first computation time and first communication time of the target model. The first computation time refers to the computation time from when the target model receives the sequence to be inferred to when the target model outputs the first token. The first communication time refers to the communication time from when the target model receives the sequence to be inferred to when the target model outputs the first token.
[0159] An initial token duration is obtained based on the first calculation duration and the first communication duration.
[0160] Specifically, the total data volume of the model parameters and initial activation values of the target model, the communication bandwidth of the target accelerator card topology architecture, and the total computing power of all accelerator cards in the target accelerator card topology architecture can be obtained, and the first calculation duration can be determined based on the total data volume and the total computing power, and the first communication duration can be determined based on the total data volume and the communication bandwidth.
[0161] For example, taking tensor parallelism as an example, the first token duration between the multiple accelerator cards can be calculated based on expression (12): .
[0162]
[0163] in, and The meaning of can be found in the above description and will not be repeated here. It is the computing power of a single accelerator card (that is, the number of multiplication and addition operations that a single accelerator card can complete per second). Indicates the calculated number of accelerator cards. It should be noted that if the number of accelerator cards needs to be increased, is the number of accelerator cards after addition. Indicates the bandwidth between accelerator cards. Indicates the bandwidth of the storage module of the accelerator card. Indicates the number of times each accelerator card needs to communicate with other accelerator cards. Indicates the storage usage corresponding to the activation initial value.
[0164] Furthermore, in the above expression (12), It can represent the first calculation duration, It can indicate the first communication duration.
[0165] In some embodiments, obtaining the average token duration when the target accelerator card topology architecture runs the target model may include:
[0166] Obtain the second calculation duration and the second communication duration of the target model, where the second calculation duration refers to the average calculation duration required for the target model to obtain other tokens after the target model outputs the first token, and the second communication duration refers to the average communication duration required for the target model to obtain other tokens after the target model outputs the first token;
[0167] An average token duration is obtained based on the second calculation duration and the second communication duration.
[0168] For example, taking tensor parallelism as an example, the first token duration between the multiple accelerator cards can be calculated based on expression (13): .
[0169]
[0170] Expression (13) is similar to Expression (12). The main difference is that in Expression (12), since the time consumption of the first token is calculated, there is no transmission time consumption of the key-value pair. That is, in Expression (12), when calculating the first communication duration, only and , does not include the storage usage of key-value pairs. However, in expression (13), since the average time consumption of each token after the second token is calculated, the key-value pairs are already included at this time, so the storage usage of key-value pairs needs to be calculated, that is, = .
[0171] In addition, it should be noted that although both expressions (12) and (13) include ,but The value of can be different. For example, in expression (12) Indicates the storage occupancy corresponding to the activation initial value, the expression (13) Indicates the storage usage corresponding to the activation value that needs to be loaded when calculating each token after the second token.
[0172] This completes the description of this application solution.
[0173] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0174] See also Figure 7 , which is a module diagram of an accelerator card deployment device provided in some embodiments of the present application. Figure 7 In the accelerator card deployment device, the accelerator card deployment device includes:
[0175] The memory usage acquisition module 701 is used to acquire the memory usage during the operation of the target model;
[0176] The accelerator card quantity determination module 702 is used to determine the number of accelerator cards required to run the target model based on the storage usage and the storage capacity of each accelerator card;
[0177] A topology construction module 703 is configured to construct multiple accelerator card topologies based on the number of accelerator cards, wherein different accelerator card topologies have different connection modes between the accelerator cards.
[0178] A calculation module 704 is used to determine the model performance index when running the target model using each accelerator card topology architecture;
[0179] The accelerator card deployment module 705 is used to determine the number of servers used to run the target model and deploy accelerator cards in the servers used to run the target model based on the first accelerator card topology architecture if there is a first accelerator card topology architecture whose model performance indicators meet preset conditions.
[0180] In some embodiments, if the first accelerator card topology architecture does not exist, the topology construction module 703 is further configured to:
[0181] Increase the number of accelerator cards required to run the target model;
[0182] Rebuild multiple different accelerator card topologies based on the increased number of accelerator cards;
[0183] Determine the model performance indicators when running the target model using each of the reconstructed accelerator card topology architectures.
[0184] If, among the reconstructed multiple accelerator card topology architectures, there is a second accelerator card topology architecture whose model performance indicators meet the preset conditions, then according to the second accelerator card topology architecture, the number of servers used to run the target model is determined, and the accelerator cards are deployed in the servers used to run the target model.
[0185] In some embodiments, the calculation module 704 is further configured to:
[0186] If there are multiple first accelerator card topology architectures or multiple second accelerator card topology architectures whose model performance indicators meet the preset conditions, then among the accelerator card topology architectures that meet the preset conditions, respectively determine the cost corresponding to each accelerator card topology architecture, where the cost includes at least one of a hardware procurement cost and a hardware operating cost;
[0187] Determine the number of servers used to run the target model and deploy accelerator cards in the servers used to run the target model based on the accelerator card topology with the lowest cost.
[0188] In some embodiments, the multiple accelerator card topology architectures constructed by the topology construction module 703 include at least two of the following topology architectures:
[0189] All accelerator cards are connected to the same central processor, and the accelerator cards are connected to each other via a first communication bus;
[0190] The accelerator cards are divided into at least two groups, and the accelerator cards in different groups are connected to different central processing units of the same server. The central processing units connected to the accelerator cards in different groups are connected via a second communication bus, the accelerator cards in the same group are connected via a third communication bus, and the accelerator cards in different groups are connected via the second communication bus.
[0191] The accelerator cards are divided into at least two groups, and the accelerator cards in different groups are connected to the same central processing unit of the same server. The accelerator cards in the same group are connected via a fourth communication bus, and the accelerator cards in different groups are connected via a fifth communication bus.
[0192] The accelerator cards are divided into at least two groups. Accelerator cards in different groups are set in different servers. The servers are connected through a sixth communication bus.
[0193] In some embodiments, the target model is a model based on a self-attention mechanism; the storage occupancy acquisition module 701 is specifically used to:
[0194] If the target model does not include an expert layer, obtaining a first storage occupancy corresponding to a model parameter of the target model, a second storage occupancy corresponding to a key-value pair generated during the operation of the target model, and a third storage occupancy corresponding to an activation value generated during the operation of the target model;
[0195] The storage occupancy during the operation of the target model is determined based on the first storage occupancy, the second storage occupancy, the third storage occupancy and the maximum number of concurrency that the target model needs to support.
[0196] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to:
[0197] Get the number of hidden layers of the target model, the first hidden dimension of a single hidden layer, the maximum sequence length to be input to the target model, and the maximum number of tokens to be output by the target model;
[0198] A second memory footprint and a third memory footprint are determined based on the number of hidden layers, the first hidden dimension, the maximum sequence length, and the maximum number of tokens.
[0199] In some embodiments, the target model is a model based on a self-attention mechanism; the storage occupancy acquisition module 701 is specifically used to:
[0200] If the target model does not include an expert layer, obtaining a first storage occupancy corresponding to a model parameter of the target model, an activation value redundancy coefficient, and a key-value pair redundancy coefficient of the target model;
[0201] A storage occupancy is determined based on the first storage occupancy, the activation value redundancy coefficient, and the key-value pair redundancy coefficient.
[0202] In some embodiments, the target model is a model based on a self-attention mechanism; the storage occupancy acquisition module 701 is specifically used to:
[0203] If the target model includes an expert layer and a non-expert layer, obtaining a first storage occupancy corresponding to the model parameters of the target model, a fourth storage occupancy corresponding to the parameters of the activated expert layer, a fifth storage occupancy corresponding to the key-value pairs and activation values of the expert layer, and a sixth storage occupancy corresponding to the key-value pairs and activation values of the non-expert layer;
[0204] Based on the first storage occupancy, the fourth storage occupancy, the fifth storage occupancy, the sixth storage occupancy and the maximum number of concurrency that the target model needs to support, the storage occupancy during the operation of the target model is determined.
[0205] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to:
[0206] The maximum number of expert layers supported by the target model to be activated and the maximum number of parameters included in a single expert layer are obtained, and a fourth storage occupancy is determined according to the maximum number of expert layers and the maximum number of parameters.
[0207] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to:
[0208] Obtain the total number of expert layers of the target model, the compression dimension of a single expert layer, the maximum number of expert layers allowed to be activated, the maximum sequence length required to be input to the target model, and the maximum number of tokens required to be output by the target model;
[0209] A fifth storage occupancy is determined based on the total number of expert layers, the compression dimension, the maximum number of expert layers allowed to be activated, the maximum sequence length, and the maximum number of tokens.
[0210] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to:
[0211] Get the number of non-expert layers of the target model, the second hidden dimension of a single non-expert layer, the maximum sequence length to be input to the target model, and the maximum number of tokens to be output by the target model;
[0212] A sixth memory footprint is determined based on the number of non-expert layers, the second hidden dimension, the maximum sequence length, and the maximum number of tokens.
[0213] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to:
[0214] Obtain the parameter quantity of the target model and the quantitative accuracy of each model parameter;
[0215] A first storage occupancy is determined based on the parameter quantity and the quantization accuracy of the model parameter.
[0216] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to:
[0217] The first token time and average token time when the accelerator card topology runs the target model. The first token time represents the time from when the target model receives the sequence to be inferred to when the target model outputs the first token. The average token time represents the average time it takes for the target model to output each token after the target model outputs the first token.
[0218] Determine the latency of the target accelerator card topology when running the target model based on the first token duration, average token duration, and the maximum number of tokens that the target model needs to output.
[0219] In some embodiments, the calculation module 704 is specifically configured to:
[0220] Obtain the first computation time and first communication time of the target model. The first computation time refers to the computation time from when the target model receives the sequence to be inferred to when the target model outputs the first token. The first communication time refers to the communication time from when the target model receives the sequence to be inferred to when the target model outputs the first token.
[0221] An initial token duration is obtained based on the first calculation duration and the first communication duration.
[0222] In some embodiments, the calculation module 704 is specifically configured to:
[0223] Obtain the total data volume of the model parameters and initial activation values of the target model, the communication bandwidth of the target accelerator card topology, and the total computing power of all accelerator cards in the target accelerator card topology;
[0224] Determine the first computing duration based on the total data volume and total computing power;
[0225] A first communication duration is determined based on the total data volume and the communication bandwidth.
[0226] In some embodiments, the calculation module 704 is specifically configured to:
[0227] Obtain the second calculation duration and the second communication duration of the target model, where the second calculation duration refers to the average calculation duration required for the target model to obtain other tokens after the target model outputs the first token, and the second communication duration refers to the average communication duration required for the target model to obtain other tokens after the target model outputs the first token;
[0228] An average token duration is obtained based on the second calculation duration and the second communication duration.
[0229] See also Figure 8 An embodiment of the present application further provides an electronic device, including a memory 10 and a processor 20, wherein the memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any of the above-mentioned accelerator card deployment method embodiments.
[0230] An embodiment of the present application further provides a computer non-volatile readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned accelerator card deployment method embodiments when running.
[0231] In an exemplary embodiment, the above-mentioned computer non-volatile readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.
[0232] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned accelerator card deployment method embodiments are implemented.
[0233] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned accelerator card deployment method embodiments are implemented.
[0234] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0235] The above is a detailed introduction to the acceleration card deployment method, device, equipment, storage medium and program product provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for deploying an accelerator card, characterized in that: The method comprises: Get the storage usage of the target model during operation; Determining the number of accelerator cards required to run the target model based on the storage occupancy and the storage capacity of each accelerator card; Constructing multiple accelerator card topology architectures based on the number of accelerator cards, wherein different accelerator card topology architectures have different connection modes between the accelerator cards; Determine model performance metrics when running the target model using various accelerator card topology architectures; If there is a first accelerator card topology architecture whose model performance indicators meet the preset conditions, determining the number of servers used to run the target model based on the first accelerator card topology architecture, and deploying accelerator cards in the servers used to run the target model; Wherein, the model performance index includes the delay duration of the target model; Determining the model performance indicators when running the target model using each accelerator card topology architecture includes: For any target accelerator card topology architecture among the multiple accelerator card topology architectures, obtaining a first token duration and an average token duration when the target accelerator card topology architecture runs the target model, wherein the first token duration represents the time taken from the target model receiving the sequence to be inferred to the target model outputting the first token, and the average token duration represents the average time taken for the target model to output other tokens after the target model outputs the first token; Based on the first token duration, the average token duration, and the maximum number of tokens required to be output by the target model, a delay duration when the target accelerator card topology architecture runs the target model is determined.
2. The method according to claim 1, characterized in that If the first accelerator card topology does not exist, the method further includes: Increase the number of accelerator cards required to run the target model; Rebuild multiple different accelerator card topologies based on the increased number of accelerator cards; Determining, among the reconstructed multiple accelerator card topology architectures, a model performance indicator when the target model is run using each accelerator card topology architecture; If, among the reconstructed multiple accelerator card topology architectures, there is a second accelerator card topology architecture whose model performance indicators meet the preset conditions, then according to the second accelerator card topology architecture, the number of servers used to run the target model is determined, and the accelerator cards are deployed in the servers used to run the target model.
3. The method according to claim 2, characterized in that The method further comprises: If there are multiple first accelerator card topology architectures or multiple second accelerator card topology architectures whose model performance indicators meet the preset conditions, determining a cost corresponding to each accelerator card topology architecture among the accelerator card topology architectures that meet the preset conditions, where the cost includes at least one of a hardware procurement cost and a hardware operating cost; According to the accelerator card topology architecture with the lowest cost, the number of servers used to run the target model is determined, and accelerator cards are deployed in the servers used to run the target model.
4. The method according to claim 1 or 2, characterized in that Among the multiple accelerator card topologies constructed, at least two of the following topologies are included: All accelerator cards are connected to the same central processor, and the accelerator cards are connected to each other via a first communication bus; The accelerator cards are divided into at least two groups, and the accelerator cards in different groups are connected to different central processing units of the same server. The central processing units connected to the accelerator cards in different groups are connected via a second communication bus, the accelerator cards in the same group are connected via a third communication bus, and the accelerator cards in different groups are connected via the second communication bus. The accelerator cards are divided into at least two groups, and the accelerator cards in different groups are connected to the same central processing unit of the same server. The accelerator cards in the same group are connected via a fourth communication bus, and the accelerator cards in different groups are connected via a fifth communication bus. The accelerator cards are divided into at least two groups. Accelerator cards in different groups are set in different servers. The servers are connected through a sixth communication bus.
5. The method according to claim 1, wherein The target model is a model based on the self-attention mechanism; The obtaining of the storage usage during the operation of the target model includes: If the target model does not include an expert layer, obtaining a first storage occupancy corresponding to a model parameter of the target model, a second storage occupancy corresponding to a key-value pair generated during the operation of the target model, and a third storage occupancy corresponding to an activation value generated during the operation of the target model; The storage occupancy during the operation of the target model is determined based on the first storage occupancy, the second storage occupancy, the third storage occupancy and the maximum number of concurrency that the target model needs to support.
6. The method according to claim 5, characterized in that Obtaining the second storage occupancy and the third storage occupancy includes: Obtain the number of hidden layers of the target model, the first hidden dimension of a single hidden layer, the maximum sequence length to be input to the target model, and the maximum number of tokens to be output by the target model; The second memory footprint and the third memory footprint are determined based on the number of hidden layers, the first hidden dimension, the maximum sequence length, and the maximum number of tokens.
7. The method according to claim 1, characterized in that The target model is a model based on the self-attention mechanism; The obtaining of the storage usage during the operation of the target model includes: If the target model does not include an expert layer, obtaining a first storage occupancy corresponding to a model parameter of the target model, an activation value redundancy coefficient, and a key-value pair redundancy coefficient of the target model; The storage occupancy is determined based on the first storage occupancy, the activation value redundancy coefficient, and the key-value pair redundancy coefficient.
8. The method according to claim 1, characterized in that The target model is a model based on the self-attention mechanism; The obtaining of the storage usage during the operation of the target model includes: If the target model includes an expert layer and a non-expert layer, obtaining a first storage occupancy corresponding to a model parameter of the target model, a fourth storage occupancy corresponding to a parameter of the activated expert layer, a fifth storage occupancy corresponding to a key-value pair and an activation value of the expert layer, and a sixth storage occupancy corresponding to the key-value pair and the activation value of the non-expert layer; The storage occupancy during the operation of the target model is determined based on the first storage occupancy, the fourth storage occupancy, the fifth storage occupancy, the sixth storage occupancy and the maximum number of concurrency that the target model needs to support.
9. The method according to claim 8, characterized in that Obtaining the fourth storage occupancy includes: The maximum number of expert layers supported by the target model to be activated and the maximum number of parameters included in a single expert layer are obtained, and the fourth storage occupancy is determined according to the maximum number of expert layers and the maximum number of parameters.
10. The method according to claim 8, characterized in that Obtaining the fifth storage occupancy includes: Obtain the total number of expert layers of the target model, the compression dimension of a single expert layer, the maximum number of expert layers allowed to be activated, the maximum sequence length required to be input into the target model, and the maximum number of tokens required to be output by the target model; The fifth storage occupancy is determined based on the total number of expert layers, the compression dimension, the maximum number of expert layers allowed to be activated, the maximum sequence length, and the maximum number of tokens.
11. The method according to claim 8, characterized in that Obtaining the sixth storage occupancy includes: Obtaining the number of non-expert layers of the target model, the second hidden dimension of a single non-expert layer, the maximum sequence length to be input into the target model, and the maximum number of tokens to be output by the target model; The sixth memory usage is determined based on the number of non-expert layers, the second hidden dimension, the maximum sequence length, and the maximum number of tokens.
12. The method according to any one of claims 5, 7 and 8, characterized in that: Obtaining a first storage occupancy corresponding to the model parameter of the target model includes: Obtaining the parameter quantity of the target model and the quantization accuracy of each model parameter; The first storage occupancy is determined based on the parameter quantity and the quantization accuracy of the model parameter.
13. The method according to claim 1, wherein Obtaining the first token duration when the target accelerator card topology architecture runs the target model includes: Obtain a first computation time and a first communication time of the target model, where the first computation time refers to the computation time from when the target model receives the sequence to be inferred to when the target model outputs the first token, and the first communication time refers to the communication time from when the target model receives the sequence to be inferred to when the target model outputs the first token; The first token duration is obtained based on the first calculation duration and the first communication duration.
14. The method according to claim 13, characterized in that The obtaining of the first calculation duration and the first communication duration of the target model includes: Obtaining the total data volume of the model parameters and initial activation values of the target model, the communication bandwidth of the target accelerator card topology, and the total computing power of all accelerator cards in the target accelerator card topology; Determining the first computing duration based on the total data volume and the total computing power; The first communication duration is determined based on the total data volume and the communication bandwidth.
15. The method according to claim 1, wherein Obtaining an average token duration when the target accelerator card topology architecture runs the target model, including: Obtain a second calculation duration and a second communication duration of the target model, where the second calculation duration refers to the average calculation duration consumed by the target model to obtain other tokens after the target model outputs the first token, and the second communication duration refers to the average communication duration consumed by the target model to obtain other tokens after the target model outputs the first token; The average token duration is obtained based on the second calculation duration and the second communication duration.
16. An accelerator card deployment device, characterized in that: The device comprises: A storage occupancy acquisition module is used to obtain the storage occupancy during the operation of the target model; an accelerator card quantity determination module, configured to determine the number of accelerator cards required to run the target model based on the storage occupancy and the storage capacity of each accelerator card; A topology construction module is used to construct multiple accelerator card topology architectures according to the number of accelerator cards, wherein the connection methods between the accelerator cards are different in different accelerator card topology architectures; A calculation module, configured to determine a model performance indicator when each accelerator card topology architecture is used to run the target model, wherein the model performance indicator includes a delay duration of the target model; when determining the model performance indicator when each accelerator card topology architecture is used to run the target model, for any target accelerator card topology architecture among the multiple accelerator card topology architectures, obtain a first token duration and an average token duration when the target accelerator card topology architecture runs the target model, and determine the delay duration when the target accelerator card topology architecture runs the target model based on the first token duration, the average token duration, and the maximum number of tokens required to be output by the target model, wherein the first token duration represents the time taken from the target model receiving the sequence to be inferred to the target model outputting the first token, and the average token duration represents the average time taken for the target model to output other tokens after the target model outputs the first token; An accelerator card deployment module is used to determine the number of servers used to run the target model and deploy accelerator cards in the servers used to run the target model based on a first accelerator card topology architecture if there is a first accelerator card topology architecture whose model performance indicators meet preset conditions.
17. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the accelerator card deployment method according to any one of claims 1 to 15 when executing the computer program.
18. A computer-readable non-volatile storage medium, characterized in that: The computer non-volatile readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the accelerator card deployment method according to any one of claims 1 to 15 is implemented.
19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the accelerator card deployment method according to any one of claims 1 to 15 is implemented.
Citation Information
Patent Citations
Array server configuration method suitable for AIPC and server
CN119473614A
GPU topology awareness scheduling method, electronic equipment and medium
CN119938289A