Accelerator card deployment method and device, equipment, storage medium and program product

By obtaining the storage occupancy of the target model and building a variety of acceleration card topology architectures, the problem of the mismatch between the acceleration card hardware resource configuration and computing requirements is solved, the model performance is optimized, resource waste is avoided and throughput is improved.

CN120407200AActive Publication Date: 2025-08-01INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510897024.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-01
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

In the prior art, the topological structure solidified configuration of the accelerator card leads to mismatch of hardware resource configuration and computing requirements, resulting in problems such as degradation of model performance, increased latency and decreased throughput.

Method used

By obtaining the storage occupancy of the target model, determining the number of acceleration cards required, and building a variety of acceleration card topology architectures, selecting topology architectures that meet preset conditions to match hardware resources and computing requirements, and optimizing model performance.

Benefits of technology

The rational allocation of accelerator card resources is realized, and resource waste or insufficient resources are avoided, ensuring that the model performance meets the expected requirements, reducing latency and improving throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407200A_ABST
    Figure CN120407200A_ABST
Patent Text Reader

Abstract

The invention discloses an accelerator card deployment method and device, equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence. In the method, on one hand, the number of accelerator cards needed for running a target model is determined based on the storage occupation amount in the running process of the target model, and therefore the number of the accelerator cards needed for running the target model is increased; it can be ensured that the sum of the storage capacities of the acceleration cards is matched with the storage occupation amount of the target model. On the other hand, according to the number of the acceleration cards, a plurality of acceleration card topology architectures are constructed, and the first acceleration card topology architecture of which the model performance index meets the preset condition is selected as the architecture reference for deploying the acceleration cards based on the model performance index when the target model is operated by each acceleration card topology architecture, so that when the target model is operated in the acceleration cards, the target model can be deployed in the acceleration cards; and it can be ensured that the performance of the target model meets expected performance requirements. Based on the two aspects, the problem that the hardware resource configuration of the model is not matched with the calculation requirement in the related technology can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to an acceleration card deployment method, device, equipment, storage medium, and program product. Background Art

[0002] Acceleration cards (such as graphics cards, intelligent network cards, etc.) are the core hardware that supports the operation of models. With the development of technologies such as the Large Language Model (LLM), the number of model parameters has increased at an extremely fast rate, and it often occurs that a single acceleration card cannot load the entire model parameters. Therefore, it is necessary to split a model and run it on multiple acceleration cards.

[0003] Currently, in some technologies, in the servers used for model operation, the number of acceleration cards and the topological structure between acceleration cards are fixedly configured. This method of fixed configuration is not reasonable enough, and problems such as the mismatch between the hardware resource configuration of the model and the computing requirements may occur, resulting in a decline in the performance of the model, such as an increase in latency, a decrease in throughput, and the model being unable to run normally. Summary of the Invention

[0004] This application provides an acceleration card deployment method, an acceleration card deployment device, an electronic device, a computer non-volatile readable storage medium, and a computer program product to at least solve the problem of the mismatch between the hardware resource configuration of the model and the computing requirements in the related technologies.

[0005] This application provides an acceleration card deployment method, including: Obtaining the storage occupancy during the operation of the target model; Determining the number of acceleration cards required to run the target model based on the storage occupancy and the storage capacity of each acceleration card; Constructing multiple acceleration card topological architectures according to the number of acceleration cards, where the connection methods between acceleration cards are different in different acceleration card topological architectures; Determining the model performance metrics when running the target model using each acceleration card topological architecture; If there is a first acceleration card topological architecture whose model performance metrics meet the preset conditions, then determining the number of servers used to run the target model based on the first acceleration card topological architecture, and deploying acceleration cards in the servers used to run the target model.

[0006] This application also provides an acceleration card deployment device, including: A storage occupancy acquisition module, configured to acquire the storage occupancy during the operation of the target model; An acceleration card quantity determination module, configured to determine the quantity of acceleration cards required for running the target model according to the storage occupancy and the storage capacity of each acceleration card; A topology construction module, configured to construct multiple acceleration card topologies according to the quantity of acceleration cards, wherein, in different acceleration card topologies, the connection manners between acceleration cards are different; A calculation module, configured to determine the model performance metrics when running the target model using each acceleration card topology; An acceleration card deployment module, configured to, if there is a first acceleration card topology whose model performance metrics meet the preset conditions, determine the quantity of servers for running the target model according to the first acceleration card topology, and deploy acceleration cards in the servers for running the target model.

[0007] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of the above acceleration card deployment method when executing the computer program.

[0008] This application also provides a computer non-volatile readable storage medium, in which a computer program is stored, and wherein the computer program implements the steps of the above acceleration card deployment method when executed by a processor.

[0009] This application also provides a computer program product, including a computer program, and the computer program implements the steps of the above acceleration card deployment method when executed by a processor.

[0010] In the technical solutions of some embodiments of this application, on the one hand, based on the storage occupancy during the running process of the target model, the quantity of acceleration cards required for running the target model is determined, so that the sum of the storage capacities of the acceleration cards can be matched with the storage occupancy of the target model, avoiding the problem of waste or insufficiency of acceleration card resources. On the other hand, according to the quantity of acceleration cards, multiple acceleration card topologies are constructed, and based on the model performance metrics when running the target model using each acceleration card topology, the first acceleration card topology whose model performance metrics meet the preset conditions is selected as the architecture reference for deploying acceleration cards. In this way, when running the target model on the acceleration cards, the performance of the target model can be ensured to meet the expected performance requirements. Based on the above two aspects, the hardware resource configuration of the target model can be matched with the computing requirements, thereby solving the problem of mismatch between the hardware resource configuration and the computing requirements of the model in the related art. Description of the Drawings

[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0012] Figure 1 It is a schematic diagram of the architecture of some servers for running large language models; Figure 2 It is a schematic diagram of one of the accelerator card topologies; Figure 3 It is a schematic diagram of another accelerator card topology; Figure 4 It is a schematic diagram of another accelerator card topology; Figure 5 It is a schematic diagram of another accelerator card topology; Figure 6 It is a schematic flowchart of the accelerator card deployment method provided by some embodiments of the present application; Figure 7 It is a schematic diagram of the modules of the accelerator card deployment device provided by some embodiments of the present application; Figure 8 It is a schematic diagram of the modules of the electronic device provided by some embodiments of the present application. Detailed implementation manners

[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0014] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0015] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following will further describe the present application in detail with reference to the drawings and specific implementation manners.

[0016] Refer to in combination Figure 1, which is a schematic diagram of the architecture of some servers for running large language models. Figure 1 In Figure 1 , the server includes a central processing unit and multiple acceleration cards. The multiple acceleration cards are connected to the central processing unit, and each acceleration card includes a storage unit and a core computing unit. The acceleration cards are connected through a communication bus. The communication bus can be a PCIe (Peripheral Component Interconnect Express) bus, an NVLink bus, etc.

[0017] When the server runs a large language model, if the number of parameters of the large language model exceeds the storage capacity limit of a single acceleration card, the central processing unit can split the large language model and run it on multiple acceleration cards. For ease of understanding, the following two examples 1) and 2) of model splitting are given.

[0018] 1) Each acceleration card runs one or more neural network layers of the large language model, that is, pipeline parallelism.

[0019] For example, assume that the large language model includes 10 neural network layers, and the model can be split as follows: Save the parameters of the 1st to 3rd neural network layers of the large language model to the storage unit S1 of acceleration card G1, save the parameters of the 4th to 8th neural network layers of the large language model to the storage unit S2 of acceleration card G2, and save the parameters of the 9th to 10th neural network layers of the large language model to the storage unit S3 of acceleration card G3. In this way, the core computing unit C1 in acceleration card 1 can load the parameters of the 1st to 3rd neural network layers of the large language model from the storage unit S1 and run the 1st to 3rd neural network layers of the large language model; the core computing unit C2 in acceleration card 2 can load the parameters of the 4th to 8th neural network layers of the large language model from the storage unit S2 and run the 4th to 8th neural network layers of the large language model; the core computing unit C3 in acceleration card 3 can load the parameters of the 9th to 10th neural network layers of the large language model from the storage unit S3 and run the 9th to 10th neural network layers of the large language model.

[0020] When using a large language model for inference, the central processing unit can preprocess the text sequence to be inferred (such as the question that needs to be answered by the large language model), so as to map the text sequence into the feature vectors required by the large language model (that is, perform Embedding processing on the text sequence). After completing the preprocessing, the central processing unit can save the feature vectors of the text sequence to the storage unit S1 of the acceleration card G1. After the core computing unit C1 of the acceleration card G1 reads the feature vectors and model parameters from the storage unit S1, it calculates the outputs of the 1st to 3rd neural network layers, and based on the communication bus between the acceleration cards, saves the output of the 3rd neural network layer (that is, the intermediate processing result) to the storage unit S2 of the acceleration card G2. The core computing unit C2 of the acceleration card G2 reads the model parameters and the output of the 3rd neural network layer from the storage unit S2, continues to calculate the outputs of the 4th to 8th neural network layers, and based on the communication bus between the acceleration cards, saves the output of the 8th neural network layer to the storage unit S3 of the acceleration card G3. According to a principle similar to that of the acceleration card G2, the acceleration card G3 can calculate the outputs of the 9th to 10th neural network layers. Among them, the output of the 10th neural network layer can be regarded as the inference result (that is, the answer to the question). The core computing unit C3 in the acceleration card G3 can return the inference result to the central processing unit. In this way, one inference of the large language model is completed.

[0021] 2) Divide the parameters of the same neural network layer into different acceleration cards for operation, that is, tensor parallelism.

[0022] For example, assume that the large language model includes 10 neural network layers, then the model can be split in the following way: divide the parameters of the 1st neural network layer of the large language model into subset 11 and subset 12, and save the parameters in subset 11 to the storage unit S11 of the acceleration card G11, and save the parameters in subset 12 to the storage unit S12 of the acceleration card G12. Divide the parameters of the 2nd neural network layer of the large language model into subset 21, subset 22 and subset 23, and save the parameters in subset 21 to the storage unit S21 of the acceleration card G21, save the parameters in subset 22 to the storage unit S22 of the acceleration card G22, and save the parameters in subset 23 to the storage unit S23 of the acceleration card G23. And so on.

[0023] When using a large language model for inference, the central processing unit can save the feature vectors of the text sequence to be inferred into the acceleration cards running the first neural network layer, that is, save the feature vectors into the storage units S11 of acceleration card G11 and the storage unit S12 of acceleration card G12. The core computing unit C11 of acceleration card G11 reads the subset 11 of model parameters and the feature vectors from the storage unit S11, and after calculation, obtains the calculation result Y11. And the core computing unit C12 of acceleration card G12 reads the subset 12 of model parameters and the feature vectors from the storage unit S12, and after calculation, obtains the calculation result Y12. Based on the communication bus between the acceleration cards, acceleration card G11 can synchronize the calculation result Y11 to acceleration card G12, and acceleration card G12 can synchronize the calculation result Y12 to acceleration card G11. Acceleration card G11 combines the calculation results Y11 and Y12 to obtain the complete calculation result of the first neural network layer. Similarly, acceleration card G12 combines the calculation results Y11 and Y12 to also obtain the complete calculation result of the first neural network layer.

[0024] For ease of understanding, the following is illustrated by an example. For instance, assume that the parameters of the first neural network layer include the weight matrix M. Then the weight matrix M can be split into sub-matrices M1 and M2. Among them, sub-matrix M1 can be regarded as the subset 11 of the above-mentioned model parameters, and sub-matrix M2 can be regarded as the subset 12 of the above-mentioned model parameters. Sub-matrix M1 can be saved into the storage unit S11 of acceleration card G11, and sub-matrix M2 can be saved into the storage unit S12 of acceleration card G12. The core computing unit C11 of acceleration card G11 can read sub-matrix M1 and the feature vectors from the storage unit S11, and perform calculations based on sub-matrix M1 and the feature vectors to obtain the calculation result Y11. And the core computing unit C12 of acceleration card G12 can read sub-matrix M2 and the feature vectors from the storage unit S12, and perform calculations based on sub-matrix M2 and the feature vectors to obtain the calculation result Y12. After acceleration cards G11 and G12 synchronize and combine the calculation results Y11 and Y12 based on the communication bus, the complete calculation result of the weight matrix M and the feature vectors can be obtained, and this complete calculation result can be used as the output of the first neural network layer.

[0025] Furthermore, the core computing units C11 and C12 can save the output of the first neural network layer to the acceleration cards for running the second neural network layer, that is, to the storage units S21 of acceleration card G21, the storage unit S22 of acceleration card G22, and the storage unit S23 of acceleration card G23. Similar to the principle of the first neural network layer, the core computing unit C21 of acceleration card G21 can read the subset 21 of model parameters and the output of the first neural network layer from the storage unit S21 to obtain the calculation result Y21. The core computing unit C22 of acceleration card G22 can read the subset 22 of model parameters and the output of the first neural network layer from the storage unit S22 to obtain the calculation result Y22. The core computing unit C23 of acceleration card G23 can read the subset 23 of model parameters and the output of the first neural network layer from the storage unit S23 to obtain the calculation result Y23. Based on the communication bus, the acceleration cards G21, G22, and G23 can synchronize the calculation results with each other to obtain the complete output of the second neural network layer. Among them, when the parameters of the same neural network layer are divided into three or more acceleration cards, the calculation results can be merged in segments to obtain the complete output of the corresponding neural network layer. For example, when synchronizing the calculation results between acceleration card G21 and acceleration card G22, the intermediate calculation result Y1 can be merged. When synchronizing the calculation results between acceleration card G21 and acceleration card G23, the intermediate calculation result Y2 can be merged. When synchronizing the intermediate calculation results Y1 and Y2 between acceleration card G21 and acceleration card G23, the complete output Y of the second neural network layer can be obtained. Compared with merging all the calculation results in the same acceleration card, this way of merging the calculation results in segments can reduce the communication pressure between the acceleration cards. For example, if both acceleration cards G21 and G22 send the calculation results to acceleration card G23 for merging, there may be a communication bottleneck in acceleration card G23, thus reducing the merging efficiency of the calculation results.

[0026] Based on the above description, it can be understood that although splitting the large language model to run on multiple acceleration cards can solve the problem that a single acceleration card cannot load all model parameters, it also introduces problems such as communication overhead and synchronization. This makes the model performance metrics of the large language model affected by both the computing speed of the acceleration cards and the communication and synchronization durations between the acceleration cards. Among them, the model performance metrics can include but are not limited to the latency duration and throughput of the large language model. The latency duration refers to the time required for the large language model to output a complete inference result after inputting the text sequence to be inferred. The throughput refers to the number of tokens output by the large language model per second. Generally, the more the number of acceleration cards, the fewer the computing tasks each acceleration card needs to execute. In this way, the computing speed can be increased and the computing duration can be shortened. However, with the increase in the number of acceleration cards, the communication and synchronization durations will also increase accordingly. Therefore, a reasonable number of acceleration cards is a key factor in reducing model latency and increasing model throughput. In addition, with the same number of acceleration cards, the communication and synchronization durations corresponding to different acceleration card topologies will also be different. Therefore, a reasonable acceleration card topology is also a key factor in reducing model latency and increasing model throughput. For ease of understanding, the following is illustrated by examples. Suppose there are 8 acceleration cards. Referring jointly to Figures 2 to 4 are schematic diagrams of different topologies composed of 8 acceleration cards.

[0027] Figure 2 In, 8 acceleration cards are deployed in the same server A. Server A includes a central processing unit A1 and a switching device A2. The 8 acceleration cards are connected to the central processing unit A1 through the switching device A2, and the acceleration cards are connected to each other through a communication bus. Based on Figure 2 the acceleration card topology shown, the central processing unit A1 can split the large language model and store different model parameters in different acceleration cards. When the large language model executes an inference task, the central processing unit A1 can preprocess the text sequence to be inferred, obtain the feature vector of the text sequence, and store the feature vector in the acceleration card running the first neural network layer. The acceleration cards can execute inference tasks in a pipelined parallel manner and communicate and synchronize data through the communication bus.

[0028] Figure 3Among them, 8 acceleration cards are deployed in the same server B, and the 8 acceleration cards are divided into two groups, with each group including 4 acceleration cards. Server B includes central processors B11, B21 and switching devices B12, B22. One group of acceleration cards is connected to central processor B11 through switching device B12, and the other group of acceleration cards is connected to central processor B21 through switching device B22. The two groups of acceleration cards are not directly connected through a communication bus. The acceleration cards within the same group are connected through a communication bus. Central processors B11 and B21 are connected through a communication bus. Based on Figure 3 the acceleration card topology shown, central processors B11 and B21 can split the large language model and save different model parameters in different acceleration cards. For example, central processor B11 can split the parameters of the 1st to 6th neural network layers of the large language model into 4 subsets 11, 12, 13, 14, and save subset 11 in the 1st acceleration card it is connected to, save subset 12 in the 2nd acceleration card it is connected to, and so on. Similarly, central processor B21 can split the parameters of the 7th to 14th neural network layers of the large language model into 4 subsets and save the parameters in these 4 subsets in the 4 acceleration cards it is connected to. When the large language model performs an inference task, central processor B11 or central processor B12 can preprocess the text sequence to be inferred to obtain the feature vector of the text sequence and save the feature vector in the acceleration card running the 1st neural network layer. The acceleration cards can perform the inference task in a pipelined parallel or tensor parallel manner. During the execution of the inference task, the acceleration cards within the same group can communicate and synchronize data through the communication bus, and the acceleration cards in different groups communicate and synchronize data through the communication bus between central processors B11 and B21.

[0029] Figure 4 Similar to Figure 3 it basically, the main difference is that the two groups of acceleration cards are both connected to the same central processor C11 through switching devices C12 and C22. Switching devices C12 and C22 are connected through a communication bus, and central processors C11 and C21 are connected through a communication bus. Central processor C11 is used to split the large language model and save the split model parameters in the two groups of acceleration cards. When the large language model performs an inference task, central processor C11 can preprocess the text sequence to be inferred to obtain the feature vector of the text sequence and save the feature vector in the acceleration card running the 1st neural network layer. The acceleration cards can perform the inference task in a pipelined parallel or tensor parallel manner. During the execution of the inference task, the acceleration cards within the same group can communicate and synchronize data through the communication bus, and the acceleration cards in different groups communicate and synchronize data through the communication bus between switching devices C12 and C22.

[0030] Figure 5 and Figure 3 Basically similar, the main difference is that, of the two groups of accelerator cards obtained, one group of accelerator cards is deployed in server E, and the other group of accelerator cards is deployed in server F. Server E includes a central processing unit E11 and a switching device E12, and the accelerator card in server E is connected to the central processing unit E11 through the switching device E12. Server F includes a central processing unit F11 and a switching device F12, and the accelerator card in server F is connected to the central processing unit F11 through the switching device F12. The central processing units E11 and F11 are connected via an inter-server communication bus. When executing inference tasks in a large language, accelerator cards in the same group communicate and synchronize data via the intra-group communication bus, and accelerator cards in different groups communicate and synchronize data via the inter-server communication bus.

[0031] It is understandable that due to the differences between the topologies, the above Figures 2 to 5 When running the same large language model using the accelerator card topology shown, the latency and throughput can vary. Therefore, a reasonable accelerator card topology is also a key factor in reducing model latency and improving model throughput.

[0032] Currently, when some server vendors sell servers for model operation, the number of accelerator cards and the topology of the accelerator cards in the server are fixed. For example, the number of accelerator cards is 10, and they are all based on Figure 2 The topology shown in the figure is used for deployment. Since the parameter quantity and structure of different large language models may be different, if the number of accelerator cards and the accelerator card topology are fixed, the hardware resource configuration of the large language model may not match the computing requirements, which may lead to performance degradation of the large language model, increased latency, decreased throughput, and model failure. Figure 2 The topology shown is illustrated by taking the deployment of 10 accelerator cards in the server as an example.

[0033] For example, suppose customer T1 needs to purchase a server for running the large language model D1. The large language model D1 has a large number of parameters and actually requires 12 accelerator cards. However, since the server is only configured with 10 accelerator cards or has only 10 accelerator card slots, customer T1 may not be able to run the large language model D1 normally based on the purchased server.

[0034] For another example, assume that customer T2 needs to purchase a server for running the large language model D2, and the number of parameters of the large language model D2 is relatively small. The actual number of accelerator cards required is 5. However, since 10 accelerator cards are actually deployed in the server, there may be a problem of waste of accelerator card resources.

[0035] For yet another example, assume that customer T3 needs to purchase a server for running the large language model D3, and the number of parameters of the large language model D3 matches 10 accelerator cards. However, the 10 accelerator cards can enable the large language model D3 to achieve better performance only when deployed according to the topology structure shown above. Figure 3 In this case, based on the purchased server, customer T3 cannot optimize the performance of the large language model D3.

[0036] In view of this, the present application provides an accelerator card deployment method. Before the server leaves the factory, based on the relevant information of the large language model to be run on the server, the hardware resource configuration (i.e., the number of accelerator cards in the server and the topology structure between the accelerator cards) adapted to the computing requirements of the large language model can be obtained, and the server can be configured according to the hardware resource configuration adapted to the computing requirements of the large language model. In this way, after selling the server to the customer, the customer can enable the large language model to achieve better performance based on the purchased server. The accelerator card deployment method can be applied to electronic devices. The electronic devices can include, but are not limited to, tablet computers, laptop computers, desktop computers, etc. Referring to Figure 6 is a schematic flowchart of the accelerator card deployment method provided in some embodiments of the present application. Figure 6 In, the accelerator card deployment method includes the following steps: Step S601, obtain the storage occupancy during the operation of the target model.

[0037] In this embodiment, the target model is the large language model to be run. Based on the above Figure 1 relevant description, during the operation of the target model, the model parameters of the target model and the intermediate operation results (such as activation values, gradients, etc.) generated during the operation of the target model need to occupy the storage space of the accelerator card. Therefore, the number of parameters, model structure, inference logic, etc. of the target model can be analyzed to obtain the storage occupancy during the operation of the target model.

[0038] In the subsequent embodiments of the present application, taking the model based on the self-attention mechanism as an example, how to obtain the storage occupancy during the operation of the target model is elaborated, which will not be repeated here.

[0039] Step S602, determine the number of accelerator cards required for running the target model according to the storage occupancy and the storage capacity of each accelerator card.

[0040] Specifically, according to the principle that the sum of the storage capacities of multiple acceleration cards is greater than or equal to the storage occupancy, the number of acceleration cards required to run the target model can be determined. Among the multiple acceleration cards, the storage capacities of different acceleration cards can be the same or different. For example, when the storage occupancy of the target model is 50GB, if the storage capacity of each acceleration card is 10GB, the number of acceleration cards required to run the target model is at least 5.

[0041] Step S603: Construct multiple acceleration card topologies according to the number of acceleration cards. Among them, in different acceleration card topologies, the connection methods between acceleration cards are different.

[0042] Specifically, among the multiple constructed acceleration card topologies, at least the following two topologies are included: 1) All acceleration cards are connected to the same central processing unit, and the acceleration cards are connected to each other through a first communication bus (i.e., the topology shown in Figure 2 ).

[0043] 2) The acceleration cards are divided into at least two groups. The acceleration cards in different groups are connected to different central processing units of the same server. The central processing units connecting the acceleration cards in different groups are connected to each other through a second communication bus. The acceleration cards in the same group are connected to each other through a third communication bus. The acceleration cards in different groups are connected to each other through a second communication bus (i.e., the topology shown in Figure 3 ).

[0044] 3) The acceleration cards are divided into at least two groups. The acceleration cards in different groups are connected to the same central processing unit of the same server. The acceleration cards in the same group are connected to each other through a fourth communication bus. The acceleration cards in different groups are connected to each other through a fifth communication bus (i.e., the topology shown in Figure 4 ).

[0045] 4) The acceleration cards are divided into at least two groups. The acceleration cards in different groups are set in different servers. The servers are connected to each other through a sixth communication bus (i.e., the topology shown in Figure 5 ).

[0046] For example, assuming that the number of acceleration cards required to run the target model is 6, then an acceleration card topology E1 including 6 acceleration cards can be constructed according to the architecture shown in Figure 2 . At the same time, the 6 acceleration cards can also be divided into two groups, with 3 acceleration cards in each group, and an acceleration card topology E2 can be constructed according to the architecture shown in Figure 3 , and an acceleration card topology E3 can be constructed according to the architecture shown in Figure 4 , and an acceleration card topology E4 can be constructed according to the architecture shown in Figure 5 .

[0047] Of course, it can be understood that the accelerator card topology architecture is not limited to Figures 2 to 5 the topology architecture shown. In practical applications, the accelerator card topology structure can be constructed according to actual needs. For example, assuming that the number of accelerator cards required when running the target model is 12, the accelerator cards can be divided into 3 groups. The first group includes 2 accelerator cards, the second group includes 4 accelerator cards, and the third group includes 6 accelerator cards. Among them, the accelerator cards in the first group are deployed in server S1, and all the accelerator cards in the first group are connected to the same central processing unit in server S1. The accelerator cards in the second group are deployed in server S2, and all the accelerator cards in the second group are connected to the same central processing unit in server S2. The accelerator cards in the third group are deployed in server S3, and the accelerator cards in the third group are divided into two subgroups. The accelerator cards in the first subgroup are connected to the central processing unit U1 in server S3, and the accelerator cards in the second subgroup are connected to the central processing unit U2 in server S3. The present application does not limit the specific form of the accelerator card topology structure.

[0048] Step S604, determine the model performance metrics when running the target model using each accelerator card topology architecture.

[0049] In this embodiment, the model performance metrics may include at least one of the latency duration and throughput of the target model. In some other embodiments, the model performance metrics may also include other metrics in addition to the latency duration and throughput, such as the hardware resource utilization rate of the target model and the queue latency (i.e., the duration of requests waiting to be executed). The following takes the model performance metrics including the latency duration and throughput of the target model as an example for illustration.

[0050] Based on Figures 2 to 5 it can be known that since the connection methods between accelerator cards are different in different accelerator card topology architectures, even if the number of accelerator cards is the same, the communication and synchronization durations in different accelerator card topology architectures can be different. Furthermore, the latency duration and throughput when running the same target model with different accelerator card topology architectures can be different.

[0051] Specifically, when calculating the latency duration and throughput when running the target model with each accelerator card topology architecture, it can be calculated based on the bandwidth, computing power of the accelerator card, the amount of data to be transmitted or loaded, etc. of each accelerator card topology structure. In the subsequent embodiments of the present application, taking the model based on the self-attention mechanism as an example, it is elaborated on how to calculate the latency duration and throughput corresponding to each accelerator card topology structure, which will not be elaborated here.

[0052] Step S605: If there is a first accelerator card topology architecture whose model performance indicators meet the preset conditions, the number of servers used to run the target model is determined based on the first accelerator card topology architecture, and accelerator cards are deployed in the servers used to run the target model.

[0053] Specifically, the model performance indicators include the target model's latency and throughput as an example. The preset conditions refer to the expected latency and throughput that the target model needs to achieve. For any accelerator card topology, if the latency corresponding to that accelerator card topology is lower than the expected latency and the throughput is higher than the expected throughput, that accelerator card topology can be considered the first accelerator card topology architecture that meets the preset conditions.

[0054] For example, if the expected latency is 30 milliseconds and the expected throughput is 20 tokens / second, the latency and throughput corresponding to the accelerator card topologies E1 to E4 are as follows: Accelerator card topology E1: latency 20 milliseconds, throughput 25 tokens / second; Accelerator card topology E2: 25 milliseconds latency, 17 tokens / second throughput; Accelerator card topology E3: 35 milliseconds latency, 10 tokens / second throughput; Accelerator card topology E4: latency is 40 milliseconds and throughput is 5 tokens / second.

[0055] In the above example, since only the latency and throughput corresponding to accelerator card topology E1 meet the preset conditions, the number of servers used to run the target model and the deployment of accelerator cards within these servers can be determined based on accelerator card topology E1. For example, since there is only one server in accelerator card topology E1, the number of servers running the target model can be determined as 1, and the connection relationships between accelerator cards within this single server can be configured based on accelerator card topology E1.

[0056] In summary, in the technical solutions of some embodiments of the present application, on the one hand, based on the storage occupancy during the operation of the target model, the number of acceleration cards required to run the target model is determined. In this way, it can be ensured that the sum of the storage capacities of the acceleration cards matches the storage occupancy of the target model, avoiding the problem of waste or insufficiency of acceleration card resources. On the other hand, according to the number of acceleration cards, multiple acceleration card topologies are constructed, and based on the model performance metrics when the target model is run on each acceleration card topology, the first acceleration card topology whose model performance metrics meet the preset conditions is selected as the architecture reference for deploying the acceleration cards. In this way, when the target model is run on the acceleration cards, it can be ensured that the performance of the target model meets the expected performance requirements. Based on the above two aspects, the hardware resource configuration of the target model can be matched with the computing requirements, thereby solving the problem of mismatch between the hardware resource configuration and the computing requirements of the model in the related art.

[0057] In some embodiments, if there is no first acceleration card topology, the method of the present application may further include: Increasing the number of acceleration cards required to run the target model; According to the increased number of acceleration cards, reconstructing multiple different acceleration card topologies; Among the reconstructed multiple acceleration card topologies, determining the model performance metrics when the target model is run using each acceleration card topology; If there is a second acceleration card topology whose model performance metrics meet the preset conditions among the reconstructed multiple acceleration card topologies, then according to the second acceleration card topology, determine the number of servers for running the target model, and deploy the acceleration cards in the servers for running the target model.

[0058] Specifically, when increasing the number of acceleration cards required to run the target model, the number of acceleration cards determined in step S602 is used as the reference number, and the newly added number of acceleration cards needs to be greater than the reference number. For example, assuming that the number of acceleration cards determined in step S602 is 10, then 10 is used as the reference number, the number of acceleration cards is increased to 12, and multiple different acceleration card topologies are reconstructed according to the specifications of 12 acceleration cards.

[0059] Since during the operation of the target model, the communication and synchronization duration is usually much lower than the computing duration, therefore, although the communication and synchronization duration between the acceleration cards will increase after increasing the number of acceleration cards, by using more acceleration cards to execute computing tasks in parallel, the computing duration can be greatly shortened, and thus the latency duration of the target model can be reduced and the throughput of the target model can be improved.

[0060] In addition, since the number of acceleration cards determined in step S602 can already meet the storage space requirements during the operation of the target model, the acceleration card topology structure is reconstructed based on the increased number of acceleration cards, and the sum of the storage capacities of the acceleration cards can still meet the storage space requirements of the target model.

[0061] In the above embodiment, by finding the second acceleration card topology structure by increasing the number of acceleration cards, the number of acceleration cards and the acceleration card topology structure can be explored to ensure that the finally determined acceleration card topology structure can meet the performance requirements and storage space requirements of the target model, and at the same time, reduce the waste of hardware resources.

[0062] In some embodiments, the method of the present application further includes: If there are multiple first acceleration card topologies or multiple second acceleration card topologies whose model performance indicators meet the preset conditions, then among the acceleration card topologies that meet the preset conditions, the costs corresponding to each acceleration card topology are determined respectively, and the costs include at least one of the hardware procurement cost and the hardware operation cost; According to the acceleration card topology structure with the lowest cost, the number of servers for running the target model is determined, and acceleration cards are deployed in the servers for running the target model.

[0063] Specifically, the hardware procurement cost includes, but is not limited to, the procurement costs of servers and acceleration cards. The hardware operation cost refers to the expenses generated during the operation of the acceleration cards or servers after the acceleration cards are deployed in the servers according to each acceleration card topology structure, such as the power consumption, heat dissipation cost, and maintenance cost of the acceleration cards. Different acceleration card architectures may have different operation costs.

[0064] In the above embodiment, among multiple acceleration card topologies whose delay duration and throughput meet the preset conditions, determining the number of servers and the acceleration card deployment method according to the acceleration card topology structure with the lowest cost can minimize the operation cost of the target model under the condition that the performance requirements of the target model meet the actual needs.

[0065] In some embodiments, considering that different types of acceleration cards have their own advantages, when constructing multiple acceleration card topologies in step S603, acceleration card topologies can also be constructed using multiple different types of acceleration cards. For example, when a Graphics Processing Unit (GPU) is used as an acceleration card, it can have the characteristic of low latency; when a Tensor Processing Unit (TPU) is used as an acceleration card, it can have the characteristic of high throughput; when a Data Processing Unit (DPU) is used as an acceleration card, by offloading tasks such as communication and storage management, the waiting time of the graphics processing unit and the tensor processing unit can be reduced. Therefore, different types of acceleration cards such as graphics processing units, tensor processing units, and data processing units can be used simultaneously in the same acceleration card topology. For example, neural network layers that need to perform data preprocessing can be deployed in the graphics processing unit, and neural network layers that need to update model parameters or perform high-throughput inference can be deployed in the tensor processing unit.

[0066] Taking the target model as a model based on the self-attention mechanism as an example, the following describes how to determine the storage occupancy during the operation of the target model, as well as how to determine the latency and throughput of the target model.

[0067] Specifically, in the target model based on the self-attention mechanism, the model architecture of the target model can be divided into a traditional Transformer architecture and a MoE-Transformer architecture. In the traditional Transformer architecture, the target model can include self-attention layers (usually multiple), a fully connected feed-forward network (FFN), a normalization layer, and an embedding layer. Among them, the parameters of the self-attention layer can include a Q matrix, a K matrix, a V matrix, and an output projection matrix. The parameters of the fully connected feed-forward network can include an upscaling matrix and a downscaling matrix. The parameters of the normalization layer can include a scaling factor and an offset term. The parameters of the embedding layer can include a positional embedding matrix and tokens. The parameters of these layers can be used as the model parameters of the target model. In the MoE-Transformer architecture, the fully connected feed-forward network is split into multiple independent network layers, and these independent network layers are also called expert layers. After the model training is completed, different expert layers tend to handle tasks in specific domains. For example, expert layer S1 tends to handle computational tasks in the code domain, and expert layer S2 tends to handle computational tasks in the mathematics domain. Tokens in the inference process will be routed to one or more of these expert layers for processing. Usually, the expert layer to which the token is routed is called the activated expert layer, and the expert layer not routed to is called the unactivated expert layer. Additionally, similar to the traditional Transformer architecture, in the MoE-Transformer architecture, neural network layers such as the attention layer, the normalization layer, and the embedding layer can also be included, and neural network layers such as the attention layer, the normalization layer, and the embedding layer can be called non-expert layers.

[0068] Furthermore, for the target model based on the traditional architecture and the target model based on the MoE-Transformer architecture, during the inference process, both models will obtain key-value pairs of the text sequence to be inferred based on the Q matrix, K matrix, and V matrix of each self-attention layer. The key-value pairs calculated by the front-end self-attention layer will be saved to the storage unit of the acceleration card. The back-end self-attention layer will read the key-value pairs calculated by the front-end self-attention layer from the storage unit and complete subsequent calculations based on the read key-value pairs. Therefore, during the operation of the target model, it is necessary to pre-plan the storage space for these key-value pairs, that is, the storage occupancy corresponding to these keys will be included in the storage occupancy of the target model. Additionally, when both models are in inference, each layer will output intermediate results, and these intermediate results include activation values. The storage occupancy corresponding to these activation values will also be included in the storage occupancy of the target model.

[0069] Based on the above description, in some embodiments, obtaining the storage occupancy during the operation of the target model in step S601 may include: If the target model does not include an expert layer (i.e., the traditional Transformer architecture), obtain the first storage occupancy corresponding to the model parameters of the target model, the second storage occupancy corresponding to the key-value pairs generated during the operation of the target model, and the third storage occupancy corresponding to the activation values generated during the operation of the target model; Based on the first storage occupancy, the second storage occupancy, the third storage occupancy, and the maximum number of concurrent requests supported by the target model, determine the storage occupancy during the operation of the target model.

[0070] Specifically, the storage occupancy required for the target model can be determined based on expression (1) .

[0071] (1) where represents the first storage occupancy corresponding to the model parameters of the target model, represents the second storage occupancy corresponding to the key-value pairs generated during the operation of the target model, represents the third storage occupancy corresponding to the activation values generated during the operation of the target model, and b represents the maximum number of concurrent requests supported by the target model.

[0072] In some embodiments, the above-mentioned obtaining of the first storage occupancy corresponding to the model parameters of the target model may include: Obtain the number of parameters of the target model and the quantization precision of each model parameter; Based on the number of parameters and the quantization precision of the model parameters, determine the first storage occupancy.

[0073] Specifically, the quantization precision is used to represent the number of bytes occupied by a single model parameter. The quantization precision may include but is not limited to FP32, FP16, and INT8. When the quantization precision is FP32, it means that a single model parameter needs to occupy 4 bytes; when the quantization precision is FP16, it means that a single model parameter needs to occupy 2 bytes; when the quantization precision is INT8, it means that a single model parameter needs to occupy 1 byte. Usually, the number of parameters and the quantization precision of the target model can be provided by the designer of the target model.

[0074] After obtaining the number of parameters and the quantization precision of the target model, the first storage occupancy corresponding to the model parameters of the target model can be determined based on expression (2) .

[0075] (2) where N represents the number of parameters of the target model, and q represents the quantization precision of the model parameters.

[0076] Further, the obtaining of the second storage occupancy and the third storage occupancy may include: Obtain the number of hidden layers of the target model, the first hidden dimension of a single hidden layer, the maximum sequence length to be input into the target model, and the maximum number of tokens to be output by the target model; Determine the second storage occupancy and the third storage occupancy based on the number of hidden layers, the first hidden dimension, the maximum sequence length, and the maximum number of tokens.

[0077] Specifically, the second storage occupancy corresponding to the key-value pairs generated during the operation of the target model can be determined based on Expression (3) .

[0078] (3) where b represents the maximum concurrency supported by the target model, represents the maximum sequence length to be input into the target model, n represents the maximum number of tokens to be output by the target model, represents the number of hidden layers of the target model, represents the first hidden dimension of a single hidden layer, and q represents the quantization precision of the model parameters.

[0079] And the third storage occupancy corresponding to the activation values generated during the operation of the target model can be determined based on Expression (4) .

[0080] (4) For the description of the relevant parameters in Expression (4), reference can be made to Expression (3), which will not be elaborated here.

[0081] In the above calculation process, the maximum sequence length, the maximum number of sequences, the maximum number of tokens, and the maximum concurrency supported by the target model are used, which can ensure that the finally determined storage occupancy meets the actual requirements of the target model and avoid problems such as insufficient storage space required by the target model.

[0082] In some embodiments, the obtaining of the storage occupancy during the operation of the target model in step S601 may include: If the target model does not include an expert layer (i.e., the traditional Transformer architecture), obtain the first storage occupancy corresponding to the model parameters of the target model, the activation value redundancy coefficient, and the key-value pair redundancy coefficient of the target model; Determine the storage occupancy based on the first storage occupancy, the activation value redundancy coefficient, and the key-value pair redundancy coefficient.

[0083] Specifically, the activation value redundancy coefficient and key-value pair redundancy coefficient of the target model can be determined based on experience. After obtaining the first storage occupancy corresponding to the model parameters of the target model, the activation value redundancy coefficient, and the key-value pair redundancy coefficient of the target model, the storage occupancy during the operation of the target model can be determined based on Expression (5). .

[0084] (5) Among them, represents the first storage occupancy corresponding to the model parameters of the target model, K1 represents the activation value redundancy coefficient of the target model, and K2 represents the key-value pair redundancy coefficient of the target model. The calculation formula for the first storage occupancy can be referred to the above Expression (2), which will not be elaborated here. The activation value redundancy coefficient and the key-value pair redundancy coefficient can be between 0.2 and 0.5.

[0085] The above Expression (1) and Expression (5) can be regarded as different calculation methods for calculating the storage occupancy of the target model under the traditional Transformer architecture. Compared with Expression (1), the calculation method in Expression (5) can avoid calculating the second storage occupancy corresponding to the key-value pairs generated during the operation of the target model and the third storage occupancy corresponding to the activation values generated during the operation of the target model, thus greatly simplifying the calculation logic.

[0086] In some embodiments, obtaining the storage occupancy during the operation of the target model in step S601 may include: If the target model includes an expert layer and a non-expert layer (i.e., the MoE-Transformer architecture), obtain the first storage occupancy corresponding to the model parameters of the target model, the fourth storage occupancy corresponding to the parameters of the activated expert layer, the fifth storage occupancy corresponding to the key-value pairs and activation values of the expert layer, and the sixth storage occupancy corresponding to the key-value pairs and activation values of the non-expert layer; Based on the first storage occupancy, the fourth storage occupancy, the fifth storage occupancy, the sixth storage occupancy, and the maximum number of concurrent requests supported by the target model, determine the storage occupancy during the operation of the target model.

[0087] Specifically, the storage occupancy required for the target model of the MoE-Transformer architecture can be calculated based on Expression (6). .

[0088]

[0089] Among them, represents the first storage occupancy corresponding to the model parameters of the target model. Represents the fourth storage occupancy corresponding to the parameters of the activated expert layer. Represents the sixth storage occupancy corresponding to the key-value pairs and activation values of the non-expert layer, where is the storage occupancy corresponding to the key-value pairs of the non-expert layer, is the storage occupancy corresponding to the activation values of the non-expert layer. Represents the fifth storage occupancy corresponding to the key-value pairs and activation values of the expert layer, where is the storage occupancy corresponding to the key-value pairs of the expert layer, is the storage occupancy corresponding to the activation values of the expert layer. b represents the maximum number of concurrent requests that the target model needs to support.

[0090] In some embodiments, the above obtaining the fourth storage occupancy includes: Obtaining the maximum number of activated expert layers supported by the target model and the maximum number of parameters included in a single expert layer, and determining the fourth storage occupancy based on the maximum number of expert layers and the maximum number of parameters.

[0091] Specifically, the maximum number of activated expert layers supported by the target model and the maximum number of parameters included in a single expert layer can be provided by the designer of the target model. Multiplying the maximum number of expert layers by the maximum number of parameters included in a single expert layer can obtain the number of parameters of the activated expert layer. Furthermore, similar to the above expression (2), based on the number of parameters of the activated expert layer and the quantization precision of the target model, the fourth storage occupancy can be determined.

[0092] In some embodiments, the above obtaining the fifth storage occupancy may include: Obtaining the total number of expert layers of the target model, the compression dimension of a single expert layer, the maximum number of allowed activated expert layers, the maximum sequence length to be input to the target model, and the maximum number of tokens to be output by the target model; Determining the fifth storage occupancy based on the total number of expert layers, the compression dimension, the maximum number of allowed activated expert layers, the maximum sequence length, and the maximum number of tokens.

[0093] Specifically, the storage occupancy corresponding to the key-value pairs of the expert layer can be determined based on expression (7) and the storage occupancy corresponding to the activation values of the expert layer can be determined based on expression (8) and adding to The resulting sum is the fifth storage occupancy.

[0094]

[0095]

[0096] In expressions (7) and (8), the parameters b, s, n, and q can be referred to the above - related descriptions and will not be elaborated here. represents the total number of expert layers of the target model, represents the maximum number of expert layers allowed to be activated, represents the compression dimension of a single expert layer.

[0097] In some embodiments, the above - mentioned obtaining the sixth storage occupancy may include: Obtaining the number of non - expert layers of the target model, the second hidden dimension of a single non - expert layer, the maximum sequence length to be input into the target model, and the maximum number of tokens to be output by the target model; Based on the number of non - expert layers, the second hidden dimension, the maximum sequence length, and the maximum number of tokens, determine the sixth storage occupancy.

[0098] Specifically, the storage occupancy corresponding to the key - value pairs of the non - expert layers can be determined based on expression (9) and the storage occupancy corresponding to the activation values of the non - expert layers can be determined based on expression (10) , and and are added together, and the resulting value is the sixth storage occupancy.

[0099]

[0100]

[0101] In expressions (9) and (10), the parameters b, s, n, and q can be referred to the above - related descriptions and will not be elaborated here. represents the total number of non - expert layers, represents the second hidden dimension of a single non - expert layer.

[0102] In some embodiments, when the model performance metric includes the latency duration of the target model, determining the model performance metric when running the target model using each accelerator card topology architecture in step S604 may include: For any target accelerator card topology architecture among multiple accelerator card topologies, obtain the first - token duration and the average - token duration when the target accelerator card topology architecture runs the target model. The first - token duration represents the time consumption from when the target model receives the sequence to be inferred until the target model outputs the first token, and the average - token duration represents the average time consumption of the target model outputting each subsequent token after the target model outputs the first token; Based on the first - token duration, the average - token duration, and the maximum number of tokens to be output by the target model, determine the latency duration when the target accelerator card topology architecture runs the target model.

[0103] Specifically, the calculation process of the above delay duration can be expressed by expression (11).

[0104] (11) Among them, represents the delay duration of the target model, represents the first token duration, represents the average token duration, represents the number of generated tokens.

[0105] In some embodiments, the above obtaining the first token duration when the target acceleration card topology runs the target model may include: Obtaining the first calculation duration and the first communication duration of the target model. The first calculation duration refers to the calculation duration required from the time when the target model receives the sequence to be inferred until the target model outputs the first token. The first communication duration refers to the communication duration required from the time when the target model receives the sequence to be inferred until the target model outputs the first token; Based on the first calculation duration and the first communication duration, obtain the first token duration.

[0106] Specifically, the total data volume of the model parameters and initial activation values of the target model, the communication bandwidth of the target acceleration card topology, and the total computing power of all acceleration cards in the target acceleration card topology can be obtained, and based on the total data volume and the total computing power, determine the first calculation duration, and based on the total data volume and the communication bandwidth, determine the first communication duration.

[0107] For example, taking tensor parallelism as an example, the first token duration between the multiple acceleration cards can be calculated based on expression (12) .

[0108]

[0109] Among them, and The meanings of can refer to the above relevant descriptions and will not be elaborated here. is the computing power of a single acceleration card (that is, the number of multiply-add operations that a single acceleration card can complete per second). represents the calculated number of acceleration cards. It should be noted that in the case where the number of acceleration cards needs to be increased, is the increased number of acceleration cards. represents the bandwidth between acceleration cards. represents the bandwidth of the storage module of the acceleration card. represents the number of times each acceleration card needs to communicate with other acceleration cards. represents the storage occupancy corresponding to the activation initial value.

[0110] Further, in the above expression (12), may represent the first calculation duration, and may represent the first communication duration.

[0111] In some embodiments, the average token duration when obtaining the target acceleration card topology to run the target model may include: Obtaining the second calculation duration and the second communication duration of the target model. The second calculation duration refers to the average calculation duration required for the target model to obtain each of the other tokens after the target model outputs the first token. The second communication duration refers to the average communication duration required for the target model to obtain each of the other tokens after the target model outputs the first token; Based on the second calculation duration and the second communication duration, obtain the average token duration.

[0112] For example, taking tensor parallelism as an example, the first token duration between the multiple acceleration cards can be calculated based on expression (13) .

[0113]

[0114] Expression (13) is basically similar to expression (12). The main difference is that in expression (12), since it is calculating the time taken for the first token, there is no transmission time for key-value pairs. That is, in expression (12), when calculating the first communication duration, there is only and , and the storage occupancy of key-value pairs is not included. However, in expression (13), since it is calculating the average time taken for each token after the second token, at this time, key-value pairs are already included, so the storage occupancy of key-value pairs needs to be counted, that is, = .

[0115] In addition, it should be noted that although both expression (12) and expression (13) include , the value can be different. For example, in expression (12) represents the storage occupancy corresponding to the activation initial value, and in expression (13) represents the storage occupancy corresponding to the activation value that needs to be loaded when calculating each token after the second token.

[0116] Thus far, the relevant description of the solution of this application is completed.

[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0118] Referring to Figure 7 , which is a schematic diagram of modules of an acceleration card deployment device provided for some embodiments of the present application. Figure 7 In , the acceleration card deployment device includes: A storage occupancy acquisition module 701, configured to acquire the storage occupancy during the operation of the target model; An acceleration card quantity determination module 702, configured to determine the number of acceleration cards required for running the target model according to the storage occupancy and the storage capacity of each acceleration card; A topology construction module 703, configured to construct multiple acceleration card topologies according to the number of acceleration cards, wherein, in different acceleration card topologies, the connection modes between acceleration cards are different; A calculation module 704, configured to determine the model performance metrics when running the target model using each acceleration card topology; An acceleration card deployment module 705, configured to, if there is a first acceleration card topology whose model performance metrics meet the preset conditions, determine the number of servers for running the target model according to the first acceleration card topology, and deploy acceleration cards in the servers for running the target model.

[0119] In some embodiments, if there is no first acceleration card topology, the topology construction module 703 is further configured to: Increase the number of acceleration cards required for running the target model; Re-construct multiple different acceleration card topologies according to the increased number of acceleration cards; In the multiple re-constructed acceleration card topologies, determine the model performance metrics when running the target model using each acceleration card topology; If there is a second acceleration card topology whose model performance metrics meet the preset conditions among the multiple re-constructed acceleration card topologies, determine the number of servers for running the target model according to the second acceleration card topology, and deploy acceleration cards in the servers for running the target model.

[0120] In some embodiments, the calculation module 704 is further configured to: If there are multiple first acceleration card topologies or multiple second acceleration card topologies whose model performance metrics meet the preset conditions, respectively determine the costs corresponding to each acceleration card topology among the acceleration card topologies that meet the preset conditions, where the costs include at least one of the hardware procurement cost and the hardware operation cost; Determine the number of servers for running the target model according to the acceleration card topology with the lowest cost, and deploy acceleration cards in the servers for running the target model.

[0121] In some embodiments, among the multiple accelerator card topologies constructed by the topology construction module 703, at least two of the following topologies are included: All accelerator cards are connected to the same central processing unit, and the accelerator cards are connected to each other through a first communication bus; The accelerator cards are divided into at least two groups. The accelerator cards in different groups are connected to different central processing units of the same server. The central processing units connecting the accelerator cards in different groups are connected through a second communication bus. The accelerator cards in the same group are connected through a third communication bus, and the accelerator cards in different groups are connected through the second communication bus; The accelerator cards are divided into at least two groups. The accelerator cards in different groups are connected to the same central processing unit of the same server. The accelerator cards in the same group are connected through a fourth communication bus, and the accelerator cards in different groups are connected through a fifth communication bus; The accelerator cards are divided into at least two groups. The accelerator cards in different groups are arranged in different servers, and the servers are connected through a sixth communication bus.

[0122] In some embodiments, the target model is a model based on the self-attention mechanism; the storage occupancy acquisition module 701 is specifically configured to: If the target model does not include an expert layer, obtain a first storage occupancy corresponding to the model parameters of the target model, a second storage occupancy corresponding to the key-value pairs generated during the operation of the target model, and a third storage occupancy corresponding to the activation values generated during the operation of the target model; Based on the first storage occupancy, the second storage occupancy, the third storage occupancy, and the maximum number of concurrent requests that the target model needs to support, determine the storage occupancy during the operation of the target model.

[0123] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to: Obtain the number of hidden layers of the target model, the first hidden dimension of a single hidden layer, the maximum sequence length to be input to the target model, and the maximum number of tokens to be output by the target model; Based on the number of hidden layers, the first hidden dimension, the maximum sequence length, and the maximum number of tokens, determine the second storage occupancy and the third storage occupancy.

[0124] In some embodiments, the target model is a model based on the self-attention mechanism; the storage occupancy acquisition module 701 is specifically configured to: If the target model does not include an expert layer, obtain a first storage occupancy corresponding to the model parameters of the target model, an activation value redundancy coefficient, and a key-value pair redundancy coefficient of the target model; Based on the first storage occupancy, the activation value redundancy coefficient, and the key-value pair redundancy coefficient, determine the storage occupancy.

[0125] In some embodiments, the target model is a model based on the self-attention mechanism; the storage occupancy acquisition module 701 is specifically configured to: If the target model includes an expert layer and a non-expert layer, obtain the first storage occupancy corresponding to the model parameters of the target model, the fourth storage occupancy corresponding to the parameters of the activated expert layer, the fifth storage occupancy corresponding to the key-value pairs and activation values of the expert layer, and the sixth storage occupancy corresponding to the key-value pairs and activation values of the non-expert layer; Based on the first storage occupancy, the fourth storage occupancy, the fifth storage occupancy, the sixth storage occupancy, and the maximum number of concurrent requests that the target model needs to support, determine the storage occupancy during the operation of the target model.

[0126] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to: Obtain the maximum number of activated expert layers supported by the target model and the maximum number of parameters included in a single expert layer, and determine the fourth storage occupancy based on the maximum number of expert layers and the maximum number of parameters.

[0127] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to: Obtain the total number of expert layers of the target model, the compression dimension of a single expert layer, the maximum number of expert layers allowed to be activated, the maximum sequence length to be input to the target model, and the maximum number of tokens to be output by the target model; Based on the total number of expert layers, the compression dimension, the maximum number of expert layers allowed to be activated, the maximum sequence length, and the maximum number of tokens, determine the fifth storage occupancy.

[0128] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to: Obtain the number of non-expert layers of the target model, the second hidden dimension of a single non-expert layer, the maximum sequence length to be input to the target model, and the maximum number of tokens to be output by the target model; Based on the number of non-expert layers, the second hidden dimension, the maximum sequence length, and the maximum number of tokens, determine the sixth storage occupancy.

[0129] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to: Obtain the number of parameters of the target model and the quantization precision of each model parameter; Based on the number of parameters and the quantization precision of the model parameters, determine the first storage occupancy.

[0130] In some embodiments, the storage occupancy acquisition module 701 is specifically configured to: The first token duration and the average token duration when the target model runs on the acceleration card topology. The first token duration represents the time taken from when the target model receives the sequence to be inferred until the target model outputs the first token. The average token duration represents the average time taken for the target model to output each subsequent token after the first token is output by the target model; Based on the first token duration, the average token duration, and the maximum number of tokens to be output by the target model, determine the latency duration when the target acceleration card topology runs the target model.

[0131] In some embodiments, the calculation module 704 is specifically configured to: Obtain the first calculation duration and the first communication duration of the target model. The first calculation duration refers to the calculation duration required from when the target model receives the sequence to be inferred until the target model outputs the first token. The first communication duration refers to the communication duration required from when the target model receives the sequence to be inferred until the target model outputs the first token; Based on the first calculation duration and the first communication duration, obtain the first token duration.

[0132] In some embodiments, the calculation module 704 is specifically configured to: Obtain the total data volume of the model parameters and the initial activation values of the target model, the communication bandwidth of the target acceleration card topology, and the total computing power of all acceleration cards in the target acceleration card topology; Based on the total data volume and the total computing power, determine the first calculation duration; Based on the total data volume and the communication bandwidth, determine the first communication duration.

[0133] In some embodiments, the calculation module 704 is specifically configured to: Obtain the second calculation duration and the second communication duration of the target model. The second calculation duration refers to the average calculation duration required for the target model to obtain each subsequent token after the first token is output. The second communication duration refers to the average communication duration required for the target model to obtain each subsequent token after the first token is output; Based on the second calculation duration and the second communication duration, obtain the average token duration.

[0134] Refer to Figure 8 , Embodiments of the present application also provide an electronic device, including a memory 10 and a processor 20. The memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any of the above-described embodiments of the acceleration card deployment method.

[0135] Embodiments of the present application also provide a computer non-volatile readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the acceleration card deployment method when running.

[0136] In an exemplary embodiment, the above computer non-volatile readable storage medium may include, but is not limited to: USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs, etc., all kinds of media that can store computer programs.

[0137] Embodiments of the present application also provide a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the acceleration card deployment method.

[0138] Embodiments of the present application also provide another computer program product, including a non-volatile computer non-volatile readable storage medium. The non-volatile computer non-volatile readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the acceleration card deployment method.

[0139] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0140] The above has introduced in detail an acceleration card deployment method, device, equipment, storage medium, and program product provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for accelerating card deployment, characterized in that, The method includes: Obtaining the storage occupancy during the operation of the target model; Determining the number of acceleration cards required to run the target model based on the storage occupancy and the storage capacity of each acceleration card; Constructing multiple acceleration card topologies based on the number of acceleration cards, where the connection methods between acceleration cards are different in different acceleration card topologies; Determining the model performance metrics when running the target model using each acceleration card topology; If there is a first acceleration card topology whose model performance metrics meet the preset conditions, determining the number of servers for running the target model based on the first acceleration card topology, and deploying acceleration cards in the servers for running the target model.

2. The method according to claim 1, wherein If there is no such first acceleration card topology, the method further includes: Increasing the number of acceleration cards required to run the target model; Reconstructing multiple different acceleration card topologies based on the increased number of acceleration cards; Determining the model performance metrics when running the target model using each acceleration card topology in the reconstructed multiple acceleration card topologies; If there is a second acceleration card topology whose model performance metrics meet the preset conditions in the reconstructed multiple acceleration card topologies, determining the number of servers for running the target model according to the second acceleration card topology, and deploying acceleration cards in the servers for running the target model.

3. The method according to claim 2, wherein The method further includes: If there are multiple first acceleration card topologies or multiple second acceleration card topologies whose model performance metrics meet the preset conditions, respectively determining the costs corresponding to each acceleration card topology among the acceleration card topologies that meet the preset conditions, where the cost includes at least one of the hardware procurement cost and the hardware operation cost; Determining the number of servers for running the target model according to the acceleration card topology with the lowest cost, and deploying acceleration cards in the servers for running the target model.

4. The method according to claim 1 or 2, characterized in that, Among the multiple constructed acceleration card topologies, at least two of the following topologies are included: All acceleration cards are connected to the same central processing unit, and the acceleration cards are connected through a first communication bus; The acceleration cards are divided into at least two groups. The acceleration cards in different groups are connected to different central processing units of the same server. The central processing units connecting the acceleration cards in different groups are connected through a second communication bus. The acceleration cards in the same group are connected through a third communication bus, and the acceleration cards in different groups are connected through the second communication bus; The acceleration cards are divided into at least two groups. The acceleration cards in different groups are connected to the same central processing unit of the same server. The acceleration cards in the same group are connected through a fourth communication bus, and the acceleration cards in different groups are connected through a fifth communication bus; The acceleration cards are divided into at least two groups. The acceleration cards in different groups are arranged in different servers, and the servers are connected through a sixth communication bus.

5. The method according to claim 1, characterized in that, The target model is a model based on the self-attention mechanism; The obtaining the storage occupancy during the operation of the target model includes: If the target model does not include an expert layer, obtain the first storage occupancy corresponding to the model parameters of the target model, the second storage occupancy corresponding to the key-value pairs generated during the operation of the target model, and the third storage occupancy corresponding to the activation values generated during the operation of the target model; Based on the first storage occupancy, the second storage occupancy, the third storage occupancy, and the maximum number of concurrent requests that the target model needs to support, determine the storage occupancy during the operation of the target model.

6. The method according to claim 5, wherein Obtaining the second storage occupancy and the third storage occupancy includes: Obtain the number of hidden layers of the target model, the first hidden dimension of a single hidden layer, the maximum sequence length to be input into the target model, and the maximum number of tokens to be output by the target model; Based on the number of hidden layers, the first hidden dimension, the maximum sequence length, and the maximum number of tokens, determine the second storage occupancy and the third storage occupancy.

7. The method according to claim 1, wherein The target model is a model based on the self-attention mechanism; The obtaining of the storage occupancy during the operation of the target model includes: If the target model does not include an expert layer, obtain the first storage occupancy corresponding to the model parameters of the target model, the activation value redundancy coefficient, and the key-value pair redundancy coefficient of the target model; Based on the first storage occupancy, the activation value redundancy coefficient, and the key-value pair redundancy coefficient, determine the storage occupancy.

8. The method according to claim 1, wherein The target model is a model based on the self-attention mechanism; The obtaining of the storage occupancy during the operation of the target model includes: If the target model includes an expert layer and a non-expert layer, obtain the first storage occupancy corresponding to the model parameters of the target model, the fourth storage occupancy corresponding to the parameters of the activated expert layer, the fifth storage occupancy corresponding to the key-value pairs and activation values of the expert layer, and the sixth storage occupancy corresponding to the key-value pairs and activation values of the non-expert layer; Based on the first storage occupancy, the fourth storage occupancy, the fifth storage occupancy, the sixth storage occupancy, and the maximum number of concurrent requests that the target model needs to support, determine the storage occupancy during the operation of the target model.

9. The method according to claim 8, characterized in that, Obtaining the fourth storage occupancy includes: Obtain the maximum number of expert layers that the target model supports to be activated and the maximum number of parameters included in a single expert layer, and determine the fourth storage occupancy based on the maximum number of expert layers and the maximum number of parameters.

10. The method according to claim 8, wherein Obtaining the fifth storage occupancy includes: Obtain the total number of expert layers of the target model, the compression dimension of a single expert layer, the maximum number of expert layers allowed to be activated, the maximum sequence length to be input into the target model, and the maximum number of tokens to be output by the target model; Based on the total number of expert layers, the compression dimension, the maximum number of expert layers allowed to be activated, the maximum sequence length, and the maximum number of tokens, determine the fifth storage occupancy.

11. The method according to claim 8, wherein Obtaining the sixth storage occupancy includes: Obtain the number of non-expert layers of the target model, the second hidden dimension of a single non-expert layer, the maximum sequence length to be input into the target model, and the maximum number of tokens to be output by the target model; Determine the sixth storage occupancy based on the number of non-expert layers, the second hidden dimension, the maximum sequence length, and the maximum number of tokens.

12. The method according to any one of claims 5, 7, and 8, characterized in that, Obtain the first storage occupancy corresponding to the model parameters of the target model, including: Obtain the number of parameters of the target model and the quantization precision of each model parameter; Determine the first storage occupancy based on the number of parameters and the quantization precision of the model parameters.

13. The method according to claim 1 or 2, characterized in that, The model performance metrics include the latency duration of the target model; The determining the model performance metrics when running the target model using each accelerator card topology includes: For any target accelerator card topology among the multiple accelerator card topologies, obtain the first token duration and the average token duration when the target accelerator card topology runs the target model. The first token duration represents the time taken from when the target model receives the sequence to be inferred until the target model outputs the first token, and the average token duration represents the average time taken for the target model to output each subsequent token after the target model outputs the first token; Determine the latency duration when the target accelerator card topology runs the target model based on the first token duration, the average token duration, and the maximum number of tokens to be output by the target model.

14. The method according to claim 13, wherein Obtain the first token duration when the target accelerator card topology runs the target model, including: Obtain the first computation duration and the first communication duration of the target model. The first computation duration refers to the computation duration required from when the target model receives the sequence to be inferred until the target model outputs the first token, and the first communication duration refers to the communication duration required from when the target model receives the sequence to be inferred until the target model outputs the first token; Obtain the first token duration based on the first computation duration and the first communication duration.

15. The method according to claim 14, wherein The obtaining the first computation duration and the first communication duration of the target model includes: Obtain the total data volume of the model parameters and the initial activation values of the target model, the communication bandwidth of the target accelerator card topology, and the total computing power of all accelerator cards in the target accelerator card topology; Determine the first computation duration based on the total data volume and the total computing power; Determine the first communication duration based on the total data volume and the communication bandwidth.

16. The method according to claim 13, wherein Obtain the average token duration when the target accelerator card topology runs the target model, including: Obtain the second computation duration and the second communication duration of the target model. The second computation duration refers to the average computation duration required for the target model to obtain each subsequent token after the target model outputs the first token, and the second communication duration refers to the average communication duration required for the target model to obtain each subsequent token after the target model outputs the first token; Based on the second calculation duration and the second communication duration, the average token duration is obtained.

17. An acceleration card deployment device, characterized in that, The device includes: A storage occupancy acquisition module, configured to acquire the storage occupancy during the operation of the target model; An accelerator card quantity determination module, configured to determine the quantity of accelerator cards required for running the target model according to the storage occupancy and the storage capacity of each accelerator card; A topology construction module, configured to construct multiple accelerator card topology architectures according to the quantity of accelerator cards, wherein, in different accelerator card topology architectures, the connection manners between accelerator cards are different; A calculation module, configured to determine the model performance metrics when running the target model using each accelerator card topology architecture; An accelerator card deployment module, configured to, if there is a first accelerator card topology architecture whose model performance metrics meet the preset conditions, determine the quantity of servers for running the target model according to the first accelerator card topology architecture, and deploy accelerator cards in the servers for running the target model.

18. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the accelerator card deployment method according to any one of claims 1 to 16 when executing the computer program.

19. A computer non-volatile readable storage medium, characterized in that, A computer program is stored in the computer non-volatile readable storage medium, wherein the computer program implements the accelerator card deployment method according to any one of claims 1 to 16 when executed by a processor.

20. A computer program product, comprising a computer program, characterized in that, The computer program implements the accelerator card deployment method according to any one of claims 1 to 16 when executed by a processor.

Citation Information

Patent Citations

  • Topology selection method and device for protocol operation, equipment and medium

    CN114707651A

  • Array server configuration method suitable for AIPC and server

    CN119473614A

  • GPU topology awareness scheduling method, electronic equipment and medium

    CN119938289A

  • Computing system, model training method and apparatus, and product

    WO2025001229A1