Model distributed reasoning test system based on multiple machines and multiple cards

Selecting the appropriate server through the multi-machine multi-card testing system, building a multi-machine multi-card environment and evaluating the calculation acceleration ratio curve, solving the problems of single-machine multi-card resource limitation and a single test strategy, and improving the test efficiency and optimization capabilities of the model in a distributed environment.

CN120386741APending Publication Date: 2025-07-29GUANGDONG POWER GRID CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510394890.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, the computing resources of single-machine multiple cards are limited, making it difficult to complete the testing tasks of super-large-scale and high-complexity models in a short time, and a single test strategy cannot be optimized for the differences between different models, affecting the testing efficiency of the model in a distributed environment.

Method used

A model distributed inference testing system with multiple machines and multiple cards is adopted. A suitable multi-card server is selected through the server matching module, a multi-card environment is built, and distributed inference testing is carried out in a multi-card environment according to different model testing strategies, and the acceleration ratio curve is evaluated to determine the optimal testing strategy.

Benefits of technology

It realizes the completion of a large number of test tasks in a short time, optimizes the test efficiency of the model in a distributed environment, provides an optimal testing strategy, and provides a basis for the inference optimization of subsequent models in a distributed environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386741A_ABST
    Figure CN120386741A_ABST
Patent Text Reader

Abstract

The invention provides a model distributed reasoning test system based on multiple machines and multiple cards. The system comprises a model reasoning test middle table, a server matching module, a test environment construction module, a reasoning test module, an index evaluation module and a test efficiency evaluation module. Obtaining a target multi-card server; constructing a multi-machine multi-card environment; starting a distributed reasoning test in a multi-machine multi-card environment to obtain model reasoning indexes of each model test strategy in different reasoning test stages; performing evaluation according to the model reasoning indexes of each model test strategy in different reasoning test stages to obtain a calculation speed-up ratio curve of each model test strategy in different reasoning test stages; and performing test efficiency evaluation according to the calculation speed-up ratio curve of each model test strategy in different reasoning test stages to obtain an optimal test strategy. According to the invention, the test efficiency of the model in the distributed environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a model distributed inference test system based on multiple machines and multiple cards. Background Art

[0002] In the field of model distributed inference testing, the currently commonly used testing method is based on a single testing strategy of a single machine with multiple cards. This method uses multiple graphics cards equipped on a single machine to perform model inference testing with a preset testing strategy, which is significantly more advanced than using a single card for testing.

[0003] However, with the continuous expansion of the scale of deep learning models and the continuous increase in task complexity, on the one hand, the computing resources of a single machine with multiple cards are ultimately limited. Facing ultra-large-scale and high-complexity models, the computing power of a single machine with multiple cards is difficult to complete a large number of testing tasks in a short time, which greatly affects the testing efficiency of the model. On the other hand, due to different models having different task complexities, some focus on computationally intensive tasks, while some have higher requirements for memory reading and writing. A single testing strategy cannot determine the testing situation of the model under different testing strategies. Therefore, a single testing strategy is difficult to optimize for these differences, resulting in the inability to analyze and evaluate the optimal testing strategy of the model in a distributed environment, and thus unable to provide a basis for subsequent inference optimization of the model in a distributed environment, affecting the testing efficiency of the model. Summary of the Invention

[0004] The present invention provides a model distributed inference test system based on multiple machines and multiple cards, aiming to improve the testing efficiency of the model in a distributed environment.

[0005] In a first aspect, the present invention provides a model distributed inference test system based on multiple machines and multiple cards, including a model inference test middle platform, a server matching module, a test environment construction module, an inference test module, a metric evaluation module, and a test efficiency evaluation module; the model inference test middle platform is respectively connected to the server matching module, the test environment construction module, the inference test module, the metric evaluation module, and the test efficiency evaluation module to manage each module;

[0006] The server matching module is used to obtain a target multi-card server according to the task complexity and the number of tasks of the model to be tested, in combination with the server computing power of each multi-card server in the multi-card server library;

[0007] The test environment construction module is used to construct a multi-machine and multi-card environment with the target multi-card server as a computing node according to the inference test target of the model to be tested and the network communication parameters of each target multi-card server;

[0008] An inference test module, which is used to start the distributed inference test of the model to be tested in the multi-machine and multi-GPU environment according to different model test strategies based on the test data set corresponding to the model to be tested, and obtain the model inference metrics of each model test strategy in different inference test stages;

[0009] An index evaluation module, which is used to evaluate according to the model inference metrics of each model test strategy in different inference test stages, and obtain the calculation speedup curve of each model test strategy in different inference test stages;

[0010] A test efficiency evaluation module, which is used to evaluate the test efficiency according to the calculation speedup curve of each model test strategy in different inference test stages, and obtain the optimal test strategy of the model to be tested.

[0011] In a second aspect, the present invention also provides a model distributed inference test method based on multi-machine and multi-GPU, which is implemented based on the model distributed inference test system based on multi-machine and multi-GPU described in the first aspect. The model distributed inference test method based on multi-machine and multi-GPU includes:

[0012] According to the task complexity and the number of tasks of the model to be tested, and combining the server computing capabilities of each multi-GPU server in the multi-GPU server library, obtain the target multi-GPU server;

[0013] Taking the target multi-GPU server as a computing node, construct a multi-machine and multi-GPU environment according to the inference test target of the model to be tested and the network communication parameters of each target multi-GPU server;

[0014] Start the distributed inference test of the model to be tested in the multi-machine and multi-GPU environment according to different model test strategies based on the test data set corresponding to the model to be tested, and obtain the model inference metrics of each model test strategy in different inference test stages;

[0015] Evaluate according to the model inference metrics of each model test strategy in different inference test stages, and obtain the calculation speedup curve of each model test strategy in different inference test stages;

[0016] Evaluate the test efficiency according to the calculation speedup curve of each model test strategy in different inference test stages, and obtain the optimal test strategy of the model to be tested.

[0017] In a third aspect, the present invention also provides an electronic device, including: a memory, which is used to store a computer software program; a processor, which is used to read and execute the computer software program, and thus implement the model distributed inference test method based on multi-machine and multi-GPU as described above.

[0018] Fourthly, the present invention further provides a non-transitory computer-readable storage medium, in which a computer software program is stored, and when the computer software program is executed by a processor, the method for distributed inference testing of a model based on multiple machines and multiple cards as described above is implemented.

[0019] Fifthly, the present invention further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method for distributed inference testing of a model based on multiple machines and multiple cards as described above is implemented.

[0020] The distributed inference testing system of a model based on multiple machines and multiple cards provided by the embodiments of the present invention matches a target multi-card server that meets the task complexity and the number of tasks according to the server computing power of the multi-card server, and then constructs a multi-machine multi-card environment according to the target multi-card server, so that the computing resources of the multi-machine multi-card environment can always meet the resources required by the model during the distributed inference testing process. Therefore, a large number of test tasks can be completed in a short time. On the other hand, during the inference process of the model in the distributed environment, the test efficiency is evaluated according to the computing acceleration ratio curve of each model test strategy in different inference test stages. Therefore, the optimal test strategy of the model in the distributed environment can be analyzed and evaluated, providing a basis for the subsequent inference optimization of the model in the distributed environment. Therefore, the embodiments of the present invention improve the test efficiency of the model. Description of the Drawings

[0021] Figure 1 is a structural diagram of the distributed inference testing system of a model based on multiple machines and multiple cards provided by the present invention;

[0022] Figure 2 is a schematic diagram of the computing acceleration ratio curve provided by the present invention;

[0023] Figure 3 is a flowchart of the method for distributed inference testing of a model based on multiple machines and multiple cards provided by the present invention;

[0024] Figure 4 is an embodiment diagram of the electronic device provided by the embodiments of the present invention;

[0025] Figure 5 is an embodiment diagram of the computer-readable storage medium provided by the embodiments of the present invention. Detailed Embodiments

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.

[0027] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.

[0028] In the description of the present invention, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present invention is not necessarily construed as being more preferred or advantageous than other embodiments. The following description is given to enable any person skilled in the art to implement and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without the use of these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0029] Optionally, refer to Figure 1 as shown in Figure 1 is a structural diagram of a model distributed inference test system based on multiple machines and multiple cards provided by the present invention. The model distributed inference test system based on multiple machines and multiple cards includes a model inference test middle platform, a server matching module, a test environment construction module, an inference test module, a metric evaluation module, and a test efficiency evaluation module. Among them, the model inference test middle platform in the embodiments of the present invention is respectively connected to the server matching module, the test environment construction module, the inference test module, the metric evaluation module, and the test efficiency evaluation module to manage each module.

[0030] Optionally, the server matching module needs to determine the task complexity and the number of tasks of the model to be tested. The task complexity can be measured by the number of parameters, the number of layers, the type of calculation (such as the scale of matrix multiplication, etc.) of the model. The number of tasks refers to the number of times the model needs to execute inference test tasks in a specific test scenario. The multi-card server library records the computing power of each multi-card server, which is usually determined by factors such as the number of servers, the GPU model, and the CPU performance, and can be quantified as the floating-point operations per second (FLOPS) metric. Further, the server matching module matches the task complexity and the number of tasks of the model to be tested with the computing power of each multi-card server, and selects the target multi-card server that can complete the tasks. For example, if the task complexity is high and the number of tasks is large, a multi-card server with strong computing power is selected.

[0031] In one embodiment, the number of parameters of the model to be tested reaches 100 million, and the number of layers is 100, belonging to a high-complexity model. This test requires 1000 inference test tasks to be executed. In the multi-GPU server library, there are server A and server B. Server A is equipped with 4 domestically controllable and indigenous TF100 GPUs, with a computing power of 1000 TFLOPS; server B is equipped with 8 domestically controllable and indigenous TF100 GPUs, with a computing power of 800 TFLOPS. The server matching module calculates that, ideally, 50 TFLOPS of computing power is required to complete one inference test task of this model. Then, a total of 50000 TFLOPS of computing power is required to complete 1000 tasks. It theoretically takes 50 seconds for server A to complete the task, and 62.5 seconds for server B. Therefore, the server matching module selects server A as the target multi-GPU server.

[0032] Optionally, when constructing a multi-machine and multi-GPU environment, the inference test objectives of the model to be tested need to be considered, such as pursuing high precision or high speed. At the same time, the network communication parameters of each target multi-GPU server need to be combined. Among them, the network communication parameters include bandwidth parameters, latency parameters, bandwidth utilization, throughput, fault tolerance ability parameters, and load balancing parameters. The bandwidth parameters determine the speed of data transmission inside and between servers, such as the amount of data transmitted per second (GB / s); the latency parameter represents the time delay of data transmission (ms); the bandwidth utilization refers to the proportion of the actually used bandwidth to the total bandwidth; the throughput is the amount of data successfully transmitted per unit time; the fault tolerance ability parameter reflects the ability of the system to maintain operation when a network failure occurs; the load balancing parameter is used to ensure uniform load distribution among GPUs within a multi-GPU server and among servers in a multi-machine environment. Therefore, the test environment construction module takes the target multi-GPU server as a computing node and constructs a multi-machine and multi-GPU environment according to the inference test objectives of the model to be tested and the network communication parameters of each target multi-GPU server, as specifically described in steps 201 to 204.

[0033] Optionally, the inference test module launches distributed inference tests in the constructed multi-machine and multi-GPU environment by using the test data set corresponding to the model to be tested according to different model test strategies. The model test strategies include data parallelism strategy and model parallelism strategy. In the data parallelism strategy, the test data set is divided into multiple subsets and calculated simultaneously on different computing nodes. Each computing node calculates different data subsets of the same model, and then the calculation results are summarized and synchronized. In the model parallelism strategy, different parts of the model (such as different layers) are assigned to different computing nodes for calculation, and the nodes cooperate with each other to complete the inference test of the entire model. During the inference test process, the inference test module monitors in real time the model inference metrics of each model test strategy at different inference test stages, such as GPU utilization (obtained by calculating the proportion of the actual GPU usage time in the total running time) and memory occupancy (obtained by monitoring the GPU memory usage), as specifically described in Steps 301 to 309.

[0034] Optionally, the calculation formula for the computing speedup (Speedup) is: Speedup = X1 / Xn, where X1 is the model inference metric of each model test strategy at different inference test stages in a single-GPU or single-machine environment, and Tn is the model inference metric of each model test strategy at different inference test stages in a multi-machine and multi-GPU environment. Therefore, the metric evaluation module calculates the speedup of each model test strategy at different inference test stages according to the model inference metrics of each model test strategy at different inference test stages in a single-GPU or single-machine environment and the model inference metrics of each model test strategy at different inference test stages in a multi-machine and multi-GPU environment, and combines the above formula. The speedups are connected in the order of the inference test stages to obtain the computing speedup curve.

[0035] In one embodiment, in a single - card environment, it takes 10 hours to complete the inference test task of an image classification model for 10,000 pictures. In a multi - machine multi - card environment adopting the data parallel strategy, the inference test process is divided into 5 stages, and the duration of each stage is 1 hour. In the first stage, by monitoring the inference metrics, the time T1 required to complete the task is calculated to be 1.5 hours. Then the speedup ratio of this stage is 10 hours / 1.5 hours ≈ 6.67. In the second stage, the task completion time T2 is 1.3 hours, and the speedup ratio is 10 hours / 1.3 hours ≈ 7.69. And so on, the speedup ratio of each stage is calculated, and then these speedup ratios (6.67, 7.69, …) are connected in the order of the stages to obtain the calculation speedup ratio curve under the data parallel strategy. For the model parallel strategy, similarly, for example, the task completion time in the single - card environment is 10 hours, and the inference test in the multi - machine multi - card environment is divided into 4 stages, with the duration of each stage being 1.2 hours. The task completion time T1 in the first stage is 1.8 hours, and the speedup ratio is 10 hours / 1.8 hours ≈ 5.56. The task completion time T2 in the second stage is 1.6 hours, and the speedup ratio is 10 hours / 1.6 hours ≈ 6.25, etc., to obtain the calculation speedup ratio curve under the model parallel strategy.

[0036] Optionally, the test efficiency evaluation module analyzes the trends of the calculation speedup ratio curves of each model test strategy in different inference test stages. For example, if in the early stage of the inference test, the speedup ratio of the data parallel strategy grows faster, and in the later stage of the inference test, the speedup ratio of the model parallel strategy is higher, then considering comprehensively, the optimal test strategy may be to adopt the data parallel strategy in the early stage and the model parallel strategy in the later stage. Therefore, it can be understood that the optimal test strategy is a combination of the optimal parallel strategies for each stage.

[0037] Continuing with the above - mentioned image classification model as an example, the calculation speedup ratio curve of the data parallel strategy has speedup ratios of 6.67, 7.69, 8.0 in the first 3 stages, and the speedup ratio gradually drops to 7.5 and 7.0 starting from the 4th stage. The calculation speedup ratio curve of the model parallel strategy has speedup ratios of 5.56 and 6.25 in the first 2 stages, and the speedup ratio significantly rises to 8.5 and 9.0 starting from the 3rd stage. The test efficiency evaluation module calculates that the average speedup ratio of the data parallel strategy in the entire 5 - stage process is (6.67 + 7.69 + 8.0 + 7.5 + 7.0) / 5 ≈ 7.37. The average speedup ratio of the model parallel strategy in the entire 4 - stage process is (5.56 + 6.25 + 8.5 + 9.0) / 4 ≈ 7.33. However, from the curve trend, the speedup ratio of the data parallel strategy grows fast in the early stage, and the speedup ratio of the model parallel strategy is high in the later stage. Therefore, the test efficiency evaluation module evaluates that the optimal test strategy for this image classification model is: adopt the data parallel strategy in the first and second stages, and adopt the model parallel strategy in the third and fourth stages.

[0038] In an embodiment of the present invention, a target multi - card server that matches the task complexity and the number of tasks is selected according to the server computing power of the multi - card server, and then a multi - machine multi - card environment is constructed based on the target multi - card server, so that the computing resources of the multi - machine multi - card environment can always meet the resources required by the model during the distributed inference test, and a large number of test tasks can be completed in a short time. During the inference process of the model in the distributed environment, the test efficiency is evaluated according to the computing acceleration ratio curve of each model test strategy in different inference test stages. Therefore, the optimal test strategy of the model in the distributed environment can be analyzed and evaluated, providing a basis for the subsequent inference optimization of the model in the distributed environment. Thus, the test efficiency of the model is improved.

[0039] In one embodiment, the descriptions of steps 201 to 204 are as follows:

[0040] Step 201: Based on the inference test target and the bandwidth parameters of the target multi - card server, each target multi - card server is divided into computing nodes with different priority levels.

[0041] Optionally, the model test system divides the priority levels of the computing nodes according to the inference test target of the model to be tested and the bandwidth parameters of the target multi - card server. If the inference test target focuses on quickly processing large - scale data, then the computing nodes with higher bandwidth should be given higher priorities. Therefore, the model test system divides the computing nodes into primary priority nodes and secondary priority nodes by comparing the bandwidth parameters of each GPU or sub - server in the target multi - card server. The primary priority nodes undertake key computing tasks, and their data transmission speed is fast, which can ensure the efficient progress of the model inference test; the secondary priority nodes assist the primary priority nodes and process some relatively less urgent tasks or tasks with slightly lower bandwidth requirements.

[0042] In one embodiment, the target multi - card server has 8 GPUs, and their bandwidth parameters are GPU1 - 100GB / s, GPU2 - 80GB / s, GPU3 - 100GB / s, GPU4 - 60GB / s, GPU5 - 90GB / s, GPU6 - 70GB / s, GPU7 - 85GB / s, and GPU8 - 50GB / s. The inference test target of the model to be tested is to quickly process a large amount of image data. The model test system sets the computing nodes with a bandwidth greater than or equal to 90GB / s as high - priority computing nodes, that is, GPU1, GPU3, and GPU5 are primary priority nodes, and the rest are secondary priority nodes. The primary priority nodes will be responsible for processing the core computing tasks of the image data, such as convolution operations, etc., while the secondary priority nodes can be used to process some auxiliary tasks, such as data transfer after the pre - processing of the image data.

[0043] Step 202: For the first computing nodes at the same level, establish a first node path based on the latency parameters between the nodes, and optimize the first node path based on the bandwidth utilization rate between the nodes to obtain the main node path.

[0044] Furthermore, for the first computing nodes within the same priority level, the model testing system establishes an initial first node path based on the latency parameters between the nodes. The latency parameter reflects the time required for data to be transmitted between nodes. Therefore, paths with low latency are preferably selected to connect the nodes. Subsequently, the model testing system optimizes the first node path according to the bandwidth utilization rate between the nodes. Paths with high bandwidth utilization rate may become congested during certain periods, affecting the data transmission efficiency. Therefore, by adjusting the path selection, the dependence on high-utilization paths is avoided or reduced, thereby obtaining a more efficient main node path.

[0045] Continuing with the above embodiment, among the main priority nodes GPU1, GPU3, and GPU5, for example, the latency between GPU1 and GPU3 is 5 ms and the bandwidth utilization rate is 40%; the latency between GPU1 and GPU5 is 8 ms and the bandwidth utilization rate is 20%; the latency between GPU3 and GPU5 is 6 ms and the bandwidth utilization rate is 30%. The model testing system establishes an initial path of GPU1 - GPU3 - GPU5 based on the latency parameters (because the latency between GPU1 and GPU3 is the lowest, and the latency between GPU3 and GPU5 is relatively low). Then, considering the bandwidth utilization rate, since the bandwidth utilization rate between GPU1 and GPU5 is the lowest, the model testing system optimizes the path to GPU1 - GPU5 - GPU3, reducing the use of high-bandwidth-utilization paths to obtain the main node path.

[0046] Step 203: For the first computing node and the second computing node at different levels, after determining that a direct connection branch needs to be established based on the throughput between the nodes, establish a second node path according to the latency parameters between the first computing node and the second computing node and the number of intermediate computing nodes passed through.

[0047] Furthermore, for the first computing node and the second computing node at different priority levels, the model testing system first determines whether a direct connection branch needs to be established according to the throughput between the nodes. If the throughput is low, a direct connection may be needed to improve the data transmission efficiency. Therefore, after determining that a direct connection branch needs to be established, the model testing system establishes a second node path according to the latency parameters between the first computing node and the second computing node and the number of intermediate computing nodes passed through, aiming to find a path with the lowest latency and a reasonable number of intermediate computing nodes passed through to balance the transmission speed and network complexity.

[0048] Continuing with the above embodiment, for the primary priority node GPU1 and the secondary priority node GPU6, the direct transmission throughput between GPU1 and GPU6 is relatively low. The latency between GPU1 and GPU6 is 15 ms. If passing through the intermediate node GPU4 (the latency between GPU1 and GPU4 is 6 ms, and the latency between GPU4 and GPU6 is 8 ms), it passes through 1 intermediate computing node. The model testing system calculates that the latency of the direct connection path is 15 ms, and the total latency of the path through the intermediate node is 6 ms + 8 ms = 14 ms, and it only passes through 1 intermediate node, and the network complexity is acceptable. Therefore, the second node path established by the model testing system is GPU1 - GPU4 - GPU6.

[0049] Step 204, based on the node main path between the first computing nodes and the second node path between the first computing nodes and the second computing nodes, construct a multi-machine and multi-GPU environment.

[0050] Furthermore, the node main path between the first computing nodes within the same priority layer of the model testing system and the second node path between the first computing nodes and the second computing nodes in different priority layers are used to connect all the computing nodes to construct a multi-machine and multi-GPU environment, specifically as in steps 2041 to 2044.

[0051] The multi-machine and multi-GPU environment constructed in the embodiments of the present invention can achieve efficient data transmission and computing task allocation. In the multi-machine and multi-GPU environment, the priority layers are divided according to the bandwidth parameters, so that the key computing tasks are undertaken by the primary priority nodes with high bandwidth, ensuring the data processing speed. The node paths are optimized based on latency and bandwidth utilization, reducing data transmission latency and congestion, and improving data transmission efficiency. A reasonable connection path between nodes of different priorities is established according to the throughput, optimizing the overall network communication, so that the constructed multi-machine and multi-GPU environment can significantly improve the inference test speed of the model to be tested.

[0052] In one embodiment, the descriptions of steps 2041 to 2044 are as follows:

[0053] Step 2041, for the first computing nodes in the same layer, determine the target computing nodes based on the fault tolerance parameter and load balancing parameter between the nodes.

[0054] Optionally, the model testing system analyzes the fault tolerance parameter and load balancing parameter between the first computing nodes in the same layer to determine two target computing nodes. Among them, the fault tolerance parameter can be quantified by factors such as the redundant hardware configuration of the node itself and the error detection and correction mechanism. For example, a node with redundant designs such as dual power supplies and hot-swappable hard disks will improve its fault tolerance score. The load balancing parameter is measured by indicators such as the current CPU and GPU utilization of the node and the task queue length.

[0055] Furthermore, the model testing system considers the mutual relationships between nodes and the overall environmental requirements, and preferentially selects two nodes with outstanding comprehensive performance in fault tolerance and load balancing as the target computing nodes.

[0056] In one embodiment, among the primary priority nodes GPU1, GPU3, and GPU5, for example, the fault tolerance score of GPU1 is 8 points (out of 10), the current GPU utilization rate is 40%, and the task queue length is 10; the fault tolerance score of GPU3 is 6 points, the current GPU utilization rate is 60%, and the task queue length is 15; the fault tolerance score of GPU5 is 7 points, the current GPU utilization rate is 30%, and the task queue length is 8. The model testing system calculates the comprehensive performance index of each primary priority node according to the above parameters. The specific formula is: S i =[F i / (1 + U i )]*[1 / (1 + O i )]. Wherein, S i represents the comprehensive performance index of the i-th primary priority node, F i represents the fault tolerance score of the i-th primary priority node, U i represents the GPU utilization rate of the i-th primary priority node, and O i represents the task queue length of the i-th primary priority node. Therefore, the comprehensive performance index of GPU1 is S1 = [8 / (1 + 0.4)]*[1 / (1 + 10)] ≈ 0.51, the comprehensive performance index of GPU3 is S3 = [6 / (1 + 0.6)]*[1 / (1 + 15)] ≈ 0.23, and the comprehensive performance index of GPU5 is S5 = [7 / (1 + 0.3)]*[1 / (1 + 8)] ≈ 0.60. By comparison, the model testing system determines GPU5 and GPU1 as the target computing nodes because their comprehensive performance indices are relatively high among these three nodes, and their comprehensive performance in terms of fault tolerance and load balancing is better.

[0057] Step 2042: Determine the second intermediate computing node based on the latency parameters between the target computing node and each first intermediate computing node; the first intermediate computing node is the node in the first computing nodes except the target computing node.

[0058] Furthermore, for the first intermediate computing nodes in the first computing nodes except the target computing node, the model testing system determines the second intermediate computing node based on the latency parameters between the target computing node and each first intermediate computing node. Among them, the latency parameters accurately reflect the time required for data transmission between the target computing node and the first intermediate computing node. Therefore, the model testing system selects the first intermediate computing node with the lowest total latency as the second intermediate computing node.

[0059] Continuing with the above embodiment, the target computing nodes are GPU5 and GPU1, and the first intermediate computing node is GPU3. The latency between GPU5 and GPU3 is 6 ms, and the latency between GPU1 and GPU3 is 5 ms. The total latency is 6 + 5 = 11 ms. The model testing system determines GPU3 as the second intermediate computing node by comparing the total latency. In this way, when constructing the backup path subsequently, the latency of data transmission from GPU5 and GPU1 to GPU3 is relatively low, which can effectively ensure the data transmission efficiency of the backup path.

[0060] Step 2043: Establish a node backup path based on the target computing nodes and the second intermediate computing node.

[0061] Furthermore, the model testing system establishes a node backup path based on the target computing nodes and the second intermediate computing node. Among them, the backup path is an important supplement to the main path. When the main path fails, such as network link interruption, computing node hardware failure, etc., it can ensure that data continues to be transmitted between computing nodes and maintain the normal operation of the multi-machine multi-GPU environment. The model testing system connects the two target computing nodes to the second intermediate computing node respectively by finely configuring network connection parameters, and sets up an intelligent detection mechanism. When the main path is detected to be abnormal, it can automatically switch to the backup path within an extremely short time.

[0062] Continuing with the above embodiment, after the model testing system determines that the target computing nodes are GPU5 and GPU1 and the second intermediate computing node is GPU3, it establishes data transmission links between GPU5 and GPU3 and between GPU1 and GPU3 respectively through the network configuration tool of the server. At the same time, the model testing system deploys a monitoring program to detect the status of the main path (such as the previously established main path of GPU1 - GPU5 - GPU3) every 5 seconds. When a fault occurs in a certain link in the main path (such as the link from GPU1 to GPU5), the system immediately automatically switches the data transmission to the backup path, such as from GPU1 through GPU3 to GPU5, to ensure that the computing task is not affected.

[0063] Step 2044: Construct a multi-machine multi-GPU environment based on the node main path and node backup path between the first computing nodes, and the second node path between the first computing node and the second computing node.

[0064] Furthermore, the model testing system connects all computing nodes based on the node main path between the first computing nodes, the newly established node backup path, and the second node path between the first computing node and the second computing node to construct a complete and efficient multi-machine and multi-GPU environment. Among them, in the constructed multi-machine and multi-GPU environment, data can be intelligently and efficiently transmitted between each computing node according to the settings of the main path, the backup path, and the paths between nodes with different priorities. Each computing node works collaboratively, not only meeting the inference test objectives of the model to be tested, but also having strong fault tolerance and excellent load balancing capabilities, providing a stable and efficient operating environment for model inference testing.

[0065] Continuing with the above embodiment, based on the results of the previous steps, the node main path is, for example, GPU1 - GPU5 - GPU3, the node backup path is GPU1 - GPU3 - GPU5, and the second node path between nodes with different priorities (such as GPU1 - GPU4 - GPU6), etc. The model testing system connects all GPUs in the server according to these paths. By precisely configuring network protocols, optimizing routing rules, etc., data can be intelligently and quickly switched between different paths. For example, under normal circumstances, data is transmitted according to the node main path. When the link between GPU1 and GPU5 in the node main path fails, data can automatically switch to the backup path GPU1 - GPU3 - GPU5 within milliseconds. At the same time, data transmission between nodes with different priorities can also be efficiently carried out according to the second node path, thus constructing a multi-machine and multi-GPU environment.

[0066] The multi-machine and multi-GPU environment constructed in the embodiment of the present invention exhibits excellent performance. In terms of fault tolerance, by reasonably constructing the backup path, when the main path fails, it can quickly and seamlessly switch to the backup path, ensuring that data transmission and computing tasks are uninterrupted, and guaranteeing the stability and reliability of model testing.

[0067] In one embodiment, the model testing strategy includes a data parallel strategy. Among them, the data parallel strategy represents dividing the test data set onto the computing cards of each computing node for backpropagation calculation and cross-node collaborative calculation. Therefore, the descriptions of steps 301 to 304 are as follows:

[0068] Step 301, for each primary priority node in each inference test stage, divide the corresponding number of test data sets onto the computing cards of each primary priority node and the computing cards of the secondary priority nodes connected to each primary priority node based on the number of secondary priority nodes connected by the second node path.

[0069] Optionally, at each primary priority node in each inference test phase, the model test system determines the number of secondary priority nodes connected to each primary priority node through the second node path, and then evenly divides the test data set into corresponding portions according to this number. Among them, the data of these portions are respectively allocated to the computing cards of each primary priority node itself and the computing cards of the secondary priority nodes connected to it. Such an allocation method ensures that each computing card can obtain data of an appropriate scale for calculation, making full use of the computing resources in a multi-machine and multi-card environment. In one embodiment, in a certain inference test phase, there are primary priority nodes GPU1, GPU3, and GPU5. Among them, GPU1 is connected to 2 secondary priority nodes (GPU4 and GPU6) through the second node path, GPU3 is connected to 1 secondary priority node (GPU7), and GPU5 is connected to 3 secondary priority nodes (GPU2, GPU8, and GPU9). The test data set contains 10,000 pieces of data. For GPU1, since it is connected to 2 secondary priority nodes and there are a total of 3 computing nodes including itself, the test data set is divided into 3 portions, with each portion having approximately 3,333 pieces of data, which are respectively allocated to the computing cards of GPU1 and the connected GPU4 and GPU6. For GPU3, the data set is divided into 2 portions (itself and the connected GPU7), with each portion having approximately 5,000 pieces of data for allocation. For GPU5, the data set is divided into 4 portions (itself and the connected GPU2, GPU8, and GPU9), and each portion of approximately 2,500 pieces of data is allocated to the corresponding computing cards.

[0070] Step 302: Based on the test data set in each computing card, start the forward propagation calculation of the model to be tested, and obtain the forward propagation results of each computing card.

[0071] Furthermore, the model test system independently starts each computing card to work in parallel. According to the structure and parameters of the model to be tested, a series of calculation operations, such as convolution and fully connected operations, are performed on the input test data, and finally the forward propagation results corresponding to each computing card are obtained. Among them, since the data processed by each computing card is different, but the calculation process is based on the same model structure and parameters, the above parallel calculation greatly accelerates the process of model inference test.

[0072] Continuing with GPU1 and its connected GPU4 and GPU6 as an example, the forward propagation calculation process on the GPU1 computing card is as follows: The 3333 pieces of data allocated are sequentially input into the model to be tested, and calculations are performed in the order of the model's convolutional layer, activation function layer, fully connected layer, etc. For example, the first layer of the model is a convolutional layer with 32 convolutional kernels, a size of 3*3, and a stride of 1. The GPU1 computing card performs a convolutional operation on the input data. After obtaining the convolutional result, it is processed by an activation function (such as ReLU) and then passed to the next layer for calculation, finally obtaining the forward propagation result of the GPU1 computing card. At the same time, the GPU4 and GPU6 computing cards also perform similar forward propagation calculations on the data allocated to them according to the same model structure and parameters.

[0073] Step 303: Based on the second node path, integrate the first forward propagation results of the computing cards in each primary priority node and the second forward propagation results of the computing cards in the secondary priority nodes connected to each primary priority node to obtain the integrated forward propagation result.

[0074] Furthermore, the model testing system integrates the first forward propagation results of the computing cards in each primary priority node through the second node main path, and also integrates the second forward propagation results of the computing cards in the secondary priority nodes connected to each primary priority node to obtain the integrated forward propagation result.

[0075] Continuing with GPU1 and its connected GPU4 and GPU6 as an example, after the GPU1 computing card completes the forward propagation calculation, it sends the result to a specified integration node (such as the central processing unit of a server) through a network link (based on the previously constructed node path). The GPU4 and GPU6 computing cards also send their respective forward propagation results to this integration node. The integration node merges the first forward propagation result of GPU1 with the second forward propagation results of GPU4 and GPU6 according to a predetermined rule. For example, arrange the corresponding data calculation results together in the order of data input to obtain the complete integrated forward propagation result for GPU1 and its connected secondary priority nodes.

[0076] Step 304: Based on the integrated forward propagation results of each primary priority node in each inference test stage, determine the model inference metrics of the data parallel strategy under different inference test stages.

[0077] Furthermore, the model testing system determines the model inference metrics of the data parallel strategy under different inference test stages according to the integrated forward propagation results of each primary priority node in each inference test stage, as specifically described in Steps 3041 to 3044.

[0078] The embodiments of the present invention achieve efficient model inference testing in a multi - machine and multi - card environment based on a data parallel strategy. In the data partitioning stage, the test data set is reasonably allocated according to the node connection situation, enabling each computing card to be fully utilized and improving the utilization rate of computing resources. The parallel forward - propagation calculation greatly accelerates the model inference process. Compared with single - card computing, the inference testing time is significantly shortened, significantly improving the testing efficiency and testing effect of the model inference testing, laying a solid foundation for the performance evaluation and optimization of the model.

[0079] In one embodiment, the descriptions of steps 3041 to 3044 are as follows:

[0080] Step 3041, for each primary - priority node in each inference testing stage, based on the forward - propagation integration result, start the back - propagation calculation of the model to be tested according to the test data set in each computing card, and obtain the gradient value of each model parameter output by each computing card.

[0081] Optionally, in each primary - priority node in each inference testing stage, the model testing system uses the forward - propagation integration result and the test data set in each computing card to start the back - propagation calculation of the model to be tested. The back - propagation algorithm is based on the chain - rule of derivative calculation. Starting from the output layer of the model, according to the error of the model output with respect to the loss function, the gradient value of each model parameter is calculated backward. Since the forward propagation has been completed on each computing card and each computing card has a corresponding test data set, the back - propagation can be performed in parallel on each computing card. Each computing card independently calculates the gradient value of the model parameters corresponding to the data it processes, which can make full use of the computing power of multiple cards and speed up the calculation.

[0082] In one embodiment, for example, in a certain inference testing stage, the primary - priority node GPU1 and its connected secondary - priority nodes GPU4 and GPU6 complete the forward propagation and obtain the integration result. Taking the GPU1 computing card as an example, the test data set it processes contains 1000 samples. In the back - propagation calculation, first calculate the error between the model output and the true label according to the loss function of the model (for example, the cross - entropy loss function). For a neural network model with n layers, the output of the l - th layer is a l , the weight matrix is W l , the bias vector is b l , and the input is z l . Then, starting from the output layer n, calculate the error

[0083] where L is the loss function, is the gradient of the loss function with respect to the activation value a n of the output layer, ⊙ represents element - wise multiplication, and f(z n ) is the activation function at z nDerivative at. Then, backpropagation calculates the error δ of the previous layer l-1 =(W l ) T δ l ⊙f'(z l-1 ) and then calculates the gradients of the weight matrix W l and the bias vector b l : The GPU1 computing card performs the above backpropagation calculation on the model parameters corresponding to 1000 samples to obtain the gradient values of each model parameter. The GPU4 and GPU6 computing cards also perform similar backpropagation calculations based on their own test datasets and output the gradient values of their respective corresponding model parameters.

[0084] Step 3042, based on the second node path, integrate the gradient values of the computing cards in each primary priority node and the gradient values of the computing cards in the secondary priority nodes connected to each primary priority node to obtain the gradient value integration result.

[0085] Furthermore, the model testing system integrates the gradient values of the computing cards in each primary priority node and the gradient values of the computing cards in the secondary priority nodes connected to each primary priority node according to the second node path. Continuing with the example of the primary priority node GPU1, after the GPU1 computing card completes the backpropagation calculation to obtain the gradient values, it sends the gradient values to the specified integration node (such as the central processing unit of the server) through the second node path (such as the paths from GPU1 to GPU4 and from GPU1 to GPU6 in GPU1 - GPU4 - GPU6). The GPU4 and GPU6 computing cards also send their respective gradient values to this integration node through the corresponding paths. The integration node merges the gradient values of the GPU1, GPU4, and GPU6 computing cards according to a predetermined rule. For example, arrange the gradient values of the corresponding parameters together in the order of the model parameters to obtain the gradient value integration result for GPU1 and its connected secondary priority nodes.

[0086] Step 3043, fuse the gradient value integration results of each primary priority node according to the node main path or the node alternative path to obtain the final gradient value result.

[0087] Furthermore, the model testing system fuses the gradient value integration results of each primary priority node according to the node main path or the node alternative path. When the node main path is normal, the model testing system collects the gradient value integration results of each primary priority node together through the node main path for fusion; if the node main path fails, the model testing system performs data transmission and fusion through the node alternative path. During the fusion process, it is necessary to ensure that the gradient values of different primary priority nodes can be correctly merged.

[0088] In one embodiment, the node main path is GPU1 - GPU5 - GPU3. When the node main path is normal, the integrated gradient value result of GPU1 is sent to GPU5 through the node main path, and then forwarded by GPU5 to GPU3. At GPU3, the integrated gradient value results of GPU1 and GPU5 are fused. For example, for a certain weight parameter W in the model, the gradient of this parameter in the integrated gradient value result of GPU1 is g 1W , and the gradient of this parameter in the integrated gradient value result of GPU5 is g 5W . When fusing, the following formula (such as simple addition) can be used: g finalW = g 1W + g 5W , to obtain the final gradient value g finalW of this parameter. If the link from GPU1 to GPU5 in the node main path fails, the transmission and fusion of the integrated gradient value result are carried out through the node alternate path (such as GPU1 - GPU3 - GPU5).

[0089] Step 3044, update the model parameters in the model to be tested based on the final gradient value results of each inference test stage, and obtain the model inference metrics for each inference test stage.

[0090] Furthermore, the model testing system updates the model parameters in the model to be tested based on the final gradient value results of each inference test stage. Usually, the Stochastic Gradient Descent (SGD) or its variant algorithms are used to adjust the model parameters according to the calculated gradient values, so that the loss of the model on the inference test data gradually decreases. After updating the model parameters, the model testing system determines the model inference metrics for each inference test stage by monitoring the performance of the model on the test data set, such as calculating the accuracy, recall rate, etc. of the model, and combining hardware metrics such as GPU utilization and memory occupancy. The model inference metrics reflect the performance of the model and the resource usage situation in this inference test stage, such as indicators like GPU utilization and memory occupancy.

[0091] In one embodiment, the model uses the Stochastic Gradient Descent algorithm with a learning rate of α. For the weight parameter W in the model, its update formula is: W = W - αg finalW . The model testing system updates all model parameters according to the above formula. After the update is completed, the model testing system uses the test data set to perform inference on the model again and calculates the accuracy of the model. For example, there are 1000 samples in the test data set, and the model correctly classifies 800 samples, then the accuracy is 800 / 1000 = 80%. At the same time, the model testing system monitors that during the process of updating the model parameters, the average utilization rate of the GPU is 70%, and the memory occupancy is stable at 5GB.

[0092] The embodiments of the present invention can efficiently update model parameters and accurately determine model inference metrics. In the stage of backpropagation to calculate gradient values, multi-card parallel computing is utilized, greatly accelerating the gradient calculation speed. The process of integrating and fusing gradient values ensures that gradient information from different computing cards and different primary priority nodes can be accurately summarized, providing a guarantee for correctly updating model parameters. By reasonably updating model parameters, the model can be continuously optimized during the inference test process, improving performance metrics such as accuracy on the test dataset, ensuring the stability and reliability of model testing, and thus improving the test efficiency of the model.

[0093] In one embodiment, the model testing strategy includes a model parallel strategy. Among them, the model parallel strategy represents that different layers are assigned to the computing cards of different computing nodes according to the model structure for collaborative model inference. Therefore, the descriptions of steps 305 to 309 are as follows:

[0094] Step 305, for each inference test stage, allocate a corresponding number of convolutional layers to the computing cards of each primary priority node, and allocate a corresponding number of fully connected layers and pooling layers to the computing cards of the first priority node and the second priority node that are connected to the primary priority node through the second node path.

[0095] Optionally, in each inference test stage, the model testing system reasonably allocates different layers of the model to be tested to each computing card according to the structure of the multi-machine multi-card environment and the characteristics of the model. Specifically, the model testing system allocates a corresponding number of convolutional layers to the computing cards of each primary priority node. Among them, the convolutional layers have a large amount of computation, and the primary priority nodes have strong computing capabilities and can process them efficiently. At the same time, allocate a corresponding number of fully connected layers and pooling layers to the computing cards of the first priority node and the second priority node that are connected to the primary priority node through the second node path.

[0096] In one embodiment, the model to be tested has 20 convolutional layers, 10 fully connected layers, and 5 pooling layers. The primary priority nodes are GPU1, GPU3, and GPU5. GPU1 is connected to the first priority node GPU4 and the second priority node GPU6 through the second node path. The model testing system allocates 8 convolutional layers to the computing card of GPU1, 6 convolutional layers to the computing card of GPU3, and 6 convolutional layers to the computing card of GPU5. 4 fully connected layers are allocated to the computing card of GPU4, and 2 pooling layers are allocated to the computing card of GPU6. Similarly, similar allocations are made to the first priority node and the second priority node connected to GPU3, and the first priority node and the second priority node connected to GPU5.

[0097] Step 306: Start the inference test for each computing card based on the test data set, transmit and integrate the inference test results of the computing cards of each primary priority node based on the node main path or the node alternate path to obtain the convolutional layer inference results, and transmit the convolutional layer inference results to the first priority nodes based on the second node path.

[0098] Furthermore, the model testing system starts the inference test for each computing card based on the test data set. On the computing cards of the primary priority nodes, the inference test is performed for the assigned convolutional layers. Each computing card works in parallel and performs convolutional operations on the input test data. After completing the inference test, the model testing system transmits and integrates the inference test results of the computing cards of each primary priority node based on the node main path or the node alternate path to obtain the convolutional layer inference results. Then, the model testing system transmits the convolutional layer inference results to the corresponding first priority nodes through the second node path.

[0099] Continuing with the computing card of GPU1 as an example, it receives the computing tasks and corresponding test data of 8 convolutional layers. During the inference test, the computing card performs operations on the input data for 8 convolutional layers in sequence. For example, the first convolutional layer has 32 convolutional kernels with a size of 3*3 and a stride of 1. The computing card performs a convolutional operation on the input data to obtain an output result and then passes it to the next convolutional layer for calculation. After completing all convolutional layer calculations, the computing card of GPU1 transmits the inference test results to the designated integration node through the node main path (e.g., the main path is normal) for integration to obtain the complete convolutional layer inference results. Then, the convolutional layer inference results are transmitted to GPU4 through the second node path (GPU1 - GPU4).

[0100] Step 307: Transmit and integrate the inference test results of the computing cards of each first priority node based on the node main path or the node alternate path to obtain the fully connected layer inference results.

[0101] Furthermore, after the computing cards of the first priority nodes receive the convolutional layer inference results from the primary priority nodes, the model testing system starts the inference test for these computing cards for the fully connected layer. The computing cards of each first priority node work in parallel and perform fully connected operations on the input convolutional layer inference results. After completing the inference test, the model testing system transmits and integrates the inference test results of the computing cards of each first priority node based on the node main path or the node alternate path to obtain the fully connected layer inference results.

[0102] Continuing with the above embodiment, after the GPU 4 receives the convolutional layer inference result from the GPU 1, its computing card performs an inference test on the allocated 4 fully connected layers. For example, the first fully connected layer has 100 neurons. The computing card performs a matrix multiplication operation on the convolutional layer inference result and the weight matrix of the fully connected layer to obtain an output result. After completing all the calculations of the fully connected layers, the computing card of the GPU 4 transmits the inference test result to the designated integration node through the node backup path (when the main path fails) for integration. Together with the calculation results of other first-priority nodes (such as GPU 7, etc.), a complete inference result of the fully connected layer is obtained.

[0103] Step 308: Transmit the inference result of the fully connected layer to the second-priority node based on the second node path, and transmit and integrate the inference test results of the computing cards of each second-priority node based on the node main path or the node backup path to obtain the inference result of the pooling layer, and complete the inference test of each inference test stage.

[0104] Furthermore, the model testing system transmits the inference result of the fully connected layer to the second-priority node through the second node path. After the computing card of the second-priority node receives the inference result of the fully connected layer, it starts the inference test for the pooling layer. The computing cards of each second-priority node work in parallel to perform a pooling operation on the input inference result of the fully connected layer. After completing the inference test, the model testing system transmits and integrates the inference test results of the computing cards of each second-priority node based on the node main path or the node backup path to obtain the inference result of the pooling layer, and complete the inference test of each inference test stage.

[0105] Continuing with the above embodiment, after the GPU 6 receives the inference result of the fully connected layer from the GPU 1 (indirectly transmitted through the GPU 4), its computing card performs an inference test on the allocated 1 pooling layer. For example, this pooling layer is a max pooling layer with a window size of 2*2 and a stride of 2. The computing card performs a max pooling operation on the inference result of the fully connected layer to obtain an output result. After completing the calculation, the computing card of the GPU 6 transmits the inference test result to the designated integration node through the node main path (or the node backup path) for integration. Together with the calculation results of other second-priority nodes (such as GPU 8, etc.), a complete inference result of the pooling layer is obtained, and the inference test of this inference test stage is completed.

[0106] Step 309: Determine the model inference metrics of the model parallel strategy under different inference test stages based on the inference test metrics during the inference test of the corresponding layer in each inference test stage.

[0107] Furthermore, based on the inference test metrics during the inference test process for the corresponding layer in each inference test phase, the model test system determines the model inference metrics of the model parallel strategy under different inference test phases. The inference test metrics may include computing time, memory occupancy, computing resource utilization rate, etc., and the metric data can be obtained by monitoring the running conditions of each computing card. For example, by recording the start and end times of each computing card executing the inference test for each layer, the computing time is calculated; by monitoring the memory usage of the computing card, the memory occupancy data is obtained; by analyzing the resource scheduling situation of the computing card, the computing resource utilization rate is evaluated.

[0108] In one embodiment, in a certain inference test phase, through monitoring, it is found that the time for the GPU1 computing card to execute the convolutional layer inference test is 30 seconds, the memory occupancy is 3GB, and the computing resource utilization rate is 80%; the time for the GPU4 computing card to execute the fully connected layer inference test is 20 seconds, the memory occupancy is 2GB, and the computing resource utilization rate is 75%; the time for the GPU6 computing card to execute the pooling layer inference test is 10 seconds, the memory occupancy is 1GB, and the computing resource utilization rate is 70%. The model test system synthesizes these data to obtain the model inference metrics under the model parallel strategy in this inference test phase. For example, the average computing time is (30 + 20 + 10) / 3 = 20 seconds, the average memory occupancy is (3 + 2 + 1) / 3 = 2GB, and the average computing resource utilization rate is (80% + 75% + 70%) / 3 ≈ 75%, etc., so as to evaluate the performance of the model in this phase.

[0109] The embodiment of the present invention realizes efficient distributed inference testing based on the model parallel strategy. In the model layer allocation phase, the computing capabilities of different priority nodes are reasonably utilized, enabling parallel computing of each layer of the model and improving the computing efficiency. During the integration of the inference test and result transmission, through the constructed node path, the efficient transmission and accurate integration of data between different computing cards and nodes are ensured, guaranteeing the smooth progress of the inference test. By monitoring the inference test metrics to determine the model inference metrics, the performance of the model in different inference test phases can be comprehensively understood, providing a basis for optimizing the model inference test process. Therefore, the test efficiency of the model inference test is improved, a high-quality model is inferred, and the effect of the model in actual applications is enhanced.

[0110] In one embodiment, the descriptions of steps 501 to 505 are as follows:

[0111] Step 501, for each inference test phase, obtain the curve intersection point between the first computing speedup curve of the data parallel strategy and the second computing speedup curve of the model parallel strategy.

[0112] Optionally, for each inference test stage, the intersection coordinates of the function expressions of the first computational speedup ratio curve of the data parallel strategy and the second computational speedup ratio curve of the model parallel strategy analyzed by the model test system are the curve intersection points. The curve intersection points indicate that the computational speedup ratios of the two strategies are the same in this inference test stage. Refer to Figure 2 , for the first computational speedup ratio curve X and the second computational speedup ratio curve Y, the curve intersection point of the first computational speedup ratio curve X and the second computational speedup ratio curve Y is p.

[0113] Step 502: Compare the computational speedup ratios in the first computational speedup ratio curve and the second computational speedup ratio curve based on the curve intersection points to obtain a comparison result.

[0114] Furthermore, the model test system compares the computational speedup ratios in the first computational speedup ratio curve and the second computational speedup ratio curve according to the curve intersection points to obtain a comparison result. Among them, the comparison result can be that the computational speedup ratio of the first computational speedup ratio curve is greater than that of the second computational speedup ratio curve, or the computational speedup ratio of the first computational speedup ratio curve is less than that of the second computational speedup ratio curve. The larger the computational speedup ratio, the better the strategy effect of the inference test model.

[0115] Step 503: If the comparison result is that the computational speedup ratio of the first computational speedup ratio curve is greater than that of the second computational speedup ratio curve, determine the data parallel strategy as the target parallel strategy.

[0116] Furthermore, if the comparison result is that the computational speedup ratio of the first computational speedup ratio curve is greater than that of the second computational speedup ratio curve, determine the data parallel strategy as the target parallel strategy.

[0117] Step 504: If the comparison result is that the computational speedup ratio of the first computational speedup ratio curve is less than that of the second computational speedup ratio curve, determine the model parallel strategy as the target parallel strategy.

[0118] Furthermore, if the comparison result is that the computational speedup ratio of the first computational speedup ratio curve is less than that of the second computational speedup ratio curve, determine the model parallel strategy as the target parallel strategy.

[0119] Step 505: Combine the strategies based on the target parallel strategy of each inference test stage to obtain the optimal test strategy for the model to be tested.

[0120] Furthermore, the model test system combines the target parallel strategies of each inference test stage to obtain the optimal test strategy for the model to be tested.

[0121] Continue to refer to Figure 2, including three inference test phases. In the first inference test phase, the computing speedup ratio of the second computing speedup ratio curve Y is greater than that of the first computing speedup ratio curve X. Therefore, the target parallel strategy in the first inference test phase is the model parallel strategy. In the second and third inference test phases, the computing speedup ratio of the second computing speedup ratio curve Y is less than that of the first computing speedup ratio curve X. Therefore, the target parallel strategies in the second and third inference test phases are the data parallel strategies. Therefore, the optimal test strategy for the model to be tested is: adopt the model parallel strategy in the first inference test phase, adopt the data parallel strategy in the second inference test phase, and adopt the data parallel strategy in the third inference test phase.

[0122] The embodiment of the present invention can accurately determine the optimal test strategy for the model to be tested according to the computing speedup ratio curves of the data parallel strategy and the model parallel strategy in different inference test phases, so that the model can make full use of the advantages of the data parallel strategy and the model parallel strategy throughout the test process. Therefore, by dynamically selecting the strategy, compared with fixedly adopting a certain parallel strategy, it can significantly improve the test efficiency of the model, reduce the inference test time, reduce the waste of computing resources, and thus obtain a model with better performance faster and improve the performance of the model in actual applications.

[0123] Optionally, referring to Figure 3 , Figure 3 is the flowchart of the model distributed inference test method based on multi-machine and multi-GPU provided by the present invention. The execution subject of the embodiment of the present invention is the model test system. Therefore, the model distributed inference test method based on multi-machine and multi-GPU includes:

[0124] Step 10, according to the task complexity and the number of tasks of the model to be tested, and combining the server computing capabilities of each multi-GPU server in the multi-GPU server library, obtain the target multi-GPU server.

[0125] Optionally, the model test system first needs to determine the task complexity and the number of tasks of the model to be tested. The task complexity can be measured by the number of parameters, the number of layers, the computing type (such as the scale of matrix multiplication, etc.) of the model. The number of tasks refers to the number of times the model needs to execute the inference test task in a specific test scenario. The server computing capabilities of each multi-GPU server are recorded in the multi-GPU server library, which are usually determined by factors such as the number of servers, the GPU model, and the CPU performance, and can be quantified as the floating-point operation per second (FLOPS) metric. Further, the model test system matches the task complexity and the number of tasks of the model to be tested with the computing capabilities of each multi-GPU server, and selects the target multi-GPU server that can complete the task. For example, if the task complexity is high and the number of tasks is large, select a multi-GPU server with strong computing capabilities.

[0126] Step 20: Using the target multi-GPU server as a computing node, construct a multi-machine multi-GPU environment according to the inference test objectives of the model to be tested and the network communication parameters of each target multi-GPU server.

[0127] Furthermore, when constructing the multi-machine multi-GPU environment, it is necessary to consider the inference test objectives of the model to be tested, such as pursuing high precision or high speed. At the same time, it is necessary to combine the network communication parameters of each target multi-GPU server. Among them, the network communication parameters include bandwidth parameters, latency parameters, bandwidth utilization rate, throughput, fault tolerance ability parameters, and load balancing parameters. The bandwidth parameters determine the speed of data transmission inside and between servers, such as the amount of data transmitted per second (GB / s); the latency parameters represent the time delay of data transmission (ms); the bandwidth utilization rate refers to the proportion of the actual used bandwidth to the total bandwidth; the throughput is the amount of data successfully transmitted per unit time; the fault tolerance ability parameters reflect the ability of the system to maintain operation when the network fails; the load balancing parameters are used to ensure uniform distribution of the load among GPUs within the multi-GPU server and among servers in the multi-machine environment.

[0128] Therefore, the model test system uses the target multi-GPU server as a computing node and constructs a multi-machine multi-GPU environment according to the inference test objectives of the model to be tested and the network communication parameters of each target multi-GPU server.

[0129] Step 30: According to different model test strategies, start the distributed inference test of the model to be tested on the test dataset corresponding to the model to be tested in the multi-machine multi-GPU environment, and obtain the model inference metrics of each model test strategy at different inference test stages.

[0130] Furthermore, the model test system starts the distributed inference test in the constructed multi-machine multi-GPU environment according to different model test strategies and using the test dataset corresponding to the model to be tested. Among them, the model test strategies include data parallelism strategy and model parallelism strategy. In the data parallelism strategy, the test dataset is divided into multiple subsets and calculated simultaneously on different computing nodes (GPUs or servers). Each computing node calculates different data subsets of the same model, and then the calculation results are summarized and synchronized. In the model parallelism strategy, different parts of the model (such as different layers) are assigned to different computing nodes for calculation, and the nodes work together to complete the inference test of the entire model. During the inference test process, the model test system monitors the model inference metrics of each model test strategy at different inference test stages in real time, such as GPU utilization rate (obtained by calculating the proportion of the actual used time of the GPU to the total running time) and memory occupancy (obtained by monitoring the usage of GPU memory).

[0131] Step 40: Evaluate according to the model inference metrics of each model testing strategy at different inference testing stages to obtain the computational speedup curves of each model testing strategy at different inference testing stages.

[0132] Further, the calculation formula for the speedup is: Speedup = X1 / Xn, where X1 is the model inference metric of each model testing strategy at different inference testing stages in a single - card or single - machine environment, and Tn is the model inference metric of each model testing strategy at different inference testing stages in a multi - machine and multi - card environment. Therefore, based on the model inference metrics of each model testing strategy at different inference testing stages in a single - card or single - machine environment and the model inference metrics of each model testing strategy at different inference testing stages in a multi - machine and multi - card environment, the model testing system calculates the speedup of each model testing strategy at different inference testing stages according to the above formula, and connects the speedups in the order of the inference testing stages to obtain the computational speedup curve.

[0133] Step 50: Evaluate the test efficiency according to the computational speedup curves of each model testing strategy at different inference testing stages to obtain the optimal testing strategy for the model to be tested.

[0134] Further, the model testing system analyzes the trends of the computational speedup curves of each model testing strategy at different inference testing stages. For example, if in the early stage of the inference test, the speedup of the data parallel strategy grows rapidly, while in the later stage of the inference test, the speedup of the model parallel strategy is higher, then considering comprehensively, an optimal testing strategy of adopting the data parallel strategy in the early stage and the model parallel strategy in the later stage may be obtained. Therefore, it can be understood that the optimal testing strategy is a combination of the optimal parallel strategies for each stage.

[0135] In the embodiment of the present invention, a target multi - card server that matches the server computing power of the multi - card server to the task complexity and the number of tasks is selected, and then a multi - machine and multi - card environment is constructed according to the target multi - card server, so that the computing resources of the multi - machine and multi - card environment can always meet the resources required by the model during the distributed inference test, and a large number of test tasks can be completed in a short time. During the process of the model's inference in the distributed environment, the test efficiency is evaluated according to the computational speedup curves of each model testing strategy at different inference testing stages. Therefore, the optimal testing strategy of the model in the distributed environment can be analyzed and evaluated, providing a basis for the subsequent inference optimization of the model in the distributed environment. Therefore, the test efficiency of the model is improved.

[0136] Please refer to Figure 4 , Figure 4 which is the embodiment diagram of the electronic device provided by the embodiment of the present invention. As Figure 4As shown in the figure, an embodiment of the present invention provides an electronic device 400, including a memory 410, a processor 420, and a computer program 411 stored on the memory 410 and executable on the processor 420. When the processor 420 executes the computer program 411, the following steps are implemented:

[0137] According to the task complexity and the number of tasks of the model to be tested, and in combination with the server computing capabilities of each multi-card server in the multi-card server library, obtain the target multi-card server;

[0138] Using the target multi-card server as a computing node, construct a multi-machine multi-card environment according to the inference test target of the model to be tested and the network communication parameters of each target multi-card server;

[0139] According to different model test strategies and the test data set corresponding to the model to be tested, start the distributed inference test of the model to be tested in the multi-machine multi-card environment, and obtain the model inference indicators of each model test strategy in different inference test stages;

[0140] Evaluate according to the model inference indicators of each model test strategy in different inference test stages, and obtain the calculation speedup ratio curves of each model test strategy in different inference test stages;

[0141] Evaluate the test efficiency according to the calculation speedup ratio curves of each model test strategy in different inference test stages, and obtain the optimal test strategy of the model to be tested.

[0142] Please refer to Figure 5 , Figure 5 which is the embodiment diagram of the computer-readable storage medium provided by the embodiment of the present invention. As Figure 5 shown, this embodiment provides a computer-readable storage medium 500, on which a computer program 411 is stored. When the computer program 411 is executed by a processor, the following steps are implemented:

[0143] According to the task complexity and the number of tasks of the model to be tested, and in combination with the server computing capabilities of each multi-card server in the multi-card server library, obtain the target multi-card server;

[0144] Using the target multi-card server as a computing node, construct a multi-machine multi-card environment according to the inference test target of the model to be tested and the network communication parameters of each target multi-card server;

[0145] According to different model test strategies and the test data set corresponding to the model to be tested, start the distributed inference test of the model to be tested in the multi-machine multi-card environment, and obtain the model inference indicators of each model test strategy in different inference test stages;

[0146] Evaluate according to the model inference metrics of each model testing strategy at different inference testing stages, and obtain the computational speedup curves of each model testing strategy at different inference testing stages;

[0147] Conduct a test efficiency evaluation based on the computational speedup curves of each model testing strategy at different inference testing stages, and obtain the optimal testing strategy for the model to be tested.

[0148] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the distributed inference testing method of the model based on multiple machines and multiple cards provided by the above-mentioned various methods. The distributed inference testing method of the model based on multiple machines and multiple cards includes:

[0149] According to the task complexity and the number of tasks of the model to be tested, and in combination with the server computing capabilities of each multi-card server in the multi-card server library, obtain the target multi-card server;

[0150] Using the target multi-card server as a computing node, construct a multi-machine multi-card environment according to the inference testing target of the model to be tested and the network communication parameters of each target multi-card server;

[0151] According to different model testing strategies, start the distributed inference testing of the model to be tested on the test data set corresponding to the model to be tested in the multi-machine multi-card environment, and obtain the model inference metrics of each model testing strategy at different inference testing stages;

[0152] Evaluate according to the model inference metrics of each model testing strategy at different inference testing stages, and obtain the computational speedup curves of each model testing strategy at different inference testing stages;

[0153] Conduct a test efficiency evaluation based on the computational speedup curves of each model testing strategy at different inference testing stages, and obtain the optimal testing strategy for the model to be tested.

[0154] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0155] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A model distributed inference test system based on multiple machines and multiple cards, characterized in that, It includes a model inference test platform, a server matching module, a test environment construction module, an inference test module, a metric evaluation module, and a test efficiency evaluation module; the model inference test platform is respectively connected to the server matching module, the test environment construction module, the inference test module, the metric evaluation module, and the test efficiency evaluation module to manage each module; The server matching module is used to obtain a target multi-card server according to the task complexity and the number of tasks of the model to be tested, in combination with the server computing power of each multi-card server in the multi-card server library; The test environment construction module is used to construct a multi-machine multi-card environment with the target multi-card server as a computing node according to the inference test target of the model to be tested and the network communication parameters of each target multi-card server; The inference test module is used to start the distributed inference test of the model to be tested in the multi-machine multi-card environment according to different model test strategies based on the test data set corresponding to the model to be tested, and obtain the model inference metrics of each model test strategy in different inference test stages; The metric evaluation module is used to evaluate according to the model inference metrics of each model test strategy in different inference test stages, and obtain the calculation speedup curve of each model test strategy in different inference test stages; The test efficiency evaluation module is used to evaluate the test efficiency according to the calculation speedup curve of each model test strategy in different inference test stages, and obtain the optimal test strategy of the model to be tested.

2. The model distributed inference test system based on multi-machine and multi-card according to claim 1, wherein The network communication parameters include bandwidth parameters, latency parameters, bandwidth utilization rate, and throughput; The constructing a multi-machine multi-card environment with the target multi-card server as a computing node according to the inference test target of the model to be tested and the network communication parameters of each target multi-card server includes: Based on the inference test target and the bandwidth parameters of the target multi-card server, each target multi-card server is divided into computing nodes of different priority levels; the computing nodes include primary priority nodes and secondary priority nodes; For the first computing nodes of the same level, a first node path is established based on the latency parameters between the nodes, and the first node path is optimized based on the bandwidth utilization rate between the nodes to obtain the node main path; For the first computing node and the second computing node of different levels, after determining that a direct connection branch needs to be established based on the throughput between the nodes, a second node path is established according to the latency parameters between the first computing node and the second computing node and the number of intermediate computing nodes passed; Based on the node main path between the first computing nodes and the second node path between the first computing node and the second computing node, the multi-machine multi-card environment is constructed.

3. The model distributed inference test system based on multiple machines and multiple cards according to claim 2, characterized in that The network communication parameters further include fault tolerance parameters and load balancing parameters; The constructing the multi-machine multi-card environment based on the node main path between the first computing nodes and the second node path between the first computing node and the second computing node includes: For the first computing nodes of the same level, the target computing nodes are determined based on the fault tolerance parameters and load balancing parameters between the nodes; Determine a second intermediate computing node based on the latency parameters between the target computing node and each first intermediate computing node; the first intermediate computing nodes are the nodes in the first computing nodes except the target computing node; Establish a node backup path based on the target computing node and the second intermediate computing node; Construct the multi-machine multi-GPU environment based on the node main paths and node backup paths among the first computing nodes, and the second node paths between the first computing nodes and the second computing nodes.

4. The multi-machine multi-card based model distributed inference test system according to claim 3, wherein, The model testing strategy includes a data parallel strategy; the data parallel strategy represents dividing the test data set onto the computing GPUs of each computing node for backpropagation calculation and cross-node collaborative calculation; According to the data parallel strategy and the test data set corresponding to the model to be tested, start the distributed inference test of the model to be tested in the multi-machine multi-GPU environment, and obtain the model inference metrics of the data parallel strategy in different inference test stages, including: For each primary priority node in each inference test stage, based on the number of secondary priority nodes connected by the second node path, divide the corresponding number of test data sets onto the computing GPUs of each primary priority node and the computing GPUs of the secondary priority nodes connected to each primary priority node; Based on the test data sets in each computing GPU, start the forward propagation calculation of the model to be tested, and obtain the forward propagation results of each computing GPU; Based on the second node path, integrate the first forward propagation results of the computing GPUs in each primary priority node and the second forward propagation results of the computing GPUs of the secondary priority nodes connected to each primary priority node to obtain the integrated forward propagation result; Based on the integrated forward propagation results of each primary priority node in each inference test stage, determine the model inference metrics of the data parallel strategy in different inference test stages.

5. The model distributed inference test system based on multiple machines and multiple cards according to claim 4, characterized in that, The determining the model inference metrics of the data parallel strategy in different inference test stages based on the integrated forward propagation results of each primary priority node in each inference test stage includes: For each primary priority node in each inference test stage, based on the integrated forward propagation result, start the backpropagation calculation of the model to be tested according to the test data sets in each computing GPU, and obtain the gradient values of each model parameter output by each computing GPU; Based on the second node path, integrate the gradient values of the computing GPUs in each primary priority node and the gradient values of the computing GPUs of the secondary priority nodes connected to each primary priority node to obtain the integrated gradient value result; Fuse the integrated gradient value results of each primary priority node according to the node main path or node backup path to obtain the final gradient value result; Based on the final gradient value results of each inference test stage, update the model parameters in the model to be tested to obtain the model inference metrics in each inference test stage.

6. The model distributed inference test system based on multiple machines and multiple cards according to claim 3, characterized in that The model testing strategy includes a model parallel strategy; the model parallel strategy characterizes that different layers are allocated to the computing cards of different computing nodes according to the model structure for collaborative model inference; according to the model parallel strategy, the distributed inference test of the model to be tested is started in the multi-machine multi-card environment, and the model inference metrics of the model parallel strategy in different inference test phases are obtained, including: For each inference test phase, a corresponding number of convolutional layers are allocated to the computing cards of each primary priority node, and a corresponding number of fully connected layers and layer pooling layers are allocated to the computing cards of the first priority node and the second priority node connected to the primary priority node through the second node path; Based on the test dataset, the inference test of each computing card is started, and the inference test results of the computing cards of each primary priority node are transmitted and integrated based on the node main path or the node backup path to obtain the convolutional layer inference results, and the convolutional layer inference results are transmitted to the first priority node based on the second node path; The inference test results of the computing cards of each first priority node are transmitted and integrated based on the node main path or the node backup path to obtain the fully connected layer inference results; The fully connected layer inference results are transmitted to the second priority node based on the second node path, and the inference test results of the computing cards of each second priority node are transmitted and integrated based on the node main path or the node backup path to obtain the layer pooling layer inference results, and the inference test of each inference test phase is completed; Based on the inference test metrics during the inference test of the corresponding layer in each inference test phase, the model inference metrics of the model parallel strategy in different inference test phases are determined.

7. The model distributed inference test system based on multi-machine and multi-card according to any one of claims 1 to 6, characterized in that The test efficiency is evaluated according to the calculation speedup curves of each model test strategy in different inference test phases to obtain the optimal test strategy of the model to be tested, including: For each inference test phase, the curve intersection point between the first calculation speedup curve of the data parallel strategy and the second calculation speedup curve of the model parallel strategy is obtained; Based on the curve intersection point, the calculation speedups in the first calculation speedup curve and the second calculation speedup curve are compared to obtain a comparison result; If the comparison result is that the calculation speedup of the first calculation speedup curve is greater than the calculation speedup of the second calculation speedup curve, the data parallel strategy is determined as the target parallel strategy; If the comparison result is that the calculation speedup of the first calculation speedup curve is less than the calculation speedup of the second calculation speedup curve, the model parallel strategy is determined as the target parallel strategy; Based on the target parallel strategy of each inference test phase, a strategy combination is performed to obtain the optimal test strategy of the model to be tested.

8. A method for distributed inference testing of a model based on multiple machines and multiple cards, implemented based on the distributed inference testing system of the model based on multiple machines and multiple cards according to any one of claims 1 to 7, characterized in that, The model distributed inference test method based on multi-machine multi-card includes: According to the task complexity and task quantity of the model to be tested, and combining the server computing capabilities of each multi-card server in the multi-card server library, the target multi-card server is obtained; Taking the target multi - GPU server as a computing node, a multi - machine and multi - GPU environment is constructed according to the inference test objective of the to - be - tested model and the network communication parameters of each target multi - GPU server. According to different model test strategies, the distributed inference test of the to - be - tested model is started in the multi - machine and multi - GPU environment based on the test data set corresponding to the to - be - tested model, and the model inference metrics at different inference test stages for each model test strategy are obtained. The model inference metrics at different inference test stages for each model test strategy are evaluated to obtain the calculation speed - up ratio curves at different inference test stages for each model test strategy. The test efficiency is evaluated according to the calculation speed - up ratio curves at different inference test stages for each model test strategy, and the optimal test strategy for the to - be - tested model is obtained.

9. An electronic device, comprising: A memory and a processor, characterized in that a computer software program is stored on the memory, and when the processor reads and executes the computer software program, the method for distributed inference testing of a model based on multi - machine and multi - GPU as claimed in claim 8 is implemented.

10. A non-transitory computer-readable storage medium, characterized in that, A computer software program is stored in the storage medium, and when the computer software program is executed by the processor, the method for distributed inference testing of a model based on multi - machine and multi - GPU as claimed in claim 8 is implemented.